# voice-pet-pipeline Specification ## Purpose TBD - created by archiving change add-voice-pet-pipeline. Update Purpose after archive. ## Requirements ### Requirement: Python desktop pet runtime The system SHALL be specified as a local Python desktop pet application that owns the voice pipeline, desktop pet state, local audio input, and local audio output in the first version. #### Scenario: App starts in local desktop mode - **WHEN** the user starts the future application on macOS - **THEN** the system SHALL initialize as a local desktop pet process rather than a remote service or browser-only tool #### Scenario: Implementation follows OpenSpec artifacts - **WHEN** this OpenSpec change is implemented - **THEN** the repository SHALL contain Python source code, automated tests, validation commands, and project-local pet assets aligned with the planning artifacts ### Requirement: Local microphone and speaker transport The system SHALL use the local microphone as the first-version input Transport and the local system speaker as the first-version output Transport. #### Scenario: Microphone input is available - **WHEN** the configured microphone is available and permitted - **THEN** the Transport SHALL provide streaming audio frames to the wakeword and VAD stages #### Scenario: Speaker output is available - **WHEN** TTS returns a valid audio segment and the configured speaker is available - **THEN** the Transport SHALL play the audio segment through the local speaker #### Scenario: Audio device is unavailable - **WHEN** the microphone or speaker is missing, denied, or unsupported - **THEN** the system SHALL expose a recoverable Transport error with a stable error code ### Requirement: Wake word detection The system SHALL listen locally for the Chinese wake word “小杰小杰” before accepting user speech for a conversation turn. #### Scenario: Wake word is detected - **WHEN** the user says “小杰小杰” and the wakeword provider returns confidence above the configured threshold - **THEN** the pipeline SHALL transition from wake listening to speech detection #### Scenario: Wake word is not detected - **WHEN** background speech or noise does not match “小杰小杰” - **THEN** the pipeline SHALL remain in wake listening and SHALL NOT invoke STT, LLM, or TTS #### Scenario: Wake model fails - **WHEN** the wakeword provider cannot load or process audio - **THEN** the system SHALL report a wakeword error and SHALL NOT crash the desktop pet process ### Requirement: VAD speech endpoint detection After wakeword detection, the system SHALL use VAD to identify when the user starts and stops speaking. #### Scenario: User begins speaking - **WHEN** VAD detects continuous speech above the configured start threshold - **THEN** the pipeline SHALL begin recording the current utterance #### Scenario: User stops speaking - **WHEN** VAD detects continuous silence above the configured end threshold - **THEN** the pipeline SHALL close the current audio segment and send it to STT #### Scenario: User says nothing after wakeword - **WHEN** no speech is detected before the configured no-speech timeout - **THEN** the pipeline SHALL return to wake listening without invoking STT or LLM #### Scenario: Recording exceeds maximum duration - **WHEN** speech continues beyond the configured maximum utterance duration - **THEN** the pipeline SHALL end the segment, mark the end reason, and continue to STT with the captured audio ### Requirement: Local STT transcription The system SHALL transcribe the captured user utterance through a local STT provider, with `sherpa-onnx` as the recommended first implementation candidate. #### Scenario: STT succeeds - **WHEN** STT returns non-empty text for a captured audio segment - **THEN** the pipeline SHALL add the trimmed text as a user message to the conversation context #### Scenario: STT returns empty text - **WHEN** STT returns empty text, punctuation-only text, or an invalid transcript - **THEN** the pipeline SHALL skip LLM invocation and return to wake listening with a recoverable status #### Scenario: STT provider fails - **WHEN** the STT provider raises an error or cannot load its local model - **THEN** the system SHALL emit an STT error code and SHALL recover to a state where future wake attempts are possible ### Requirement: Conversation context management The system SHALL maintain an in-memory conversation context for the active desktop pet session. #### Scenario: User transcript is accepted - **WHEN** a valid STT transcript is produced - **THEN** the system SHALL append it to the context as a user message before invoking the LLM #### Scenario: Assistant reply completes - **WHEN** the LLM and TTS stages complete a reply - **THEN** the system SHALL append the assistant text to the context #### Scenario: Context exceeds configured budget - **WHEN** the context exceeds the configured message, character, or token budget - **THEN** the system SHALL preserve the system prompt and most recent conversation turns while removing older ordinary messages ### Requirement: Cloud LLM streaming reply The system SHALL use a cloud LLM provider for reply generation, with OpenAI Responses API streaming as the recommended first implementation. #### Scenario: LLM stream begins - **WHEN** the LLM provider receives valid context messages and configuration - **THEN** it SHALL stream reply deltas rather than waiting for the full reply before producing output #### Scenario: LLM model is configured - **WHEN** the system starts - **THEN** the LLM model name SHALL be read from configuration and SHALL NOT be hard-coded in pipeline logic #### Scenario: LLM request fails - **WHEN** the LLM provider times out, is rate limited, lacks credentials, or encounters a network error - **THEN** the pipeline SHALL emit a structured LLM error and transition through error recovery to wake listening ### Requirement: Local TTS synthesis and playback The system SHALL synthesize assistant replies through a local TTS provider, with `sherpa-onnx` as the recommended first implementation candidate. #### Scenario: Sentence is ready for speech - **WHEN** the LLM stream produces a complete sentence or configured text chunk - **THEN** the TTS provider SHALL synthesize that text into an audio segment #### Scenario: First TTS segment is ready - **WHEN** the first synthesized audio segment is available - **THEN** the output Transport SHALL begin playback without waiting for every remaining LLM delta #### Scenario: TTS fails - **WHEN** TTS returns empty audio or raises an error - **THEN** the pipeline SHALL emit a structured TTS error and SHALL recover without terminating the application ### Requirement: Pipeline state machine The system SHALL expose deterministic pipeline states for `idle`, `wake_listening`, `speech_detecting`, `recording`, `transcribing`, `thinking`, `speaking`, `interrupted`, and `error_recovering`. #### Scenario: Normal conversation turn - **WHEN** wakeword, VAD, STT, LLM, TTS, and playback all succeed - **THEN** the state sequence SHALL progress through wake listening, speech detection, recording, transcribing, thinking, speaking, and back to wake listening #### Scenario: Recoverable provider error - **WHEN** any provider fails during a conversation turn - **THEN** the state machine SHALL enter error recovery and then return to wake listening after cleanup #### Scenario: UI subscribes to state - **WHEN** the pipeline state changes - **THEN** the desktop pet UI SHALL be able to update its visible state without directly invoking audio or model providers ### Requirement: Desktop pet visual states The system SHALL define desktop pet visual states for idle, listening, recording, thinking, speaking, and error feedback. #### Scenario: Wakeword is detected - **WHEN** the pipeline enters speech detection or recording - **THEN** the desktop pet SHALL show a listening or recording visual state #### Scenario: LLM is generating - **WHEN** the pipeline enters thinking - **THEN** the desktop pet SHALL show a thinking visual state #### Scenario: TTS is playing - **WHEN** the pipeline enters speaking - **THEN** the desktop pet SHALL show a speaking visual state #### Scenario: Error occurs - **WHEN** the pipeline enters error recovery - **THEN** the desktop pet SHALL show a concise visible error state and then return to idle or wake listening when recovered ### Requirement: Generated pet asset specification The system SHALL specify that production desktop pet images are generated with the image generation skill when usable, or with a reproducible project-local fallback when image generation output is unavailable or fails validation, and saved inside the project as transparent PNG assets. #### Scenario: Image generation output is usable - **WHEN** pet images are generated with `imagegen` and pass role and transparency validation - **THEN** the images SHALL be processed for transparency, validated, and saved to a project asset directory #### Scenario: Image generation output is unavailable or invalid - **WHEN** image generation output is unavailable, inaccessible as a project file, or fails role/transparency validation - **THEN** the system SHALL use a reproducible project-local fallback to create transparent PNG pet state assets and SHALL validate them before use #### Scenario: Planning phase is executed - **WHEN** this OpenSpec change is implemented - **THEN** generated or generated-derived pet assets SHALL be saved inside the project and SHALL NOT be referenced from a temporary generation directory ### Requirement: Audio feedback suppression The system SHALL prevent the desktop pet's own TTS playback from being treated as a new user wakeword or speech input. #### Scenario: TTS playback is active - **WHEN** the output Transport is playing synthesized assistant speech - **THEN** wakeword detection and VAD input processing SHALL be suppressed or ignored for that playback window by default #### Scenario: User speaks during playback - **WHEN** the user speaks while TTS playback is active in the first version - **THEN** the system SHALL ignore that speech unless a future interruption feature is explicitly specified ### Requirement: Structured errors and logging The system SHALL represent stage failures with structured error codes and log state transitions, latency metrics, and provider names. #### Scenario: Stage transition occurs - **WHEN** the pipeline changes state - **THEN** the system SHALL log the turn ID, old state, new state, timestamp, and triggering event #### Scenario: Provider call completes - **WHEN** a provider call succeeds or fails - **THEN** the system SHALL log provider name, stage, duration, and error code if present #### Scenario: Sensitive data is present - **WHEN** logs are written - **THEN** the system SHALL NOT log API keys, authorization headers, or raw credentials ### Requirement: Performance targets The system SHALL define measurable first-version latency and reliability targets for wakeword response, endpoint detection, LLM first output, and speech playback. #### Scenario: Wakeword latency is measured - **WHEN** wakeword detection succeeds in a normal local environment - **THEN** the system SHALL target visible listening feedback within 800 ms #### Scenario: User stops speaking - **WHEN** the user stops speaking after a normal utterance - **THEN** VAD endpoint detection SHALL target transition to transcription within 900 ms #### Scenario: LLM begins responding - **WHEN** a valid transcript has been sent to the LLM provider - **THEN** the system SHALL target first playable reply text within 3.5 seconds #### Scenario: Normal reply is spoken - **WHEN** a normal Chinese question under 10 seconds is processed - **THEN** the system SHALL target the start of spoken playback within 5 seconds P95 after the user stops speaking ### Requirement: Security and privacy The system SHALL protect credentials and minimize audio persistence. #### Scenario: API key is required - **WHEN** the LLM provider needs an OpenAI API key - **THEN** the key SHALL be loaded from environment or local uncommitted configuration and SHALL NOT be committed #### Scenario: Audio is processed - **WHEN** user speech is captured for STT - **THEN** raw audio SHALL NOT be persisted by default #### Scenario: Cloud LLM is invoked - **WHEN** the system sends data to the cloud LLM - **THEN** it SHALL send transcript text and necessary conversation context, not raw microphone audio ### Requirement: Testability The system SHALL be designed so each stage can be tested with mock providers and file-based audio fixtures. #### Scenario: Pipeline is unit tested - **WHEN** mock providers produce deterministic events - **THEN** tests SHALL verify state transitions, context updates, and error recovery #### Scenario: Audio fixture is replayed - **WHEN** a file-based audio fixture containing “小杰小杰” and user speech is replayed in a test Transport - **THEN** the pipeline SHALL be testable without using a live microphone #### Scenario: OpenSpec planning validation runs - **WHEN** this planning change is complete - **THEN** `openspec validate add-voice-pet-pipeline --strict` and `openspec validate --all --strict` SHALL pass ### Requirement: Module-level git commit gates The system implementation process SHALL require an immediate Git commit after each major module or milestone is completed, where a major module means a top-level task group from `tasks.md` or a phase milestone from `proposal.md`. #### Scenario: Major module is completed - **WHEN** an implementer completes a top-level task group or milestone phase - **THEN** the implementer SHALL run the applicable build, tests, lint/type checks, and OpenSpec validation before committing #### Scenario: Build validation fails before commit - **WHEN** the applicable build, tests, lint/type checks, or OpenSpec validation fail - **THEN** the implementer SHALL NOT create the module commit until the failure is fixed and validation passes #### Scenario: Module commit is created - **WHEN** validation passes for the completed module - **THEN** the implementer SHALL immediately create a Git commit using the Chinese message format “[模块名]:完成[具体功能描述],包含[关键变更]” #### Scenario: Intermediate work remains after module completion - **WHEN** module-related changes remain unstaged or uncommitted after the module is completed - **THEN** the implementer SHALL include those changes in the module commit or explicitly separate unrelated external changes before continuing to the next module