14 KiB
voice-pet-pipeline Specification
Purpose
TBD - created by archiving change add-voice-pet-pipeline. Update Purpose after archive.
Requirements
Requirement: Python desktop pet runtime
The system SHALL be specified as a local Python desktop pet application that owns the voice pipeline, desktop pet state, local audio input, and local audio output in the first version.
Scenario: App starts in local desktop mode
- WHEN the user starts the future application on macOS
- THEN the system SHALL initialize as a local desktop pet process rather than a remote service or browser-only tool
Scenario: Implementation follows OpenSpec artifacts
- WHEN this OpenSpec change is implemented
- THEN the repository SHALL contain Python source code, automated tests, validation commands, and project-local pet assets aligned with the planning artifacts
Requirement: Local microphone and speaker transport
The system SHALL use the local microphone as the first-version input Transport and the local system speaker as the first-version output Transport.
Scenario: Microphone input is available
- WHEN the configured microphone is available and permitted
- THEN the Transport SHALL provide streaming audio frames to the wakeword and VAD stages
Scenario: Speaker output is available
- WHEN TTS returns a valid audio segment and the configured speaker is available
- THEN the Transport SHALL play the audio segment through the local speaker
Scenario: Audio device is unavailable
- WHEN the microphone or speaker is missing, denied, or unsupported
- THEN the system SHALL expose a recoverable Transport error with a stable error code
Requirement: Wake word detection
The system SHALL listen locally for the Chinese wake word “小杰小杰” before accepting user speech for a conversation turn.
Scenario: Wake word is detected
- WHEN the user says “小杰小杰” and the wakeword provider returns confidence above the configured threshold
- THEN the pipeline SHALL transition from wake listening to speech detection
Scenario: Wake word is not detected
- WHEN background speech or noise does not match “小杰小杰”
- THEN the pipeline SHALL remain in wake listening and SHALL NOT invoke STT, LLM, or TTS
Scenario: Wake model fails
- WHEN the wakeword provider cannot load or process audio
- THEN the system SHALL report a wakeword error and SHALL NOT crash the desktop pet process
Requirement: VAD speech endpoint detection
After wakeword detection, the system SHALL use VAD to identify when the user starts and stops speaking.
Scenario: User begins speaking
- WHEN VAD detects continuous speech above the configured start threshold
- THEN the pipeline SHALL begin recording the current utterance
Scenario: User stops speaking
- WHEN VAD detects continuous silence above the configured end threshold
- THEN the pipeline SHALL close the current audio segment and send it to STT
Scenario: User says nothing after wakeword
- WHEN no speech is detected before the configured no-speech timeout
- THEN the pipeline SHALL return to wake listening without invoking STT or LLM
Scenario: Recording exceeds maximum duration
- WHEN speech continues beyond the configured maximum utterance duration
- THEN the pipeline SHALL end the segment, mark the end reason, and continue to STT with the captured audio
Requirement: Local STT transcription
The system SHALL transcribe the captured user utterance through a local STT provider, with sherpa-onnx as the recommended first implementation candidate.
Scenario: STT succeeds
- WHEN STT returns non-empty text for a captured audio segment
- THEN the pipeline SHALL add the trimmed text as a user message to the conversation context
Scenario: STT returns empty text
- WHEN STT returns empty text, punctuation-only text, or an invalid transcript
- THEN the pipeline SHALL skip LLM invocation and return to wake listening with a recoverable status
Scenario: STT provider fails
- WHEN the STT provider raises an error or cannot load its local model
- THEN the system SHALL emit an STT error code and SHALL recover to a state where future wake attempts are possible
Requirement: Conversation context management
The system SHALL maintain an in-memory conversation context for the active desktop pet session.
Scenario: User transcript is accepted
- WHEN a valid STT transcript is produced
- THEN the system SHALL append it to the context as a user message before invoking the LLM
Scenario: Assistant reply completes
- WHEN the LLM and TTS stages complete a reply
- THEN the system SHALL append the assistant text to the context
Scenario: Context exceeds configured budget
- WHEN the context exceeds the configured message, character, or token budget
- THEN the system SHALL preserve the system prompt and most recent conversation turns while removing older ordinary messages
Requirement: Cloud LLM streaming reply
The system SHALL use a cloud LLM provider for reply generation, with OpenAI Responses API streaming as the recommended first implementation.
Scenario: LLM stream begins
- WHEN the LLM provider receives valid context messages and configuration
- THEN it SHALL stream reply deltas rather than waiting for the full reply before producing output
Scenario: LLM model is configured
- WHEN the system starts
- THEN the LLM model name SHALL be read from configuration and SHALL NOT be hard-coded in pipeline logic
Scenario: LLM request fails
- WHEN the LLM provider times out, is rate limited, lacks credentials, or encounters a network error
- THEN the pipeline SHALL emit a structured LLM error and transition through error recovery to wake listening
Requirement: Local TTS synthesis and playback
The system SHALL synthesize assistant replies through a local TTS provider, with sherpa-onnx as the recommended first implementation candidate.
Scenario: Sentence is ready for speech
- WHEN the LLM stream produces a complete sentence or configured text chunk
- THEN the TTS provider SHALL synthesize that text into an audio segment
Scenario: First TTS segment is ready
- WHEN the first synthesized audio segment is available
- THEN the output Transport SHALL begin playback without waiting for every remaining LLM delta
Scenario: TTS fails
- WHEN TTS returns empty audio or raises an error
- THEN the pipeline SHALL emit a structured TTS error and SHALL recover without terminating the application
Requirement: Pipeline state machine
The system SHALL expose deterministic pipeline states for idle, wake_listening, speech_detecting, recording, transcribing, thinking, speaking, interrupted, and error_recovering.
Scenario: Normal conversation turn
- WHEN wakeword, VAD, STT, LLM, TTS, and playback all succeed
- THEN the state sequence SHALL progress through wake listening, speech detection, recording, transcribing, thinking, speaking, and back to wake listening
Scenario: Recoverable provider error
- WHEN any provider fails during a conversation turn
- THEN the state machine SHALL enter error recovery and then return to wake listening after cleanup
Scenario: UI subscribes to state
- WHEN the pipeline state changes
- THEN the desktop pet UI SHALL be able to update its visible state without directly invoking audio or model providers
Requirement: Desktop pet visual states
The system SHALL define desktop pet visual states for idle, listening, recording, thinking, speaking, and error feedback.
Scenario: Wakeword is detected
- WHEN the pipeline enters speech detection or recording
- THEN the desktop pet SHALL show a listening or recording visual state
Scenario: LLM is generating
- WHEN the pipeline enters thinking
- THEN the desktop pet SHALL show a thinking visual state
Scenario: TTS is playing
- WHEN the pipeline enters speaking
- THEN the desktop pet SHALL show a speaking visual state
Scenario: Error occurs
- WHEN the pipeline enters error recovery
- THEN the desktop pet SHALL show a concise visible error state and then return to idle or wake listening when recovered
Requirement: Generated pet asset specification
The system SHALL specify that production desktop pet images are generated with the image generation skill when usable, or with a reproducible project-local fallback when image generation output is unavailable or fails validation, and saved inside the project as transparent PNG assets.
Scenario: Image generation output is usable
- WHEN pet images are generated with
imagegenand pass role and transparency validation - THEN the images SHALL be processed for transparency, validated, and saved to a project asset directory
Scenario: Image generation output is unavailable or invalid
- WHEN image generation output is unavailable, inaccessible as a project file, or fails role/transparency validation
- THEN the system SHALL use a reproducible project-local fallback to create transparent PNG pet state assets and SHALL validate them before use
Scenario: Planning phase is executed
- WHEN this OpenSpec change is implemented
- THEN generated or generated-derived pet assets SHALL be saved inside the project and SHALL NOT be referenced from a temporary generation directory
Requirement: Audio feedback suppression
The system SHALL prevent the desktop pet's own TTS playback from being treated as a new user wakeword or speech input.
Scenario: TTS playback is active
- WHEN the output Transport is playing synthesized assistant speech
- THEN wakeword detection and VAD input processing SHALL be suppressed or ignored for that playback window by default
Scenario: User speaks during playback
- WHEN the user speaks while TTS playback is active in the first version
- THEN the system SHALL ignore that speech unless a future interruption feature is explicitly specified
Requirement: Structured errors and logging
The system SHALL represent stage failures with structured error codes and log state transitions, latency metrics, and provider names.
Scenario: Stage transition occurs
- WHEN the pipeline changes state
- THEN the system SHALL log the turn ID, old state, new state, timestamp, and triggering event
Scenario: Provider call completes
- WHEN a provider call succeeds or fails
- THEN the system SHALL log provider name, stage, duration, and error code if present
Scenario: Sensitive data is present
- WHEN logs are written
- THEN the system SHALL NOT log API keys, authorization headers, or raw credentials
Requirement: Performance targets
The system SHALL define measurable first-version latency and reliability targets for wakeword response, endpoint detection, LLM first output, and speech playback.
Scenario: Wakeword latency is measured
- WHEN wakeword detection succeeds in a normal local environment
- THEN the system SHALL target visible listening feedback within 800 ms
Scenario: User stops speaking
- WHEN the user stops speaking after a normal utterance
- THEN VAD endpoint detection SHALL target transition to transcription within 900 ms
Scenario: LLM begins responding
- WHEN a valid transcript has been sent to the LLM provider
- THEN the system SHALL target first playable reply text within 3.5 seconds
Scenario: Normal reply is spoken
- WHEN a normal Chinese question under 10 seconds is processed
- THEN the system SHALL target the start of spoken playback within 5 seconds P95 after the user stops speaking
Requirement: Security and privacy
The system SHALL protect credentials and minimize audio persistence.
Scenario: API key is required
- WHEN the LLM provider needs an OpenAI API key
- THEN the key SHALL be loaded from environment or local uncommitted configuration and SHALL NOT be committed
Scenario: Audio is processed
- WHEN user speech is captured for STT
- THEN raw audio SHALL NOT be persisted by default
Scenario: Cloud LLM is invoked
- WHEN the system sends data to the cloud LLM
- THEN it SHALL send transcript text and necessary conversation context, not raw microphone audio
Requirement: Testability
The system SHALL be designed so each stage can be tested with mock providers and file-based audio fixtures.
Scenario: Pipeline is unit tested
- WHEN mock providers produce deterministic events
- THEN tests SHALL verify state transitions, context updates, and error recovery
Scenario: Audio fixture is replayed
- WHEN a file-based audio fixture containing “小杰小杰” and user speech is replayed in a test Transport
- THEN the pipeline SHALL be testable without using a live microphone
Scenario: OpenSpec planning validation runs
- WHEN this planning change is complete
- THEN
openspec validate add-voice-pet-pipeline --strictandopenspec validate --all --strictSHALL pass
Requirement: Module-level git commit gates
The system implementation process SHALL require an immediate Git commit after each major module or milestone is completed, where a major module means a top-level task group from tasks.md or a phase milestone from proposal.md.
Scenario: Major module is completed
- WHEN an implementer completes a top-level task group or milestone phase
- THEN the implementer SHALL run the applicable build, tests, lint/type checks, and OpenSpec validation before committing
Scenario: Build validation fails before commit
- WHEN the applicable build, tests, lint/type checks, or OpenSpec validation fail
- THEN the implementer SHALL NOT create the module commit until the failure is fixed and validation passes
Scenario: Module commit is created
- WHEN validation passes for the completed module
- THEN the implementer SHALL immediately create a Git commit using the Chinese message format “[模块名]:完成[具体功能描述],包含[关键变更]”
Scenario: Intermediate work remains after module completion
- WHEN module-related changes remain unstaged or uncommitted after the module is completed
- THEN the implementer SHALL include those changes in the module commit or explicitly separate unrelated external changes before continuing to the next module