# voice-pet-pipeline Specification ## Purpose TBD - created by archiving change add-voice-pet-pipeline. Update Purpose after archive. ## Requirements ### Requirement: Python desktop pet runtime The system SHALL be specified as a local Python desktop pet application that owns the voice pipeline, desktop pet state, local audio input, and local audio output in the first version. #### Scenario: App starts in local desktop mode - **WHEN** the user starts the future application on macOS - **THEN** the system SHALL initialize as a local desktop pet process rather than a remote service or browser-only tool #### Scenario: Implementation follows OpenSpec artifacts - **WHEN** this OpenSpec change is implemented - **THEN** the repository SHALL contain Python source code, automated tests, validation commands, and project-local pet assets aligned with the planning artifacts ### Requirement: Local microphone and speaker transport The system SHALL use the local microphone as the first-version input Transport and the local system speaker as the first-version output Transport, and `run-live` SHALL exercise those real local devices by default. #### Scenario: Microphone input is available - **WHEN** the configured microphone is available and permitted - **THEN** the live Transport SHALL provide streaming audio frames to wake detection, VAD, and STT stages #### Scenario: Speaker output is available - **WHEN** TTS returns a valid local audio segment or audio file and the configured speaker is available - **THEN** the live Transport or macOS playback command SHALL play the audio through the local speaker #### Scenario: Audio device is unavailable - **WHEN** the microphone or speaker is missing, denied, unsupported, or inaccessible through `sounddevice` - **THEN** the system SHALL expose a recoverable Transport error with a stable error code and SHALL NOT silently fall back to fixture audio in live mode ### Requirement: Wake word detection The system SHALL listen locally for the Chinese wake word “小杰小杰” before accepting user speech for a conversation turn. #### Scenario: Wake word is detected - **WHEN** the user says “小杰小杰” and the wakeword provider returns confidence above the configured threshold - **THEN** the pipeline SHALL transition from wake listening to speech detection #### Scenario: Wake word is not detected - **WHEN** background speech or noise does not match “小杰小杰” - **THEN** the pipeline SHALL remain in wake listening and SHALL NOT invoke STT, LLM, or TTS #### Scenario: Wake model fails - **WHEN** the wakeword provider cannot load or process audio - **THEN** the system SHALL report a wakeword error and SHALL NOT crash the desktop pet process ### Requirement: VAD speech endpoint detection After wakeword detection, the system SHALL use VAD to identify when the user starts and stops speaking. #### Scenario: User begins speaking - **WHEN** VAD detects continuous speech above the configured start threshold - **THEN** the pipeline SHALL begin recording the current utterance #### Scenario: User stops speaking - **WHEN** VAD detects continuous silence above the configured end threshold - **THEN** the pipeline SHALL close the current audio segment and send it to STT #### Scenario: User says nothing after wakeword - **WHEN** no speech is detected before the configured no-speech timeout - **THEN** the pipeline SHALL return to wake listening without invoking STT or LLM #### Scenario: Recording exceeds maximum duration - **WHEN** speech continues beyond the configured maximum utterance duration - **THEN** the pipeline SHALL end the segment, mark the end reason, and continue to STT with the captured audio ### Requirement: Local STT transcription The system SHALL transcribe captured user utterances through the configured STT provider; cloud mode SHALL use NewAPI-compatible ASR and local mode SHALL use project-local `sherpa-onnx` model assets under `models/`. #### Scenario: STT succeeds - **WHEN** configured STT returns non-empty text for a captured live audio segment - **THEN** the pipeline SHALL add the trimmed text as a user message to the current process conversation context #### Scenario: STT returns empty text - **WHEN** STT returns empty text, punctuation-only text, or an invalid transcript - **THEN** the live runtime SHALL skip LLM invocation and return to standby with a recoverable status #### Scenario: STT provider fails - **WHEN** the STT provider raises an error, cloud ASR fails, or local model loading fails - **THEN** the system SHALL emit an STT or model error code and SHALL recover to a state where future wake attempts are possible if startup can continue safely ### Requirement: Conversation context management The system SHALL maintain an in-memory conversation context for the active desktop pet session, and live repeated voice runtime SHALL define that session as the current `run-live` process only. #### Scenario: User transcript is accepted - **WHEN** a valid STT transcript is produced during a live turn - **THEN** the system SHALL append it to the current process context as a user message before invoking the LLM #### Scenario: Assistant reply completes - **WHEN** the LLM and TTS stages complete a non-empty reply during a live turn - **THEN** the system SHALL append the assistant text to the current process context #### Scenario: Context exceeds configured budget - **WHEN** the context exceeds the configured message or character budget - **THEN** the system SHALL preserve the system prompt and most recent conversation turns while removing older ordinary messages #### Scenario: Runtime exits - **WHEN** the `run-live` process exits - **THEN** the system SHALL discard the context and SHALL NOT persist it across process restarts ### Requirement: Cloud LLM streaming reply The system SHALL use a cloud LLM provider for reply generation, with an OpenAI/NewAPI-compatible chat completions endpoint as the first-version runtime target. #### Scenario: LLM stream begins - **WHEN** the LLM provider receives valid context messages and configuration - **THEN** it SHALL stream reply deltas rather than waiting for the full reply before producing output #### Scenario: LLM model is configured - **WHEN** the system starts - **THEN** the LLM base URL, API key, and model name SHALL be read from `.env` or local uncommitted configuration and SHALL NOT be hard-coded in pipeline logic #### Scenario: LLM request fails - **WHEN** the LLM provider times out, is rate limited, lacks credentials, or encounters a network error - **THEN** the pipeline SHALL emit a structured LLM error and transition through error recovery to wake listening ### Requirement: Local TTS synthesis and playback The system SHALL synthesize assistant replies through the configured TTS provider and SHALL play synthesized speech through the local system output path; cloud mode SHALL use NewAPI-compatible TTS and local mode SHALL use a local playback-capable TTS provider. #### Scenario: Reply text is ready for speech - **WHEN** the LLM produces a non-empty reply for a live turn - **THEN** the TTS provider SHALL synthesize that text into locally playable speech #### Scenario: Playback starts - **WHEN** the first synthesized audio output is available - **THEN** the output Transport or macOS playback command SHALL begin playback through the local speaker #### Scenario: TTS fails - **WHEN** TTS returns empty audio, fails to create an audio file, or playback command fails - **THEN** the live runtime SHALL emit a structured TTS or playback error and SHALL recover without terminating default loop mode ### Requirement: Pipeline state machine The system SHALL expose deterministic pipeline states for `idle`, `wake_listening`, `speech_detecting`, `recording`, `transcribing`, `thinking`, `speaking`, `interrupted`, and `error_recovering`. #### Scenario: Normal conversation turn - **WHEN** wakeword, VAD, STT, LLM, TTS, and playback all succeed - **THEN** the state sequence SHALL progress through wake listening, speech detection, recording, transcribing, thinking, speaking, and back to wake listening #### Scenario: Recoverable provider error - **WHEN** any provider fails during a conversation turn - **THEN** the state machine SHALL enter error recovery and then return to wake listening after cleanup #### Scenario: UI subscribes to state - **WHEN** the pipeline state changes - **THEN** the desktop pet UI SHALL be able to update its visible state without directly invoking audio or model providers ### Requirement: Desktop pet visual states The system SHALL define desktop pet visual states for idle, listening, recording, thinking, speaking, and error feedback. #### Scenario: Wakeword is detected - **WHEN** the pipeline enters speech detection or recording - **THEN** the desktop pet SHALL show a listening or recording visual state #### Scenario: LLM is generating - **WHEN** the pipeline enters thinking - **THEN** the desktop pet SHALL show a thinking visual state #### Scenario: TTS is playing - **WHEN** the pipeline enters speaking - **THEN** the desktop pet SHALL show a speaking visual state #### Scenario: Error occurs - **WHEN** the pipeline enters error recovery - **THEN** the desktop pet SHALL show a concise visible error state and then return to idle or wake listening when recovered ### Requirement: Generated pet asset specification The system SHALL specify that production desktop pet images are generated with the image generation skill when usable, or with a reproducible project-local fallback when image generation output is unavailable or fails validation, and saved inside the project as transparent PNG assets. #### Scenario: Image generation output is usable - **WHEN** pet images are generated with `imagegen` and pass role and transparency validation - **THEN** the images SHALL be processed for transparency, validated, and saved to a project asset directory #### Scenario: Image generation output is unavailable or invalid - **WHEN** image generation output is unavailable, inaccessible as a project file, or fails role/transparency validation - **THEN** the system SHALL use a reproducible project-local fallback to create transparent PNG pet state assets and SHALL validate them before use #### Scenario: Planning phase is executed - **WHEN** this OpenSpec change is implemented - **THEN** generated or generated-derived pet assets SHALL be saved inside the project and SHALL NOT be referenced from a temporary generation directory ### Requirement: Audio feedback suppression The system SHALL prevent the desktop pet's own TTS playback from being treated as a new user wakeword or speech input. #### Scenario: TTS playback is active - **WHEN** the output Transport is playing synthesized assistant speech - **THEN** wakeword detection and VAD input processing SHALL be suppressed or ignored for that playback window by default #### Scenario: User speaks during playback - **WHEN** the user speaks while TTS playback is active in the first version - **THEN** the system SHALL ignore that speech unless a future interruption feature is explicitly specified ### Requirement: Structured errors and logging The system SHALL represent stage failures with structured error codes and log state transitions, latency metrics, and provider names. #### Scenario: Stage transition occurs - **WHEN** the pipeline changes state - **THEN** the system SHALL log the turn ID, old state, new state, timestamp, and triggering event #### Scenario: Provider call completes - **WHEN** a provider call succeeds or fails - **THEN** the system SHALL log provider name, stage, duration, and error code if present #### Scenario: Sensitive data is present - **WHEN** logs are written - **THEN** the system SHALL NOT log API keys, authorization headers, or raw credentials ### Requirement: Performance targets The system SHALL define measurable first-version latency and reliability targets for wakeword response, endpoint detection, LLM first output, and speech playback. #### Scenario: Wakeword latency is measured - **WHEN** wakeword detection succeeds in a normal local environment - **THEN** the system SHALL target visible listening feedback within 800 ms #### Scenario: User stops speaking - **WHEN** the user stops speaking after a normal utterance - **THEN** VAD endpoint detection SHALL target transition to transcription within 900 ms #### Scenario: LLM begins responding - **WHEN** a valid transcript has been sent to the LLM provider - **THEN** the system SHALL target first playable reply text within 3.5 seconds #### Scenario: Normal reply is spoken - **WHEN** a normal Chinese question under 10 seconds is processed - **THEN** the system SHALL target the start of spoken playback within 5 seconds P95 after the user stops speaking ### Requirement: Security and privacy The system SHALL protect credentials and minimize audio persistence. #### Scenario: API key is required - **WHEN** the LLM provider needs an OpenAI API key - **THEN** the key SHALL be loaded from `.env` or local uncommitted configuration and SHALL NOT be committed #### Scenario: Audio is processed - **WHEN** user speech is captured for STT - **THEN** raw audio SHALL NOT be persisted by default #### Scenario: Cloud LLM is invoked - **WHEN** the system sends data to the cloud LLM - **THEN** it SHALL send transcript text and necessary conversation context, not raw microphone audio ### Requirement: Testability The system SHALL be designed so each stage can be tested with mock providers, file-based audio fixtures, and fake live runtime components without requiring real devices in automated tests. #### Scenario: Repeated runtime is unit tested - **WHEN** fake live providers produce two deterministic turns - **THEN** tests SHALL verify two STT calls, two LLM calls, two TTS calls, two playback calls, and final return to standby #### Scenario: Temporary context is unit tested - **WHEN** fake live providers run two turns in one runtime instance - **THEN** tests SHALL verify the second LLM request includes the first turn user and assistant messages #### Scenario: Process-local context is unit tested - **WHEN** a second runtime instance is created after a first instance has conversation history - **THEN** tests SHALL verify the second instance starts with no user/assistant history #### Scenario: OpenSpec planning validation runs - **WHEN** this change is complete before implementation - **THEN** `openspec validate complete-live-repeat-voice-runtime --strict` and `openspec validate --all --strict` SHALL pass ### Requirement: Module-level git commit gates The system implementation process SHALL require an immediate Git commit after each major module or milestone is completed, where a major module means a top-level task group from `tasks.md` or a phase milestone from `proposal.md`. #### Scenario: Major module is completed - **WHEN** an implementer completes a top-level task group or milestone phase - **THEN** the implementer SHALL run the applicable build, tests, lint/type checks, and OpenSpec validation before committing #### Scenario: Build validation fails before commit - **WHEN** the applicable build, tests, lint/type checks, or OpenSpec validation fail - **THEN** the implementer SHALL NOT create the module commit until the failure is fixed and validation passes #### Scenario: Module commit is created - **WHEN** validation passes for the completed module - **THEN** the implementer SHALL immediately create a Git commit using the Chinese message format “[模块名]:完成[具体功能描述],包含[关键变更]” #### Scenario: Intermediate work remains after module completion - **WHEN** module-related changes remain unstaged or uncommitted after the module is completed - **THEN** the implementer SHALL include those changes in the module commit or explicitly separate unrelated external changes before continuing to the next module ### Requirement: Live repeat voice runtime The system SHALL provide a `run-live` command that performs real repeated voice conversation with local microphone input, configured speech recognition and speech synthesis providers, cloud LLM reply generation, local speaker playback, and automatic return to standby. #### Scenario: Live runtime starts in standby - **WHEN** the user runs `PYTHONPATH=src python3.11 -m owner_voice_pet run-live` - **THEN** the system SHALL load `.env`, validate required live dependencies, initialize local audio/model providers, and enter a standby listening loop #### Scenario: Live runtime completes two turns - **WHEN** the user wakes the system with “小杰小杰”, asks a question, hears the reply, then wakes it again and asks another question - **THEN** the system SHALL complete wake, recording, STT, LLM, TTS, playback for both turns and SHALL return to standby after each turn #### Scenario: Once mode completes one turn - **WHEN** the user runs `PYTHONPATH=src python3.11 -m owner_voice_pet run-live --once` - **THEN** the system SHALL run at most one wake-to-playback turn and exit after the turn completes or fails with a documented live error ### Requirement: Temporary in-process conversation history The live runtime SHALL maintain conversation history only in memory for the current process and SHALL include retained user/assistant history in later LLM requests during that same process. #### Scenario: Second turn uses first turn history - **WHEN** the first live turn appends a user transcript and an assistant reply - **AND** a second live turn sends an LLM request - **THEN** the second request SHALL include the first turn user message and assistant message unless configured context limits require truncation #### Scenario: New runtime starts empty - **WHEN** a new `run-live` process or new live runtime instance starts - **THEN** it SHALL NOT load user/assistant history from a previous process or previous runtime instance #### Scenario: Process exits - **WHEN** the live runtime exits for any reason - **THEN** conversation history SHALL be discarded and SHALL NOT be written to disk, database, logs, or model files ### Requirement: Local speech model management The system SHALL provide project-local speech model preparation and diagnostics for live VAD/STT operation, storing downloaded model artifacts under `models/` without committing them to Git. #### Scenario: Models are downloaded - **WHEN** the user runs `python3.11 scripts/download_speech_models.py --dir models` - **THEN** the script SHALL create or update a project-local model directory with the files required by the configured VAD/STT providers #### Scenario: Model check succeeds - **WHEN** required dependencies and model files are available - **THEN** `PYTHONPATH=src python3.11 -m owner_voice_pet model-check` SHALL exit successfully and report the model directory and checked providers #### Scenario: Model check fails - **WHEN** `sherpa-onnx` is unavailable, a model file is missing, or a model cannot be loaded - **THEN** `model-check` SHALL fail with a structured model error and SHALL NOT start live microphone listening ### Requirement: Configurable speech provider mode The system SHALL read `OWNER_SPEECH_PROVIDER` from `.env` to choose between cloud speech providers and local speech providers for first-version live runtime. #### Scenario: Cloud speech provider is selected - **WHEN** `OWNER_SPEECH_PROVIDER=cloud` - **THEN** live ASR SHALL use the configured cloud model from `OWNER_ASR_MODEL`, live TTS SHALL use the configured cloud model from `OWNER_TTS_MODEL`, and the implementation SHALL default those models to `mimo-v2.5-asr` and `mimo-v2.5-tts` #### Scenario: Local speech provider is selected - **WHEN** `OWNER_SPEECH_PROVIDER=local` - **THEN** live ASR/VAD SHALL use project-local speech model assets and local TTS SHALL use a local playback-capable provider #### Scenario: Speech provider is invalid - **WHEN** `OWNER_SPEECH_PROVIDER` is neither `cloud` nor `local` - **THEN** startup validation SHALL fail with a structured configuration error ### Requirement: Live audio device readiness The system SHALL provide live audio device diagnostics and SHALL use `sounddevice` for first-version real microphone and speaker access. #### Scenario: Device check succeeds - **WHEN** `sounddevice` can be imported and at least one input and one output device are available - **THEN** `PYTHONPATH=src python3.11 -m owner_voice_pet device-check` SHALL exit successfully and report usable audio devices #### Scenario: Device check fails - **WHEN** `sounddevice` is missing, device query fails, microphone permission is denied, or no usable input/output device exists - **THEN** `device-check` SHALL fail with a structured audio device error #### Scenario: Live runtime uses real devices - **WHEN** `run-live` starts successfully - **THEN** it SHALL use the local microphone as its default input and the local speaker or macOS playback command as its default output rather than fixture audio ### Requirement: Live terminal state reporting The live runtime SHALL emit concise Chinese terminal status messages for observable runtime states. #### Scenario: Normal turn status - **WHEN** a live turn succeeds - **THEN** terminal output SHALL include states equivalent to standby, wake hit, recording, transcribing, thinking, speaking, and returning to standby #### Scenario: Recoverable error status - **WHEN** a live turn encounters empty STT, LLM failure, TTS failure, or playback failure - **THEN** terminal output SHALL include the failing stage and a stable error code or recoverable explanation ### Requirement: Live runtime error recovery The live runtime SHALL recover from turn-level failures and continue listening in default loop mode. #### Scenario: LLM fails during default loop - **WHEN** the LLM provider times out, is rate limited, or returns a network error during a live turn - **THEN** the system SHALL report `LIVE_LLM_FAILED` or an equivalent structured error and SHALL return to standby without terminating the process #### Scenario: TTS or playback fails during default loop - **WHEN** local TTS or audio playback fails during a live turn - **THEN** the system SHALL report the failure and SHALL return to standby without terminating the process #### Scenario: Startup dependency fails - **WHEN** `.env`, required models, or audio devices are unavailable at startup - **THEN** the system SHALL exit with a documented non-zero status instead of entering a fake live loop