22 KiB
voice-pet-pipeline Specification
Purpose
TBD - created by archiving change add-voice-pet-pipeline. Update Purpose after archive.
Requirements
Requirement: Python desktop pet runtime
The system SHALL be specified as a local Python desktop pet application that owns the voice pipeline, desktop pet state, local audio input, and local audio output in the first version.
Scenario: App starts in local desktop mode
- WHEN the user starts the future application on macOS
- THEN the system SHALL initialize as a local desktop pet process rather than a remote service or browser-only tool
Scenario: Implementation follows OpenSpec artifacts
- WHEN this OpenSpec change is implemented
- THEN the repository SHALL contain Python source code, automated tests, validation commands, and project-local pet assets aligned with the planning artifacts
Requirement: Local microphone and speaker transport
The system SHALL use the local microphone as the first-version input Transport and the local system speaker as the first-version output Transport, and run-live SHALL exercise those real local devices by default.
Scenario: Microphone input is available
- WHEN the configured microphone is available and permitted
- THEN the live Transport SHALL provide streaming audio frames to wake detection, VAD, and STT stages
Scenario: Speaker output is available
- WHEN TTS returns a valid local audio segment or audio file and the configured speaker is available
- THEN the live Transport or macOS playback command SHALL play the audio through the local speaker
Scenario: Audio device is unavailable
- WHEN the microphone or speaker is missing, denied, unsupported, or inaccessible through
sounddevice - THEN the system SHALL expose a recoverable Transport error with a stable error code and SHALL NOT silently fall back to fixture audio in live mode
Requirement: Wake word detection
The system SHALL listen locally for the Chinese wake word “小杰小杰” before accepting user speech for a conversation turn.
Scenario: Wake word is detected
- WHEN the user says “小杰小杰” and the wakeword provider returns confidence above the configured threshold
- THEN the pipeline SHALL transition from wake listening to speech detection
Scenario: Wake word is not detected
- WHEN background speech or noise does not match “小杰小杰”
- THEN the pipeline SHALL remain in wake listening and SHALL NOT invoke STT, LLM, or TTS
Scenario: Wake model fails
- WHEN the wakeword provider cannot load or process audio
- THEN the system SHALL report a wakeword error and SHALL NOT crash the desktop pet process
Requirement: VAD speech endpoint detection
After wakeword detection, the system SHALL use VAD to identify when the user starts and stops speaking.
Scenario: User begins speaking
- WHEN VAD detects continuous speech above the configured start threshold
- THEN the pipeline SHALL begin recording the current utterance
Scenario: User stops speaking
- WHEN VAD detects continuous silence above the configured end threshold
- THEN the pipeline SHALL close the current audio segment and send it to STT
Scenario: User says nothing after wakeword
- WHEN no speech is detected before the configured no-speech timeout
- THEN the pipeline SHALL return to wake listening without invoking STT or LLM
Scenario: Recording exceeds maximum duration
- WHEN speech continues beyond the configured maximum utterance duration
- THEN the pipeline SHALL end the segment, mark the end reason, and continue to STT with the captured audio
Requirement: Local STT transcription
The system SHALL transcribe captured user utterances through the configured STT provider; cloud mode SHALL use NewAPI-compatible ASR and local mode SHALL use project-local sherpa-onnx model assets under models/.
Scenario: STT succeeds
- WHEN configured STT returns non-empty text for a captured live audio segment
- THEN the pipeline SHALL add the trimmed text as a user message to the current process conversation context
Scenario: STT returns empty text
- WHEN STT returns empty text, punctuation-only text, or an invalid transcript
- THEN the live runtime SHALL skip LLM invocation and return to standby with a recoverable status
Scenario: STT provider fails
- WHEN the STT provider raises an error, cloud ASR fails, or local model loading fails
- THEN the system SHALL emit an STT or model error code and SHALL recover to a state where future wake attempts are possible if startup can continue safely
Requirement: Conversation context management
The system SHALL maintain an in-memory conversation context for the active desktop pet session, and live repeated voice runtime SHALL define that session as the current run-live process only.
Scenario: User transcript is accepted
- WHEN a valid STT transcript is produced during a live turn
- THEN the system SHALL append it to the current process context as a user message before invoking the LLM
Scenario: Assistant reply completes
- WHEN the LLM and TTS stages complete a non-empty reply during a live turn
- THEN the system SHALL append the assistant text to the current process context
Scenario: Context exceeds configured budget
- WHEN the context exceeds the configured message or character budget
- THEN the system SHALL preserve the system prompt and most recent conversation turns while removing older ordinary messages
Scenario: Runtime exits
- WHEN the
run-liveprocess exits - THEN the system SHALL discard the context and SHALL NOT persist it across process restarts
Requirement: Cloud LLM streaming reply
The system SHALL use a cloud LLM provider for reply generation, with an OpenAI/NewAPI-compatible chat completions endpoint as the first-version runtime target.
Scenario: LLM stream begins
- WHEN the LLM provider receives valid context messages and configuration
- THEN it SHALL stream reply deltas rather than waiting for the full reply before producing output
Scenario: LLM model is configured
- WHEN the system starts
- THEN the LLM base URL, API key, and model name SHALL be read from
.envor local uncommitted configuration and SHALL NOT be hard-coded in pipeline logic
Scenario: LLM request fails
- WHEN the LLM provider times out, is rate limited, lacks credentials, or encounters a network error
- THEN the pipeline SHALL emit a structured LLM error and transition through error recovery to wake listening
Requirement: Local TTS synthesis and playback
The system SHALL synthesize assistant replies through the configured TTS provider and SHALL play synthesized speech through the local system output path; cloud mode SHALL use NewAPI-compatible TTS and local mode SHALL use a local playback-capable TTS provider.
Scenario: Reply text is ready for speech
- WHEN the LLM produces a non-empty reply for a live turn
- THEN the TTS provider SHALL synthesize that text into locally playable speech
Scenario: Playback starts
- WHEN the first synthesized audio output is available
- THEN the output Transport or macOS playback command SHALL begin playback through the local speaker
Scenario: TTS fails
- WHEN TTS returns empty audio, fails to create an audio file, or playback command fails
- THEN the live runtime SHALL emit a structured TTS or playback error and SHALL recover without terminating default loop mode
Requirement: Pipeline state machine
The system SHALL expose deterministic pipeline states for idle, wake_listening, speech_detecting, recording, transcribing, thinking, speaking, interrupted, and error_recovering.
Scenario: Normal conversation turn
- WHEN wakeword, VAD, STT, LLM, TTS, and playback all succeed
- THEN the state sequence SHALL progress through wake listening, speech detection, recording, transcribing, thinking, speaking, and back to wake listening
Scenario: Recoverable provider error
- WHEN any provider fails during a conversation turn
- THEN the state machine SHALL enter error recovery and then return to wake listening after cleanup
Scenario: UI subscribes to state
- WHEN the pipeline state changes
- THEN the desktop pet UI SHALL be able to update its visible state without directly invoking audio or model providers
Requirement: Desktop pet visual states
The system SHALL define desktop pet visual states for idle, listening, recording, thinking, speaking, and error feedback.
Scenario: Wakeword is detected
- WHEN the pipeline enters speech detection or recording
- THEN the desktop pet SHALL show a listening or recording visual state
Scenario: LLM is generating
- WHEN the pipeline enters thinking
- THEN the desktop pet SHALL show a thinking visual state
Scenario: TTS is playing
- WHEN the pipeline enters speaking
- THEN the desktop pet SHALL show a speaking visual state
Scenario: Error occurs
- WHEN the pipeline enters error recovery
- THEN the desktop pet SHALL show a concise visible error state and then return to idle or wake listening when recovered
Requirement: Generated pet asset specification
The system SHALL specify that production desktop pet images are generated with the image generation skill when usable, or with a reproducible project-local fallback when image generation output is unavailable or fails validation, and saved inside the project as transparent PNG assets.
Scenario: Image generation output is usable
- WHEN pet images are generated with
imagegenand pass role and transparency validation - THEN the images SHALL be processed for transparency, validated, and saved to a project asset directory
Scenario: Image generation output is unavailable or invalid
- WHEN image generation output is unavailable, inaccessible as a project file, or fails role/transparency validation
- THEN the system SHALL use a reproducible project-local fallback to create transparent PNG pet state assets and SHALL validate them before use
Scenario: Planning phase is executed
- WHEN this OpenSpec change is implemented
- THEN generated or generated-derived pet assets SHALL be saved inside the project and SHALL NOT be referenced from a temporary generation directory
Requirement: Audio feedback suppression
The system SHALL prevent the desktop pet's own TTS playback from being treated as a new user wakeword or speech input.
Scenario: TTS playback is active
- WHEN the output Transport is playing synthesized assistant speech
- THEN wakeword detection and VAD input processing SHALL be suppressed or ignored for that playback window by default
Scenario: User speaks during playback
- WHEN the user speaks while TTS playback is active in the first version
- THEN the system SHALL ignore that speech unless a future interruption feature is explicitly specified
Requirement: Structured errors and logging
The system SHALL represent stage failures with structured error codes and log state transitions, latency metrics, and provider names.
Scenario: Stage transition occurs
- WHEN the pipeline changes state
- THEN the system SHALL log the turn ID, old state, new state, timestamp, and triggering event
Scenario: Provider call completes
- WHEN a provider call succeeds or fails
- THEN the system SHALL log provider name, stage, duration, and error code if present
Scenario: Sensitive data is present
- WHEN logs are written
- THEN the system SHALL NOT log API keys, authorization headers, or raw credentials
Requirement: Performance targets
The system SHALL define measurable first-version latency and reliability targets for wakeword response, endpoint detection, LLM first output, and speech playback.
Scenario: Wakeword latency is measured
- WHEN wakeword detection succeeds in a normal local environment
- THEN the system SHALL target visible listening feedback within 800 ms
Scenario: User stops speaking
- WHEN the user stops speaking after a normal utterance
- THEN VAD endpoint detection SHALL target transition to transcription within 900 ms
Scenario: LLM begins responding
- WHEN a valid transcript has been sent to the LLM provider
- THEN the system SHALL target first playable reply text within 3.5 seconds
Scenario: Normal reply is spoken
- WHEN a normal Chinese question under 10 seconds is processed
- THEN the system SHALL target the start of spoken playback within 5 seconds P95 after the user stops speaking
Requirement: Security and privacy
The system SHALL protect credentials and minimize audio persistence.
Scenario: API key is required
- WHEN the LLM provider needs an OpenAI API key
- THEN the key SHALL be loaded from
.envor local uncommitted configuration and SHALL NOT be committed
Scenario: Audio is processed
- WHEN user speech is captured for STT
- THEN raw audio SHALL NOT be persisted by default
Scenario: Cloud LLM is invoked
- WHEN the system sends data to the cloud LLM
- THEN it SHALL send transcript text and necessary conversation context, not raw microphone audio
Requirement: Testability
The system SHALL be designed so each stage can be tested with mock providers, file-based audio fixtures, and fake live runtime components without requiring real devices in automated tests.
Scenario: Repeated runtime is unit tested
- WHEN fake live providers produce two deterministic turns
- THEN tests SHALL verify two STT calls, two LLM calls, two TTS calls, two playback calls, and final return to standby
Scenario: Temporary context is unit tested
- WHEN fake live providers run two turns in one runtime instance
- THEN tests SHALL verify the second LLM request includes the first turn user and assistant messages
Scenario: Process-local context is unit tested
- WHEN a second runtime instance is created after a first instance has conversation history
- THEN tests SHALL verify the second instance starts with no user/assistant history
Scenario: OpenSpec planning validation runs
- WHEN this change is complete before implementation
- THEN
openspec validate complete-live-repeat-voice-runtime --strictandopenspec validate --all --strictSHALL pass
Requirement: Module-level git commit gates
The system implementation process SHALL require an immediate Git commit after each major module or milestone is completed, where a major module means a top-level task group from tasks.md or a phase milestone from proposal.md.
Scenario: Major module is completed
- WHEN an implementer completes a top-level task group or milestone phase
- THEN the implementer SHALL run the applicable build, tests, lint/type checks, and OpenSpec validation before committing
Scenario: Build validation fails before commit
- WHEN the applicable build, tests, lint/type checks, or OpenSpec validation fail
- THEN the implementer SHALL NOT create the module commit until the failure is fixed and validation passes
Scenario: Module commit is created
- WHEN validation passes for the completed module
- THEN the implementer SHALL immediately create a Git commit using the Chinese message format “[模块名]:完成[具体功能描述],包含[关键变更]”
Scenario: Intermediate work remains after module completion
- WHEN module-related changes remain unstaged or uncommitted after the module is completed
- THEN the implementer SHALL include those changes in the module commit or explicitly separate unrelated external changes before continuing to the next module
Requirement: Live repeat voice runtime
The system SHALL provide a run-live command that performs real repeated voice conversation with local microphone input, configured speech recognition and speech synthesis providers, cloud LLM reply generation, local speaker playback, and automatic return to standby.
Scenario: Live runtime starts in standby
- WHEN the user runs
PYTHONPATH=src python3.11 -m owner_voice_pet run-live - THEN the system SHALL load
.env, validate required live dependencies, initialize local audio/model providers, and enter a standby listening loop
Scenario: Live runtime completes two turns
- WHEN the user wakes the system with “小杰小杰”, asks a question, hears the reply, then wakes it again and asks another question
- THEN the system SHALL complete wake, recording, STT, LLM, TTS, playback for both turns and SHALL return to standby after each turn
Scenario: Once mode completes one turn
- WHEN the user runs
PYTHONPATH=src python3.11 -m owner_voice_pet run-live --once - THEN the system SHALL run at most one wake-to-playback turn and exit after the turn completes or fails with a documented live error
Requirement: Temporary in-process conversation history
The live runtime SHALL maintain conversation history only in memory for the current process and SHALL include retained user/assistant history in later LLM requests during that same process.
Scenario: Second turn uses first turn history
- WHEN the first live turn appends a user transcript and an assistant reply
- AND a second live turn sends an LLM request
- THEN the second request SHALL include the first turn user message and assistant message unless configured context limits require truncation
Scenario: New runtime starts empty
- WHEN a new
run-liveprocess or new live runtime instance starts - THEN it SHALL NOT load user/assistant history from a previous process or previous runtime instance
Scenario: Process exits
- WHEN the live runtime exits for any reason
- THEN conversation history SHALL be discarded and SHALL NOT be written to disk, database, logs, or model files
Requirement: Local speech model management
The system SHALL provide project-local speech model preparation and diagnostics for live VAD/STT operation, storing downloaded model artifacts under models/ without committing them to Git.
Scenario: Models are downloaded
- WHEN the user runs
python3.11 scripts/download_speech_models.py --dir models - THEN the script SHALL create or update a project-local model directory with the files required by the configured VAD/STT providers
Scenario: Model check succeeds
- WHEN required dependencies and model files are available
- THEN
PYTHONPATH=src python3.11 -m owner_voice_pet model-checkSHALL exit successfully and report the model directory and checked providers
Scenario: Model check fails
- WHEN
sherpa-onnxis unavailable, a model file is missing, or a model cannot be loaded - THEN
model-checkSHALL fail with a structured model error and SHALL NOT start live microphone listening
Requirement: Configurable speech provider mode
The system SHALL read OWNER_SPEECH_PROVIDER from .env to choose between cloud speech providers and local speech providers for first-version live runtime.
Scenario: Cloud speech provider is selected
- WHEN
OWNER_SPEECH_PROVIDER=cloud - THEN live ASR SHALL use the configured cloud model from
OWNER_ASR_MODEL, live TTS SHALL use the configured cloud model fromOWNER_TTS_MODEL, and the implementation SHALL default those models tomimo-v2.5-asrandmimo-v2.5-tts
Scenario: Local speech provider is selected
- WHEN
OWNER_SPEECH_PROVIDER=local - THEN live ASR/VAD SHALL use project-local speech model assets and local TTS SHALL use a local playback-capable provider
Scenario: Speech provider is invalid
- WHEN
OWNER_SPEECH_PROVIDERis neithercloudnorlocal - THEN startup validation SHALL fail with a structured configuration error
Requirement: Live audio device readiness
The system SHALL provide live audio device diagnostics and SHALL use sounddevice for first-version real microphone and speaker access.
Scenario: Device check succeeds
- WHEN
sounddevicecan be imported and at least one input and one output device are available - THEN
PYTHONPATH=src python3.11 -m owner_voice_pet device-checkSHALL exit successfully and report usable audio devices
Scenario: Device check fails
- WHEN
sounddeviceis missing, device query fails, microphone permission is denied, or no usable input/output device exists - THEN
device-checkSHALL fail with a structured audio device error
Scenario: Live runtime uses real devices
- WHEN
run-livestarts successfully - THEN it SHALL use the local microphone as its default input and the local speaker or macOS playback command as its default output rather than fixture audio
Requirement: Live terminal state reporting
The live runtime SHALL emit concise Chinese terminal status messages for observable runtime states.
Scenario: Normal turn status
- WHEN a live turn succeeds
- THEN terminal output SHALL include states equivalent to standby, wake hit, recording, transcribing, thinking, speaking, and returning to standby
Scenario: Recoverable error status
- WHEN a live turn encounters empty STT, LLM failure, TTS failure, or playback failure
- THEN terminal output SHALL include the failing stage and a stable error code or recoverable explanation
Requirement: Live runtime error recovery
The live runtime SHALL recover from turn-level failures and continue listening in default loop mode.
Scenario: LLM fails during default loop
- WHEN the LLM provider times out, is rate limited, or returns a network error during a live turn
- THEN the system SHALL report
LIVE_LLM_FAILEDor an equivalent structured error and SHALL return to standby without terminating the process
Scenario: TTS or playback fails during default loop
- WHEN local TTS or audio playback fails during a live turn
- THEN the system SHALL report the failure and SHALL return to standby without terminating the process
Scenario: Startup dependency fails
- WHEN
.env, required models, or audio devices are unavailable at startup - THEN the system SHALL exit with a documented non-zero status instead of entering a fake live loop