Files
Owner/openspec/changes/archive/2026-06-17-add-voice-pet-pipeline/specs/voice-pet-pipeline/spec.md
T

14 KiB

ADDED Requirements

Requirement: Python desktop pet runtime

The system SHALL be specified as a local Python desktop pet application that owns the voice pipeline, desktop pet state, local audio input, and local audio output in the first version.

Scenario: App starts in local desktop mode

  • WHEN the user starts the future application on macOS
  • THEN the system SHALL initialize as a local desktop pet process rather than a remote service or browser-only tool

Scenario: Implementation follows OpenSpec artifacts

  • WHEN this OpenSpec change is implemented
  • THEN the repository SHALL contain Python source code, automated tests, validation commands, and project-local pet assets aligned with the planning artifacts

Requirement: Local microphone and speaker transport

The system SHALL use the local microphone as the first-version input Transport and the local system speaker as the first-version output Transport.

Scenario: Microphone input is available

  • WHEN the configured microphone is available and permitted
  • THEN the Transport SHALL provide streaming audio frames to the wakeword and VAD stages

Scenario: Speaker output is available

  • WHEN TTS returns a valid audio segment and the configured speaker is available
  • THEN the Transport SHALL play the audio segment through the local speaker

Scenario: Audio device is unavailable

  • WHEN the microphone or speaker is missing, denied, or unsupported
  • THEN the system SHALL expose a recoverable Transport error with a stable error code

Requirement: Wake word detection

The system SHALL listen locally for the Chinese wake word “小杰小杰” before accepting user speech for a conversation turn.

Scenario: Wake word is detected

  • WHEN the user says “小杰小杰” and the wakeword provider returns confidence above the configured threshold
  • THEN the pipeline SHALL transition from wake listening to speech detection

Scenario: Wake word is not detected

  • WHEN background speech or noise does not match “小杰小杰”
  • THEN the pipeline SHALL remain in wake listening and SHALL NOT invoke STT, LLM, or TTS

Scenario: Wake model fails

  • WHEN the wakeword provider cannot load or process audio
  • THEN the system SHALL report a wakeword error and SHALL NOT crash the desktop pet process

Requirement: VAD speech endpoint detection

After wakeword detection, the system SHALL use VAD to identify when the user starts and stops speaking.

Scenario: User begins speaking

  • WHEN VAD detects continuous speech above the configured start threshold
  • THEN the pipeline SHALL begin recording the current utterance

Scenario: User stops speaking

  • WHEN VAD detects continuous silence above the configured end threshold
  • THEN the pipeline SHALL close the current audio segment and send it to STT

Scenario: User says nothing after wakeword

  • WHEN no speech is detected before the configured no-speech timeout
  • THEN the pipeline SHALL return to wake listening without invoking STT or LLM

Scenario: Recording exceeds maximum duration

  • WHEN speech continues beyond the configured maximum utterance duration
  • THEN the pipeline SHALL end the segment, mark the end reason, and continue to STT with the captured audio

Requirement: Local STT transcription

The system SHALL transcribe the captured user utterance through a local STT provider, with sherpa-onnx as the recommended first implementation candidate.

Scenario: STT succeeds

  • WHEN STT returns non-empty text for a captured audio segment
  • THEN the pipeline SHALL add the trimmed text as a user message to the conversation context

Scenario: STT returns empty text

  • WHEN STT returns empty text, punctuation-only text, or an invalid transcript
  • THEN the pipeline SHALL skip LLM invocation and return to wake listening with a recoverable status

Scenario: STT provider fails

  • WHEN the STT provider raises an error or cannot load its local model
  • THEN the system SHALL emit an STT error code and SHALL recover to a state where future wake attempts are possible

Requirement: Conversation context management

The system SHALL maintain an in-memory conversation context for the active desktop pet session.

Scenario: User transcript is accepted

  • WHEN a valid STT transcript is produced
  • THEN the system SHALL append it to the context as a user message before invoking the LLM

Scenario: Assistant reply completes

  • WHEN the LLM and TTS stages complete a reply
  • THEN the system SHALL append the assistant text to the context

Scenario: Context exceeds configured budget

  • WHEN the context exceeds the configured message, character, or token budget
  • THEN the system SHALL preserve the system prompt and most recent conversation turns while removing older ordinary messages

Requirement: Cloud LLM streaming reply

The system SHALL use a cloud LLM provider for reply generation, with OpenAI Responses API streaming as the recommended first implementation.

Scenario: LLM stream begins

  • WHEN the LLM provider receives valid context messages and configuration
  • THEN it SHALL stream reply deltas rather than waiting for the full reply before producing output

Scenario: LLM model is configured

  • WHEN the system starts
  • THEN the LLM model name SHALL be read from configuration and SHALL NOT be hard-coded in pipeline logic

Scenario: LLM request fails

  • WHEN the LLM provider times out, is rate limited, lacks credentials, or encounters a network error
  • THEN the pipeline SHALL emit a structured LLM error and transition through error recovery to wake listening

Requirement: Local TTS synthesis and playback

The system SHALL synthesize assistant replies through a local TTS provider, with sherpa-onnx as the recommended first implementation candidate.

Scenario: Sentence is ready for speech

  • WHEN the LLM stream produces a complete sentence or configured text chunk
  • THEN the TTS provider SHALL synthesize that text into an audio segment

Scenario: First TTS segment is ready

  • WHEN the first synthesized audio segment is available
  • THEN the output Transport SHALL begin playback without waiting for every remaining LLM delta

Scenario: TTS fails

  • WHEN TTS returns empty audio or raises an error
  • THEN the pipeline SHALL emit a structured TTS error and SHALL recover without terminating the application

Requirement: Pipeline state machine

The system SHALL expose deterministic pipeline states for idle, wake_listening, speech_detecting, recording, transcribing, thinking, speaking, interrupted, and error_recovering.

Scenario: Normal conversation turn

  • WHEN wakeword, VAD, STT, LLM, TTS, and playback all succeed
  • THEN the state sequence SHALL progress through wake listening, speech detection, recording, transcribing, thinking, speaking, and back to wake listening

Scenario: Recoverable provider error

  • WHEN any provider fails during a conversation turn
  • THEN the state machine SHALL enter error recovery and then return to wake listening after cleanup

Scenario: UI subscribes to state

  • WHEN the pipeline state changes
  • THEN the desktop pet UI SHALL be able to update its visible state without directly invoking audio or model providers

Requirement: Desktop pet visual states

The system SHALL define desktop pet visual states for idle, listening, recording, thinking, speaking, and error feedback.

Scenario: Wakeword is detected

  • WHEN the pipeline enters speech detection or recording
  • THEN the desktop pet SHALL show a listening or recording visual state

Scenario: LLM is generating

  • WHEN the pipeline enters thinking
  • THEN the desktop pet SHALL show a thinking visual state

Scenario: TTS is playing

  • WHEN the pipeline enters speaking
  • THEN the desktop pet SHALL show a speaking visual state

Scenario: Error occurs

  • WHEN the pipeline enters error recovery
  • THEN the desktop pet SHALL show a concise visible error state and then return to idle or wake listening when recovered

Requirement: Generated pet asset specification

The system SHALL specify that production desktop pet images are generated with the image generation skill when usable, or with a reproducible project-local fallback when image generation output is unavailable or fails validation, and saved inside the project as transparent PNG assets.

Scenario: Image generation output is usable

  • WHEN pet images are generated with imagegen and pass role and transparency validation
  • THEN the images SHALL be processed for transparency, validated, and saved to a project asset directory

Scenario: Image generation output is unavailable or invalid

  • WHEN image generation output is unavailable, inaccessible as a project file, or fails role/transparency validation
  • THEN the system SHALL use a reproducible project-local fallback to create transparent PNG pet state assets and SHALL validate them before use

Scenario: Planning phase is executed

  • WHEN this OpenSpec change is implemented
  • THEN generated or generated-derived pet assets SHALL be saved inside the project and SHALL NOT be referenced from a temporary generation directory

Requirement: Audio feedback suppression

The system SHALL prevent the desktop pet's own TTS playback from being treated as a new user wakeword or speech input.

Scenario: TTS playback is active

  • WHEN the output Transport is playing synthesized assistant speech
  • THEN wakeword detection and VAD input processing SHALL be suppressed or ignored for that playback window by default

Scenario: User speaks during playback

  • WHEN the user speaks while TTS playback is active in the first version
  • THEN the system SHALL ignore that speech unless a future interruption feature is explicitly specified

Requirement: Structured errors and logging

The system SHALL represent stage failures with structured error codes and log state transitions, latency metrics, and provider names.

Scenario: Stage transition occurs

  • WHEN the pipeline changes state
  • THEN the system SHALL log the turn ID, old state, new state, timestamp, and triggering event

Scenario: Provider call completes

  • WHEN a provider call succeeds or fails
  • THEN the system SHALL log provider name, stage, duration, and error code if present

Scenario: Sensitive data is present

  • WHEN logs are written
  • THEN the system SHALL NOT log API keys, authorization headers, or raw credentials

Requirement: Performance targets

The system SHALL define measurable first-version latency and reliability targets for wakeword response, endpoint detection, LLM first output, and speech playback.

Scenario: Wakeword latency is measured

  • WHEN wakeword detection succeeds in a normal local environment
  • THEN the system SHALL target visible listening feedback within 800 ms

Scenario: User stops speaking

  • WHEN the user stops speaking after a normal utterance
  • THEN VAD endpoint detection SHALL target transition to transcription within 900 ms

Scenario: LLM begins responding

  • WHEN a valid transcript has been sent to the LLM provider
  • THEN the system SHALL target first playable reply text within 3.5 seconds

Scenario: Normal reply is spoken

  • WHEN a normal Chinese question under 10 seconds is processed
  • THEN the system SHALL target the start of spoken playback within 5 seconds P95 after the user stops speaking

Requirement: Security and privacy

The system SHALL protect credentials and minimize audio persistence.

Scenario: API key is required

  • WHEN the LLM provider needs an OpenAI API key
  • THEN the key SHALL be loaded from environment or local uncommitted configuration and SHALL NOT be committed

Scenario: Audio is processed

  • WHEN user speech is captured for STT
  • THEN raw audio SHALL NOT be persisted by default

Scenario: Cloud LLM is invoked

  • WHEN the system sends data to the cloud LLM
  • THEN it SHALL send transcript text and necessary conversation context, not raw microphone audio

Requirement: Testability

The system SHALL be designed so each stage can be tested with mock providers and file-based audio fixtures.

Scenario: Pipeline is unit tested

  • WHEN mock providers produce deterministic events
  • THEN tests SHALL verify state transitions, context updates, and error recovery

Scenario: Audio fixture is replayed

  • WHEN a file-based audio fixture containing “小杰小杰” and user speech is replayed in a test Transport
  • THEN the pipeline SHALL be testable without using a live microphone

Scenario: OpenSpec planning validation runs

  • WHEN this planning change is complete
  • THEN openspec validate add-voice-pet-pipeline --strict and openspec validate --all --strict SHALL pass

Requirement: Module-level git commit gates

The system implementation process SHALL require an immediate Git commit after each major module or milestone is completed, where a major module means a top-level task group from tasks.md or a phase milestone from proposal.md.

Scenario: Major module is completed

  • WHEN an implementer completes a top-level task group or milestone phase
  • THEN the implementer SHALL run the applicable build, tests, lint/type checks, and OpenSpec validation before committing

Scenario: Build validation fails before commit

  • WHEN the applicable build, tests, lint/type checks, or OpenSpec validation fail
  • THEN the implementer SHALL NOT create the module commit until the failure is fixed and validation passes

Scenario: Module commit is created

  • WHEN validation passes for the completed module
  • THEN the implementer SHALL immediately create a Git commit using the Chinese message format “[模块名]:完成[具体功能描述],包含[关键变更]”

Scenario: Intermediate work remains after module completion

  • WHEN module-related changes remain unstaged or uncommitted after the module is completed
  • THEN the implementer SHALL include those changes in the module commit or explicitly separate unrelated external changes before continuing to the next module