Files
Owner/openspec/changes/complete-live-repeat-voice-runtime/specs/voice-pet-pipeline/spec.md
T

10 KiB

ADDED Requirements

Requirement: Live repeat voice runtime

The system SHALL provide a run-live command that performs real repeated voice conversation with local microphone input, local speech processing, cloud LLM reply generation, local TTS, local speaker playback, and automatic return to standby.

Scenario: Live runtime starts in standby

  • WHEN the user runs PYTHONPATH=src python3.11 -m owner_voice_pet run-live
  • THEN the system SHALL load .env, validate required live dependencies, initialize local audio/model providers, and enter a standby listening loop

Scenario: Live runtime completes two turns

  • WHEN the user wakes the system with “小杰小杰”, asks a question, hears the reply, then wakes it again and asks another question
  • THEN the system SHALL complete wake, recording, STT, LLM, TTS, playback for both turns and SHALL return to standby after each turn

Scenario: Once mode completes one turn

  • WHEN the user runs PYTHONPATH=src python3.11 -m owner_voice_pet run-live --once
  • THEN the system SHALL run at most one wake-to-playback turn and exit after the turn completes or fails with a documented live error

Requirement: Temporary in-process conversation history

The live runtime SHALL maintain conversation history only in memory for the current process and SHALL include retained user/assistant history in later LLM requests during that same process.

Scenario: Second turn uses first turn history

  • WHEN the first live turn appends a user transcript and an assistant reply
  • AND a second live turn sends an LLM request
  • THEN the second request SHALL include the first turn user message and assistant message unless configured context limits require truncation

Scenario: New runtime starts empty

  • WHEN a new run-live process or new live runtime instance starts
  • THEN it SHALL NOT load user/assistant history from a previous process or previous runtime instance

Scenario: Process exits

  • WHEN the live runtime exits for any reason
  • THEN conversation history SHALL be discarded and SHALL NOT be written to disk, database, logs, or model files

Requirement: Local speech model management

The system SHALL provide project-local speech model preparation and diagnostics for live VAD/STT operation, storing downloaded model artifacts under models/ without committing them to Git.

Scenario: Models are downloaded

  • WHEN the user runs python3.11 scripts/download_speech_models.py --dir models
  • THEN the script SHALL create or update a project-local model directory with the files required by the configured VAD/STT providers

Scenario: Model check succeeds

  • WHEN required dependencies and model files are available
  • THEN PYTHONPATH=src python3.11 -m owner_voice_pet model-check SHALL exit successfully and report the model directory and checked providers

Scenario: Model check fails

  • WHEN sherpa-onnx is unavailable, a model file is missing, or a model cannot be loaded
  • THEN model-check SHALL fail with a structured model error and SHALL NOT start live microphone listening

Requirement: Live audio device readiness

The system SHALL provide live audio device diagnostics and SHALL use sounddevice for first-version real microphone and speaker access.

Scenario: Device check succeeds

  • WHEN sounddevice can be imported and at least one input and one output device are available
  • THEN PYTHONPATH=src python3.11 -m owner_voice_pet device-check SHALL exit successfully and report usable audio devices

Scenario: Device check fails

  • WHEN sounddevice is missing, device query fails, microphone permission is denied, or no usable input/output device exists
  • THEN device-check SHALL fail with a structured audio device error

Scenario: Live runtime uses real devices

  • WHEN run-live starts successfully
  • THEN it SHALL use the local microphone as its default input and the local speaker or macOS playback command as its default output rather than fixture audio

Requirement: Live terminal state reporting

The live runtime SHALL emit concise Chinese terminal status messages for observable runtime states.

Scenario: Normal turn status

  • WHEN a live turn succeeds
  • THEN terminal output SHALL include states equivalent to standby, wake hit, recording, transcribing, thinking, speaking, and returning to standby

Scenario: Recoverable error status

  • WHEN a live turn encounters empty STT, LLM failure, TTS failure, or playback failure
  • THEN terminal output SHALL include the failing stage and a stable error code or recoverable explanation

Requirement: Live runtime error recovery

The live runtime SHALL recover from turn-level failures and continue listening in default loop mode.

Scenario: LLM fails during default loop

  • WHEN the LLM provider times out, is rate limited, or returns a network error during a live turn
  • THEN the system SHALL report LIVE_LLM_FAILED or an equivalent structured error and SHALL return to standby without terminating the process

Scenario: TTS or playback fails during default loop

  • WHEN local TTS or audio playback fails during a live turn
  • THEN the system SHALL report the failure and SHALL return to standby without terminating the process

Scenario: Startup dependency fails

  • WHEN .env, required models, or audio devices are unavailable at startup
  • THEN the system SHALL exit with a documented non-zero status instead of entering a fake live loop

MODIFIED Requirements

Requirement: Conversation context management

The system SHALL maintain an in-memory conversation context for the active desktop pet session, and live repeated voice runtime SHALL define that session as the current run-live process only.

Scenario: User transcript is accepted

  • WHEN a valid STT transcript is produced during a live turn
  • THEN the system SHALL append it to the current process context as a user message before invoking the LLM

Scenario: Assistant reply completes

  • WHEN the LLM and TTS stages complete a non-empty reply during a live turn
  • THEN the system SHALL append the assistant text to the current process context

Scenario: Context exceeds configured budget

  • WHEN the context exceeds the configured message or character budget
  • THEN the system SHALL preserve the system prompt and most recent conversation turns while removing older ordinary messages

Scenario: Runtime exits

  • WHEN the run-live process exits
  • THEN the system SHALL discard the context and SHALL NOT persist it across process restarts

Requirement: Local microphone and speaker transport

The system SHALL use the local microphone as the first-version input Transport and the local system speaker as the first-version output Transport, and run-live SHALL exercise those real local devices by default.

Scenario: Microphone input is available

  • WHEN the configured microphone is available and permitted
  • THEN the live Transport SHALL provide streaming audio frames to wake detection, VAD, and STT stages

Scenario: Speaker output is available

  • WHEN TTS returns a valid local audio segment or audio file and the configured speaker is available
  • THEN the live Transport or macOS playback command SHALL play the audio through the local speaker

Scenario: Audio device is unavailable

  • WHEN the microphone or speaker is missing, denied, unsupported, or inaccessible through sounddevice
  • THEN the system SHALL expose a recoverable Transport error with a stable error code and SHALL NOT silently fall back to fixture audio in live mode

Requirement: Local STT transcription

The system SHALL transcribe captured user utterances through a local STT provider, with sherpa-onnx as the first live implementation candidate and project-local models under models/.

Scenario: STT succeeds

  • WHEN local STT returns non-empty text for a captured live audio segment
  • THEN the pipeline SHALL add the trimmed text as a user message to the current process conversation context

Scenario: STT returns empty text

  • WHEN STT returns empty text, punctuation-only text, or an invalid transcript
  • THEN the live runtime SHALL skip LLM invocation and return to standby with a recoverable status

Scenario: STT provider fails

  • WHEN the STT provider raises an error or cannot load its local model
  • THEN the system SHALL emit an STT or model error code and SHALL recover to a state where future wake attempts are possible if startup can continue safely

Requirement: Local TTS synthesis and playback

The system SHALL synthesize assistant replies through a local TTS provider and SHALL play synthesized speech through the local system output path, with macOS say/afplay accepted as the first live implementation.

Scenario: Reply text is ready for speech

  • WHEN the LLM produces a non-empty reply for a live turn
  • THEN the TTS provider SHALL synthesize that text into locally playable speech

Scenario: Playback starts

  • WHEN the first synthesized audio output is available
  • THEN the output Transport or macOS playback command SHALL begin playback through the local speaker

Scenario: TTS fails

  • WHEN TTS returns empty audio, fails to create an audio file, or playback command fails
  • THEN the live runtime SHALL emit a structured TTS or playback error and SHALL recover without terminating default loop mode

Requirement: Testability

The system SHALL be designed so each stage can be tested with mock providers, file-based audio fixtures, and fake live runtime components without requiring real devices in automated tests.

Scenario: Repeated runtime is unit tested

  • WHEN fake live providers produce two deterministic turns
  • THEN tests SHALL verify two STT calls, two LLM calls, two TTS calls, two playback calls, and final return to standby

Scenario: Temporary context is unit tested

  • WHEN fake live providers run two turns in one runtime instance
  • THEN tests SHALL verify the second LLM request includes the first turn user and assistant messages

Scenario: Process-local context is unit tested

  • WHEN a second runtime instance is created after a first instance has conversation history
  • THEN tests SHALL verify the second instance starts with no user/assistant history

Scenario: OpenSpec planning validation runs

  • WHEN this change is complete before implementation
  • THEN openspec validate complete-live-repeat-voice-runtime --strict and openspec validate --all --strict SHALL pass