## ADDED Requirements ### Requirement: Live repeat voice runtime The system SHALL provide a `run-live` command that performs real repeated voice conversation with local microphone input, configured speech recognition and speech synthesis providers, cloud LLM reply generation, local speaker playback, and automatic return to standby. #### Scenario: Live runtime starts in standby - **WHEN** the user runs `PYTHONPATH=src python3.11 -m owner_voice_pet run-live` - **THEN** the system SHALL load `.env`, validate required live dependencies, initialize local audio/model providers, and enter a standby listening loop #### Scenario: Live runtime completes two turns - **WHEN** the user wakes the system with “小杰小杰”, asks a question, hears the reply, then wakes it again and asks another question - **THEN** the system SHALL complete wake, recording, STT, LLM, TTS, playback for both turns and SHALL return to standby after each turn #### Scenario: Once mode completes one turn - **WHEN** the user runs `PYTHONPATH=src python3.11 -m owner_voice_pet run-live --once` - **THEN** the system SHALL run at most one wake-to-playback turn and exit after the turn completes or fails with a documented live error ### Requirement: Temporary in-process conversation history The live runtime SHALL maintain conversation history only in memory for the current process and SHALL include retained user/assistant history in later LLM requests during that same process. #### Scenario: Second turn uses first turn history - **WHEN** the first live turn appends a user transcript and an assistant reply - **AND** a second live turn sends an LLM request - **THEN** the second request SHALL include the first turn user message and assistant message unless configured context limits require truncation #### Scenario: New runtime starts empty - **WHEN** a new `run-live` process or new live runtime instance starts - **THEN** it SHALL NOT load user/assistant history from a previous process or previous runtime instance #### Scenario: Process exits - **WHEN** the live runtime exits for any reason - **THEN** conversation history SHALL be discarded and SHALL NOT be written to disk, database, logs, or model files ### Requirement: Local speech model management The system SHALL provide project-local speech model preparation and diagnostics for live VAD/STT operation, storing downloaded model artifacts under `models/` without committing them to Git. #### Scenario: Models are downloaded - **WHEN** the user runs `python3.11 scripts/download_speech_models.py --dir models` - **THEN** the script SHALL create or update a project-local model directory with the files required by the configured VAD/STT providers #### Scenario: Model check succeeds - **WHEN** required dependencies and model files are available - **THEN** `PYTHONPATH=src python3.11 -m owner_voice_pet model-check` SHALL exit successfully and report the model directory and checked providers #### Scenario: Model check fails - **WHEN** `sherpa-onnx` is unavailable, a model file is missing, or a model cannot be loaded - **THEN** `model-check` SHALL fail with a structured model error and SHALL NOT start live microphone listening ### Requirement: Configurable speech provider mode The system SHALL read `OWNER_SPEECH_PROVIDER` from `.env` to choose between cloud speech providers and local speech providers for first-version live runtime. #### Scenario: Cloud speech provider is selected - **WHEN** `OWNER_SPEECH_PROVIDER=cloud` - **THEN** live ASR SHALL use the configured cloud model from `OWNER_ASR_MODEL`, live TTS SHALL use the configured cloud model from `OWNER_TTS_MODEL`, and the implementation SHALL default those models to `mimo-v2.5-asr` and `mimo-v2.5-tts` #### Scenario: Local speech provider is selected - **WHEN** `OWNER_SPEECH_PROVIDER=local` - **THEN** live ASR/VAD SHALL use project-local speech model assets and local TTS SHALL use a local playback-capable provider #### Scenario: Speech provider is invalid - **WHEN** `OWNER_SPEECH_PROVIDER` is neither `cloud` nor `local` - **THEN** startup validation SHALL fail with a structured configuration error ### Requirement: Live audio device readiness The system SHALL provide live audio device diagnostics and SHALL use `sounddevice` for first-version real microphone and speaker access. #### Scenario: Device check succeeds - **WHEN** `sounddevice` can be imported and at least one input and one output device are available - **THEN** `PYTHONPATH=src python3.11 -m owner_voice_pet device-check` SHALL exit successfully and report usable audio devices #### Scenario: Device check fails - **WHEN** `sounddevice` is missing, device query fails, microphone permission is denied, or no usable input/output device exists - **THEN** `device-check` SHALL fail with a structured audio device error #### Scenario: Live runtime uses real devices - **WHEN** `run-live` starts successfully - **THEN** it SHALL use the local microphone as its default input and the local speaker or macOS playback command as its default output rather than fixture audio ### Requirement: Live terminal state reporting The live runtime SHALL emit concise Chinese terminal status messages for observable runtime states. #### Scenario: Normal turn status - **WHEN** a live turn succeeds - **THEN** terminal output SHALL include states equivalent to standby, wake hit, recording, transcribing, thinking, speaking, and returning to standby #### Scenario: Recoverable error status - **WHEN** a live turn encounters empty STT, LLM failure, TTS failure, or playback failure - **THEN** terminal output SHALL include the failing stage and a stable error code or recoverable explanation ### Requirement: Live runtime error recovery The live runtime SHALL recover from turn-level failures and continue listening in default loop mode. #### Scenario: LLM fails during default loop - **WHEN** the LLM provider times out, is rate limited, or returns a network error during a live turn - **THEN** the system SHALL report `LIVE_LLM_FAILED` or an equivalent structured error and SHALL return to standby without terminating the process #### Scenario: TTS or playback fails during default loop - **WHEN** local TTS or audio playback fails during a live turn - **THEN** the system SHALL report the failure and SHALL return to standby without terminating the process #### Scenario: Startup dependency fails - **WHEN** `.env`, required models, or audio devices are unavailable at startup - **THEN** the system SHALL exit with a documented non-zero status instead of entering a fake live loop ## MODIFIED Requirements ### Requirement: Conversation context management The system SHALL maintain an in-memory conversation context for the active desktop pet session, and live repeated voice runtime SHALL define that session as the current `run-live` process only. #### Scenario: User transcript is accepted - **WHEN** a valid STT transcript is produced during a live turn - **THEN** the system SHALL append it to the current process context as a user message before invoking the LLM #### Scenario: Assistant reply completes - **WHEN** the LLM and TTS stages complete a non-empty reply during a live turn - **THEN** the system SHALL append the assistant text to the current process context #### Scenario: Context exceeds configured budget - **WHEN** the context exceeds the configured message or character budget - **THEN** the system SHALL preserve the system prompt and most recent conversation turns while removing older ordinary messages #### Scenario: Runtime exits - **WHEN** the `run-live` process exits - **THEN** the system SHALL discard the context and SHALL NOT persist it across process restarts ### Requirement: Local microphone and speaker transport The system SHALL use the local microphone as the first-version input Transport and the local system speaker as the first-version output Transport, and `run-live` SHALL exercise those real local devices by default. #### Scenario: Microphone input is available - **WHEN** the configured microphone is available and permitted - **THEN** the live Transport SHALL provide streaming audio frames to wake detection, VAD, and STT stages #### Scenario: Speaker output is available - **WHEN** TTS returns a valid local audio segment or audio file and the configured speaker is available - **THEN** the live Transport or macOS playback command SHALL play the audio through the local speaker #### Scenario: Audio device is unavailable - **WHEN** the microphone or speaker is missing, denied, unsupported, or inaccessible through `sounddevice` - **THEN** the system SHALL expose a recoverable Transport error with a stable error code and SHALL NOT silently fall back to fixture audio in live mode ### Requirement: Local STT transcription The system SHALL transcribe captured user utterances through the configured STT provider; cloud mode SHALL use NewAPI-compatible ASR and local mode SHALL use project-local `sherpa-onnx` model assets under `models/`. #### Scenario: STT succeeds - **WHEN** configured STT returns non-empty text for a captured live audio segment - **THEN** the pipeline SHALL add the trimmed text as a user message to the current process conversation context #### Scenario: STT returns empty text - **WHEN** STT returns empty text, punctuation-only text, or an invalid transcript - **THEN** the live runtime SHALL skip LLM invocation and return to standby with a recoverable status #### Scenario: STT provider fails - **WHEN** the STT provider raises an error, cloud ASR fails, or local model loading fails - **THEN** the system SHALL emit an STT or model error code and SHALL recover to a state where future wake attempts are possible if startup can continue safely ### Requirement: Local TTS synthesis and playback The system SHALL synthesize assistant replies through the configured TTS provider and SHALL play synthesized speech through the local system output path; cloud mode SHALL use NewAPI-compatible TTS and local mode SHALL use a local playback-capable TTS provider. #### Scenario: Reply text is ready for speech - **WHEN** the LLM produces a non-empty reply for a live turn - **THEN** the TTS provider SHALL synthesize that text into locally playable speech #### Scenario: Playback starts - **WHEN** the first synthesized audio output is available - **THEN** the output Transport or macOS playback command SHALL begin playback through the local speaker #### Scenario: TTS fails - **WHEN** TTS returns empty audio, fails to create an audio file, or playback command fails - **THEN** the live runtime SHALL emit a structured TTS or playback error and SHALL recover without terminating default loop mode ### Requirement: Testability The system SHALL be designed so each stage can be tested with mock providers, file-based audio fixtures, and fake live runtime components without requiring real devices in automated tests. #### Scenario: Repeated runtime is unit tested - **WHEN** fake live providers produce two deterministic turns - **THEN** tests SHALL verify two STT calls, two LLM calls, two TTS calls, two playback calls, and final return to standby #### Scenario: Temporary context is unit tested - **WHEN** fake live providers run two turns in one runtime instance - **THEN** tests SHALL verify the second LLM request includes the first turn user and assistant messages #### Scenario: Process-local context is unit tested - **WHEN** a second runtime instance is created after a first instance has conversation history - **THEN** tests SHALL verify the second instance starts with no user/assistant history #### Scenario: OpenSpec planning validation runs - **WHEN** this change is complete before implementation - **THEN** `openspec validate complete-live-repeat-voice-runtime --strict` and `openspec validate --all --strict` SHALL pass