97 lines
5.6 KiB
Markdown
97 lines
5.6 KiB
Markdown
## ADDED Requirements
|
|
|
|
### Requirement: Complete full-duplex agent runtime
|
|
The system SHALL provide a complete `run-agent-live` runtime that uses a dedicated full-duplex agent architecture instead of the legacy turn-based `VoiceAssistantPipeline.run_agent_turn()` loop.
|
|
|
|
#### Scenario: Full-duplex entry starts the new runtime
|
|
- **WHEN** the user runs `.venv/bin/python -m owner_voice_pet run-agent-live`
|
|
- **THEN** the command SHALL construct `FullDuplexAgentRuntime`
|
|
- **AND** it SHALL NOT call `VoiceAssistantPipeline.run_agent_turn()` as the primary runtime path
|
|
|
|
#### Scenario: Legacy runtime remains available
|
|
- **WHEN** the user runs `.venv/bin/python -m owner_voice_pet run-live`
|
|
- **THEN** the legacy wake-word turn-based runtime SHALL remain available
|
|
|
|
### Requirement: AudioHub fanout
|
|
The full-duplex runtime SHALL use a single AudioHub to own microphone input and distribute frames to independent consumers without frame stealing.
|
|
|
|
#### Scenario: Multiple consumers subscribe
|
|
- **WHEN** VAD, streaming STT, interrupt detection, and diagnostics subscribe to processed capture audio
|
|
- **THEN** each consumer SHALL receive the same ordered audio frame sequence from its own cursor
|
|
|
|
#### Scenario: Ring buffer overflows
|
|
- **WHEN** capture or render ring buffer exceeds configured capacity
|
|
- **THEN** the hub SHALL drop oldest safe frames, emit an overrun diagnostic, and continue without blocking audio callbacks
|
|
|
|
### Requirement: Required WebRTC APM in full-duplex mode
|
|
The full-duplex runtime SHALL process microphone capture through WebRTC AEC, NS, and AGC using playback render reference before VAD, STT, or interrupt detection consume frames.
|
|
|
|
#### Scenario: APM is unavailable
|
|
- **WHEN** `OWNER_AUDIO_APM_PROVIDER=webrtc`, `OWNER_AUDIO_APM_REQUIRED=1`, and no real WebRTC APM provider is available
|
|
- **THEN** `run-agent-live` SHALL fail startup with a structured audio APM error
|
|
- **AND** it SHALL NOT silently run the software barge-in fallback as complete full-duplex mode
|
|
|
|
#### Scenario: Fake APM is used in self-test
|
|
- **WHEN** `agent-self-test` or unit tests request fake APM
|
|
- **THEN** fake APM SHALL be allowed and SHALL be clearly reported as test-only
|
|
|
|
#### Scenario: Pure assistant echo occurs
|
|
- **WHEN** microphone input is only assistant render echo
|
|
- **THEN** APM-processed audio SHALL NOT trigger valid interruption or valid user transcript
|
|
|
|
### Requirement: Always-on interruption controller
|
|
The full-duplex runtime SHALL run a continuous interruption controller that can cancel active response work from `thinking`, `speaking`, or cancellable `tool_running` states.
|
|
|
|
#### Scenario: User interrupts while speaking
|
|
- **WHEN** the assistant is playing audio and processed capture contains valid user speech
|
|
- **THEN** playback SHALL stop, active LLM/TTS work SHALL be cancelled, and the state SHALL transition through `interrupted` back to `listening`
|
|
|
|
#### Scenario: Interruption latency is measured
|
|
- **WHEN** self-test injects user speech during playback
|
|
- **THEN** the runtime SHALL report VAD-to-playback-stop latency and target P95 under 200 ms
|
|
|
|
#### Scenario: Interruption does not require STT partial
|
|
- **WHEN** valid user speech is detected before streaming STT emits partial text
|
|
- **THEN** playback cancellation SHALL still occur
|
|
|
|
### Requirement: Streaming STT and streaming TTS runtime
|
|
The full-duplex runtime SHALL use streaming recognition and streaming speech output contracts for the agent path.
|
|
|
|
#### Scenario: User speaks during listening
|
|
- **WHEN** processed capture contains user speech
|
|
- **THEN** streaming STT SHALL emit partial or stable transcript events while speech continues and final transcript when the utterance ends
|
|
|
|
#### Scenario: LLM emits a sentence
|
|
- **WHEN** the LLM stream produces a complete safe speech segment
|
|
- **THEN** streaming TTS SHALL synthesize playable PCM chunks without waiting for the full assistant response
|
|
|
|
#### Scenario: Playback is interrupted
|
|
- **WHEN** cancellation occurs during streaming TTS or playback
|
|
- **THEN** unplayed audio chunks SHALL be discarded and unspoken assistant text SHALL NOT be committed to context
|
|
|
|
### Requirement: Full-duplex Conversation Manager
|
|
The full-duplex runtime SHALL coordinate final user transcripts, session context, long-term memory retrieval, LLM streaming, tool calls, TTS playback, interruption, and recovery.
|
|
|
|
#### Scenario: Memory is enabled
|
|
- **WHEN** a final user transcript is accepted and memory is enabled
|
|
- **THEN** the Conversation Manager SHALL retrieve Top-K relevant memories and provide them as separate context to the LLM
|
|
|
|
#### Scenario: Tool call is requested
|
|
- **WHEN** the LLM stream requests a supported tool call
|
|
- **THEN** the Tool Router SHALL classify, approve, reject, or request confirmation before execution
|
|
|
|
#### Scenario: Tool is high risk
|
|
- **WHEN** a tool request involves deletion, upload, payment, account changes, installation, or other high-risk action
|
|
- **THEN** the runtime SHALL NOT execute it automatically
|
|
|
|
### Requirement: Full-duplex self-test commands
|
|
The system SHALL provide deterministic self-test commands for the complete full-duplex agent runtime.
|
|
|
|
#### Scenario: Agent self-test runs
|
|
- **WHEN** the developer runs `.venv/bin/python -m owner_voice_pet agent-self-test --profile full-duplex --turns 3`
|
|
- **THEN** the command SHALL exercise STT, LLM, TTS, interruption, memory, and Tool Router paths and return JSON with `success=true` only when all checks pass
|
|
|
|
#### Scenario: Audio self-test runs
|
|
- **WHEN** the developer runs `.venv/bin/python -m owner_voice_pet audio-self-test --duration 10 --check-echo`
|
|
- **THEN** the command SHALL report device availability, APM provider health, echo suppression result, and interruption latency statistics
|