25 KiB
ADDED Requirements
Requirement: Full-duplex agent mode
The system SHALL support a planned full-duplex agent voice assistant mode that runs continuous microphone listening, streaming recognition, streaming response generation, streaming speech playback, interruption handling, memory retrieval, and safe tool routing as one coordinated pipeline.
Scenario: Full-duplex mode starts
- WHEN the future implementation starts with
OWNER_ASSISTANT_MODE=full_duplex_agent - THEN the system SHALL initialize continuous audio capture, audio playback, WebRTC audio processing, VAD, streaming STT, streaming TTS, conversation management, memory, and tool routing before accepting user interaction
Scenario: Turn-based mode remains available during migration
- WHEN the future implementation starts with the existing turn-based mode
- THEN the system SHALL preserve the existing wake-to-capture-to-STT-to-LLM-to-TTS behavior unless full-duplex mode is explicitly enabled
Scenario: Full-duplex provider is unavailable
- WHEN a required full-duplex provider cannot be initialized and no configured fallback is allowed
- THEN the system SHALL fail startup with a structured provider error rather than entering a fake full-duplex loop
Requirement: WebRTC audio processing foundation
The system SHALL define a WebRTC audio processing stage for full-duplex mode that applies acoustic echo cancellation, noise suppression, and automatic gain control to microphone capture audio using speaker playback audio as render reference.
Scenario: AEC uses playback reference
- WHEN TTS playback writes PCM frames to the speaker output path
- THEN the same render audio SHALL be provided to the WebRTC audio processing stage as echo cancellation reference
Scenario: Capture audio is processed
- WHEN microphone frames are captured during full-duplex mode
- THEN the frames SHALL pass through AEC, noise suppression, and automatic gain control before VAD, interruption detection, and streaming STT consume them
Scenario: Pure assistant echo is present
- WHEN the microphone receives only the assistant's own playback echo and no user speech
- THEN the processed audio SHALL NOT produce a valid user interruption or valid user transcript
Scenario: APM format mismatch occurs
- WHEN capture audio and render reference audio have incompatible sample rate, channel count, or frame size
- THEN the system SHALL either explicitly resample to the configured internal format or fail with
AUDIO_APM_FORMAT_MISMATCH
Scenario: WebRTC APM fails at runtime
- WHEN the WebRTC audio processing stage raises a processing error
- THEN the pipeline SHALL emit a structured audio processing error and transition through recovery without corrupting conversation history
Requirement: Full-duplex audio ring buffers
The system SHALL use bounded capture and render reference ring buffers with timestamps so full-duplex capture, playback, AEC, VAD, STT, and interruption detection can run concurrently without unbounded memory growth.
Scenario: Capture frame is received
- WHEN the microphone callback receives a PCM frame
- THEN the frame SHALL be written to the capture ring buffer with a monotonic timestamp, frame id, sample rate, channel count, and duration
Scenario: Render frame is queued
- WHEN TTS or playback queues a PCM frame for speaker output
- THEN the frame SHALL be written to the render reference ring buffer with timing metadata usable by AEC
Scenario: Ring buffer overflows
- WHEN an audio ring buffer exceeds its configured capacity
- THEN the system SHALL drop the oldest safe frames, emit an
audio_buffer_overrunevent, and continue without blocking the audio callback indefinitely
Scenario: Reference drift is detected
- WHEN capture and render timestamps drift beyond the configured tolerance
- THEN the system SHALL emit an
audio_reference_driftevent and reduce confidence in interruption decisions until alignment recovers
Requirement: Full-duplex agent state machine
The system SHALL define a deterministic full-duplex state machine with idle, listening, thinking, speaking, interrupted, tool_running, and recovering states.
Scenario: User input reaches the LLM
- WHEN full-duplex listening produces a valid final user transcript
- THEN the state machine SHALL transition from
listeningtothinking
Scenario: First playable response is ready
- WHEN the LLM response stream and TTS pipeline produce the first playable audio chunk
- THEN the state machine SHALL transition from
thinkingtospeaking
Scenario: User interrupts playback
- WHEN valid user speech is detected while the assistant is speaking
- THEN the state machine SHALL transition from
speakingtointerrupted
Scenario: Tool call starts
- WHEN the LLM requests an approved tool call
- THEN the state machine SHALL transition to
tool_runninguntil the tool result is returned, rejected, cancelled, or fails
Scenario: Recoverable error occurs
- WHEN a recoverable provider, memory, tool, STT, TTS, LLM, or playback error occurs
- THEN the state machine SHALL transition to
recoveringand then to a safelisteningoridlestate after cleanup
Requirement: Pipeline event bus for agent stages
The system SHALL expose a stable pipeline event bus for full-duplex audio, recognition, generation, playback, interruption, memory, tools, confirmation, and recovery events.
Scenario: Terminal reporter displays state
- WHEN the terminal displays full-duplex runtime status
- THEN it SHALL derive its output from pipeline events rather than directly reading provider internals
Scenario: GUI subscribes to the same events
- WHEN a future desktop pet GUI is implemented
- THEN it SHALL be able to subscribe to the same event bus used by the terminal reporter and tests
Scenario: Event contains diagnostics
- WHEN a stage emits an event
- THEN the event SHALL include the session id or turn id when available, stage name, event type, timestamp, and sanitized payload
Scenario: Sensitive payload is present
- WHEN an event involves credentials, raw audio, memory content, or tool output
- THEN the event payload SHALL omit secrets and SHALL truncate or summarize sensitive content according to security policy
Requirement: Low-latency interruption and cancellation
The full-duplex mode SHALL detect valid user speech during assistant playback and cancel the active response pipeline with a target interruption latency under 200 ms.
Scenario: User speaks while assistant is speaking
- WHEN the assistant is playing TTS audio and APM-processed microphone audio contains valid user speech
- THEN the interruption detector SHALL emit
interrupt_detected, stop remaining playback, and cancel active LLM and TTS work
Scenario: Interruption latency is measured
- WHEN a test fixture injects user speech during
speaking - THEN the system SHALL record the latency from VAD speech start to playback stop and SHALL target P95 under 200 ms
Scenario: Echo does not interrupt
- WHEN microphone input during playback is explained by assistant render reference echo after AEC
- THEN the system SHALL NOT emit
interrupt_detected
Scenario: Cancellation is idempotent
- WHEN interruption or recovery triggers cancellation more than once
- THEN LLM, TTS, playback, STT, and cancellable tool tasks SHALL handle repeated cancellation without unhandled exceptions
Scenario: Interrupted assistant text is stored
- WHEN assistant playback is interrupted before the full response is spoken
- THEN only assistant text that was fully spoken before interruption SHALL be eligible for short-term context or long-term memory
Requirement: Streaming STT and realtime transcript
The full-duplex mode SHALL use a streaming speech-to-text provider that emits partial, stable partial, and final transcripts while the user is speaking.
Scenario: Partial transcript is produced
- WHEN the user is speaking and streaming STT emits a changed partial transcript
- THEN the event bus SHALL emit a realtime transcript event for user-visible feedback
Scenario: Stable partial is produced
- WHEN a partial transcript remains stable according to the configured stability filter
- THEN the system SHALL mark it as stable partial while still withholding it from the LLM unless finalization rules allow submission
Scenario: Final transcript is produced
- WHEN the current user utterance ends
- THEN streaming STT SHALL emit a final transcript that becomes the user message candidate for conversation processing
Scenario: Empty transcript is produced
- WHEN final STT returns empty text, punctuation-only text, or text below confidence thresholds
- THEN the system SHALL discard the utterance, avoid invoking the LLM, and recover to listening
Scenario: Streaming STT provider changes
- WHEN the configured provider changes from
faster_whispertosensevoiceorsherpa_onnx - THEN the state machine and Conversation Manager SHALL keep the same transcript event contract
Requirement: Streaming response and TTS playback
The full-duplex mode SHALL stream LLM output into sentence segmentation, TTS synthesis, and playback without waiting for the full assistant reply to complete.
Scenario: First sentence is ready
- WHEN the LLM stream produces a complete sentence or configured safe speech segment
- THEN the sentence segmenter SHALL send sanitized text to the streaming TTS provider immediately
Scenario: TTS chunk is ready
- WHEN the TTS provider produces the first playable PCM chunk
- THEN playback SHALL begin without waiting for all remaining assistant text
Scenario: TTS text contains emoji
- WHEN LLM output contains emoji, sticker markers, Markdown image syntax, or other non-speech content
- THEN the text sent to TTS SHALL be sanitized before synthesis
Scenario: Playback is cancelled
- WHEN cancellation is requested during response playback
- THEN playback SHALL stop at the next configured audio chunk boundary and SHALL clear unplayed audio chunks
Scenario: TTS provider fails
- WHEN streaming TTS fails before or during playback
- THEN the system SHALL emit a structured TTS error and enter recovery without appending unspoken text to context
Requirement: Conversation manager for agent turns
The system SHALL define a Conversation Manager that coordinates short-term context, long-term memory retrieval, LLM streaming, tool calls, response playback, interruption, and recovery.
Scenario: User final transcript is accepted
- WHEN a valid final transcript is produced
- THEN the Conversation Manager SHALL create a user message candidate and gather short-term context and relevant long-term memory before invoking the LLM
Scenario: Assistant response completes
- WHEN the assistant response is fully spoken and no tool call or interruption is pending
- THEN the Conversation Manager SHALL commit the spoken assistant text to short-term context
Scenario: Response is interrupted
- WHEN user speech interrupts the assistant response
- THEN the Conversation Manager SHALL cancel the current response, preserve only committed spoken text, and process buffered user speech as the next input
Scenario: Conversation manager recovers
- WHEN a provider error occurs
- THEN the Conversation Manager SHALL emit a recoverable error event and return the agent to a safe listening or idle state when possible
Requirement: Long-term memory with FAISS and SQLite
The system SHALL define a long-term memory layer where SQLite stores text records and metadata while FAISS stores vector indexes for semantic retrieval.
Scenario: Memory is saved
- WHEN the user explicitly asks the assistant to remember a non-sensitive preference, fact, project detail, or task summary
- THEN the Memory Manager SHALL save a typed memory record with metadata in SQLite and an embedding vector in FAISS
Scenario: Memory is retrieved
- WHEN a later user input is semantically related to stored memories
- THEN the Memory Manager SHALL retrieve Top-K relevant memories and provide them to the Conversation Manager as separate memory context
Scenario: Runtime restarts
- WHEN the assistant process restarts and memory is enabled
- THEN saved non-sensitive memories SHALL remain available through SQLite and FAISS
Scenario: Memory is disabled
- WHEN
OWNER_MEMORY_ENABLED=0 - THEN the system SHALL NOT read from or write to SQLite or FAISS memory storage
Scenario: Sensitive content is detected
- WHEN a memory candidate contains credentials, payment data, private identifiers, account secrets, or other sensitive content
- THEN the system SHALL NOT automatically save it and SHALL require explicit user confirmation or reject the save
Scenario: Memory index is inconsistent
- WHEN SQLite metadata and FAISS index records are missing, mismatched, or fail checksum validation
- THEN the Memory Manager SHALL report a health error and SHALL NOT return unverified memory records
Requirement: Tool Router and structured tool execution
The system SHALL define a Tool Router that validates structured tool calls, applies security policy, routes approved tools, enforces execution limits, and returns sanitized tool results to the LLM.
Scenario: Allowed memory search runs
- WHEN the LLM requests
memory.searchwith valid arguments - THEN the Tool Router SHALL execute the memory search within configured limits and return sanitized results
Scenario: Tool schema is invalid
- WHEN a tool call has an unknown name, missing required argument, invalid type, or unknown field
- THEN the Tool Router SHALL reject it with a structured validation error
Scenario: Tool exceeds call budget
- WHEN a turn exceeds the configured maximum number of tool calls or total tool runtime
- THEN the Tool Router SHALL reject further tool calls for that turn and return a budget error
Scenario: Tool output is too long
- WHEN a tool returns output larger than the configured limit
- THEN the Tool Router SHALL truncate the output, mark it as truncated, and avoid sending full oversized content to the LLM
Scenario: Tool execution fails
- WHEN an approved tool raises an error or times out
- THEN the Tool Router SHALL return a structured tool error and the Conversation Manager SHALL recover or ask the user for next steps
Requirement: Safe first-version tools
The first full-duplex Agent tool set SHALL prioritize low-risk tools and SHALL include only memory.search, memory.save, shell.readonly, openinterpreter.run, and browser.playwright as planned first-version adapters.
Scenario: Read-only shell command is allowed
- WHEN
shell.readonlyreceives an allowlisted read-only command within allowed directories - THEN the tool SHALL execute with timeout and output truncation
Scenario: Write shell command is rejected
- WHEN
shell.readonlyreceives a command that writes, deletes, installs dependencies, changes permissions, uploads data, or mutates external state - THEN the Tool Router SHALL reject the command
Scenario: Memory save is requested
- WHEN
memory.saveis requested for a non-sensitive user-approved memory - THEN the tool SHALL save through the Memory Manager and return the new memory id
Scenario: High-risk action is requested
- WHEN any planned first-version tool request involves deletion, upload, payment, trading, account changes, permission changes, secret access, or arbitrary code execution
- THEN the Tool Router SHALL reject it or require explicit confirmation according to policy
Requirement: Open Interpreter external adapter
The system SHALL treat Open Interpreter as an optional external CLI or subprocess backend and SHALL NOT copy the local openinterpreter/ clone into the Owner source tree.
Scenario: Open Interpreter is disabled
- WHEN
OWNER_OPENINTERPRETER_ENABLED=0 - THEN the Tool Router SHALL mark
openinterpreter.rununavailable and SHALL NOT invoke any Open Interpreter command
Scenario: Open Interpreter path is missing
- WHEN
OWNER_OPENINTERPRETER_ENABLED=1but the configured command cannot be found - THEN the adapter SHALL return
OPENINTERPRETER_UNAVAILABLEwithout affecting other tools
Scenario: Low-risk Open Interpreter task is approved
- WHEN a low-risk Open Interpreter task passes directory, timeout, and argument validation
- THEN the adapter SHALL run it as an external process with bounded runtime and sanitized output
Scenario: High-risk Open Interpreter task is requested
- WHEN an Open Interpreter task would write files, delete data, change accounts, install dependencies, upload data, or execute broad arbitrary code
- THEN the Tool Router SHALL require explicit confirmation or reject the task in first-version policy
Requirement: Browser automation boundary
The system SHALL define browser.playwright as a planned browser automation adapter with explicit safety boundaries.
Scenario: Safe browser read task is requested
- WHEN the user asks for a low-risk browser read or extraction task and Playwright is enabled
- THEN the adapter SHALL run in an approved browser context with timeout, navigation limits, and sanitized output
Scenario: Sensitive browser action is requested
- WHEN a browser task involves login, payment, purchase, form submission, account changes, private messages, uploads, or irreversible actions
- THEN the adapter SHALL require explicit confirmation or reject the task according to policy
Scenario: Playwright is unavailable
- WHEN Playwright or its browser runtime is not installed
- THEN
browser.playwrightSHALL report a structured unavailable error and SHALL NOT affect voice conversation
Requirement: Computer control reservation
The system SHALL reserve a future ComputerControlProvider boundary for GUI control while excluding GUI click, keyboard, screen, and Accessibility automation from the first full-duplex Agent implementation.
Scenario: GUI control is requested in first version
- WHEN a tool call requests direct GUI click, keyboard input, screen control, or Accessibility automation
- THEN the first-version Tool Router SHALL reject it as unsupported
Scenario: Future provider is designed
- WHEN a future change implements computer control
- THEN it SHALL use public capabilities such as macOS Accessibility, Playwright, or trycua-style providers and SHALL define separate confirmation and audit requirements
Scenario: Codex Computer Use is referenced
- WHEN design work references Codex Computer Use behavior
- THEN the system SHALL use it only as a safety policy reference and SHALL NOT copy private or bundled implementation details
Requirement: Agent security and privacy
The full-duplex Agent architecture SHALL protect credentials, raw audio, memory data, and tool execution boundaries by default.
Scenario: Credentials are configured
- WHEN LLM, embedding, tool, or external adapter credentials are needed
- THEN they SHALL be read from local uncommitted configuration and SHALL NOT be hard-coded or logged
Scenario: Raw audio is processed
- WHEN microphone or render audio is processed for APM, VAD, STT, or interruption
- THEN raw audio and derived voice features SHALL remain process-local unless the user explicitly enables a documented recording feature
Scenario: Long-term memory stores text
- WHEN memory is saved
- THEN the memory record SHALL include sensitivity metadata and SHALL NOT include raw audio or credentials
Scenario: Tool audit is recorded
- WHEN a tool executes
- THEN the audit record SHALL include tool name, risk level, confirmation status, duration, status, and sanitized summary without secrets
Scenario: User disables tools
- WHEN
OWNER_TOOL_ROUTER_ENABLED=0 - THEN no tool calls SHALL execute and LLM tool requests SHALL be rejected or ignored with a structured message
Requirement: Full-duplex performance targets
The full-duplex Agent architecture SHALL define measurable latency and reliability targets for audio processing, interruption, streaming STT, LLM-to-TTS, and tool routing.
Scenario: Audio frame is processed
- WHEN WebRTC APM receives a normal 20 ms frame
- THEN processing SHALL target completion before the next frame interval in normal local runtime conditions
Scenario: User begins speaking during playback
- WHEN valid user speech begins while assistant speech is playing
- THEN playback stop SHALL target P95 latency under 200 ms from speech start
Scenario: User starts speaking in listening mode
- WHEN the user begins speaking while the system is listening
- THEN streaming STT SHALL target first stable user-visible transcript within 800 ms P95
Scenario: LLM produces first sentence
- WHEN the LLM stream yields a complete sentence
- THEN TTS SHALL target first playable audio chunk within 1000 ms P95
Scenario: Tool call runs
- WHEN an approved low-risk tool runs
- THEN it SHALL obey configured timeout, output limit, and maximum calls per turn
Requirement: Full-duplex testability
The full-duplex Agent architecture SHALL be testable with fake audio devices, fake WebRTC APM reference audio, fake VAD, fake streaming STT/TTS, fake memory stores, fake tools, and deterministic virtual time.
Scenario: Fake APM suppresses echo
- WHEN a test injects playback reference audio and matching microphone echo
- THEN the fake or real APM test SHALL verify that echo does not trigger user transcript or interruption
Scenario: Speaking interruption is tested
- WHEN a test injects user speech during assistant playback
- THEN the state sequence SHALL include
speaking -> interrupted -> listeningand active LLM/TTS/playback tasks SHALL be cancelled
Scenario: Streaming STT and TTS are tested
- WHEN fake providers emit partial transcripts, final transcript, LLM deltas, sentence segments, and TTS chunks
- THEN tests SHALL verify event order, context writes, and playback order
Scenario: Memory persists across restart
- WHEN a memory is saved and a new runtime instance starts with memory enabled
- THEN the memory SHALL be retrievable from SQLite and FAISS
Scenario: Tools are tested
- WHEN fake tool calls cover allowed, rejected, confirmation-required, timeout, and oversized output paths
- THEN Tool Router tests SHALL verify decisions, audit events, and sanitized tool results
Scenario: OpenSpec validation runs
- WHEN this planning change is complete
- THEN
openspec validate add-full-duplex-agent-voice-assistant --strictandopenspec validate --all --strictSHALL pass
Requirement: Live run-agent-live runtime
The system SHALL provide a real run-agent-live runtime entry point that starts continuous listening, handles user speech without a wake word, streams the assistant reply through interruptible playback, and returns to listening after completion or interruption.
Scenario: Agent live runtime starts
- WHEN the user runs
.venv/bin/python -m owner_voice_pet run-agent-live - THEN the command SHALL start the full-duplex Agent runtime instead of returning
FULL_DUPLEX_RUNTIME_NOT_IMPLEMENTED
Scenario: Agent check-config remains diagnostic
- WHEN the user runs
.venv/bin/python -m owner_voice_pet run-agent-live --check-config - THEN the command SHALL validate and print non-secret full-duplex configuration without opening the microphone
Scenario: Legacy live runtime remains available
- WHEN the user runs
.venv/bin/python -m owner_voice_pet run-live - THEN the existing wake-word turn-based runtime SHALL remain available during migration
Requirement: Software render-reference interruption gate
The first live full-duplex runtime SHALL use software render-reference gating to prevent assistant playback from triggering user interruption while still stopping playback quickly when real user speech is detected.
Scenario: User interrupts assistant playback
- WHEN valid user speech is detected during assistant playback and it is not classified as assistant render-reference echo
- THEN playback SHALL stop at the next configured chunk boundary and the captured user frames SHALL be retained for the next STT pass
Scenario: STT partial is unavailable during interruption
- WHEN VAD and render-reference gating indicate valid user speech but realtime STT has not produced a partial transcript yet
- THEN the runtime SHALL still stop playback and SHALL use final STT on the buffered user audio afterward
Scenario: Assistant echo is present
- WHEN microphone input during playback matches recent assistant render-reference audio and no user speech is present
- THEN the runtime SHALL NOT emit
barge_in_detectedand SHALL continue playback
Scenario: Interrupted assistant text is committed
- WHEN playback is stopped before a sentence is fully played
- THEN only fully played assistant text SHALL be written to short-term context