[测试、性能、安全、验收与归档]:完成全量测试验收与 OpenSpec 归档,包含 CLI 验收、NewAPI smoke、安全扫描和主 spec 更新
This commit is contained in:
+6
-6
@@ -43,9 +43,9 @@
|
||||
|
||||
## 6. 测试、性能、安全、验收与归档
|
||||
|
||||
- [ ] 6.1 完成自动化测试套件;前置条件:所有模块已实现;验收标准:覆盖核心模型、配置、Transport、Wake/VAD/STT、Context、LLM/TTS、Pipeline、UI、资产;测试要点:`python3.11 -m unittest discover -s tests` 通过;优先级:P0;预计:60 分钟。
|
||||
- [ ] 6.2 完成端到端验收命令;前置条件:Pipeline 可用;验收标准:fixture 音频链路可从“小杰小杰”走到播放输出和上下文更新;测试要点:CLI acceptance 命令退出码为 0;优先级:P0;预计:45 分钟。
|
||||
- [ ] 6.3 完成 OpenAI/NewAPI smoke 验收;前置条件:提供临时 API key 环境变量;验收标准:调用配置的 base URL 获得非空 LLM 回复,失败时输出结构化诊断且不泄露 key;测试要点:真实 smoke 或可解释的网络失败证据;优先级:P0;预计:45 分钟。
|
||||
- [ ] 6.4 完成安全检查;前置条件:全部代码已实现;验收标准:仓库中不包含 API key,日志脱敏,默认不持久化原始音频;测试要点:secret grep、配置测试和日志脱敏测试通过;优先级:P0;预计:30 分钟。
|
||||
- [ ] 6.5 完成 OpenSpec archive;前置条件:所有任务完成并验证通过;验收标准:`openspec archive add-voice-pet-pipeline` 后主 spec 更新,change 进入 archive;测试要点:`openspec validate --all --strict` 通过;优先级:P0;预计:30 分钟。
|
||||
- [ ] 6.6 完成最终验收提交;前置条件:6.1 至 6.5 已完成;验收标准:完整测试、OpenSpec strict 校验、git status 审计通过后立即 commit;测试要点:提交信息使用“`[测试、性能、安全、验收与归档]:完成[具体功能描述],包含[关键变更]`”格式;优先级:P0;预计:20 分钟。
|
||||
- [x] 6.1 完成自动化测试套件;前置条件:所有模块已实现;验收标准:覆盖核心模型、配置、Transport、Wake/VAD/STT、Context、LLM/TTS、Pipeline、UI、资产;测试要点:`python3.11 -m unittest discover -s tests` 通过;优先级:P0;预计:60 分钟。
|
||||
- [x] 6.2 完成端到端验收命令;前置条件:Pipeline 可用;验收标准:fixture 音频链路可从“小杰小杰”走到播放输出和上下文更新;测试要点:CLI acceptance 命令退出码为 0;优先级:P0;预计:45 分钟。
|
||||
- [x] 6.3 完成 OpenAI/NewAPI smoke 验收;前置条件:提供临时 API key 环境变量;验收标准:调用配置的 base URL 获得非空 LLM 回复,失败时输出结构化诊断且不泄露 key;测试要点:真实 smoke 或可解释的网络失败证据;优先级:P0;预计:45 分钟。
|
||||
- [x] 6.4 完成安全检查;前置条件:全部代码已实现;验收标准:仓库中不包含 API key,日志脱敏,默认不持久化原始音频;测试要点:secret grep、配置测试和日志脱敏测试通过;优先级:P0;预计:30 分钟。
|
||||
- [x] 6.5 完成 OpenSpec archive;前置条件:所有任务完成并验证通过;验收标准:`openspec archive add-voice-pet-pipeline` 后主 spec 更新,change 进入 archive;测试要点:`openspec validate --all --strict` 通过;优先级:P0;预计:30 分钟。
|
||||
- [x] 6.6 完成最终验收提交;前置条件:6.1 至 6.5 已完成;验收标准:完整测试、OpenSpec strict 校验、git status 审计通过后立即 commit;测试要点:提交信息使用“`[测试、性能、安全、验收与归档]:完成[具体功能描述],包含[关键变更]`”格式;优先级:P0;预计:20 分钟。
|
||||
@@ -0,0 +1,268 @@
|
||||
# voice-pet-pipeline Specification
|
||||
|
||||
## Purpose
|
||||
TBD - created by archiving change add-voice-pet-pipeline. Update Purpose after archive.
|
||||
## Requirements
|
||||
### Requirement: Python desktop pet runtime
|
||||
The system SHALL be specified as a local Python desktop pet application that owns the voice pipeline, desktop pet state, local audio input, and local audio output in the first version.
|
||||
|
||||
#### Scenario: App starts in local desktop mode
|
||||
- **WHEN** the user starts the future application on macOS
|
||||
- **THEN** the system SHALL initialize as a local desktop pet process rather than a remote service or browser-only tool
|
||||
|
||||
#### Scenario: Implementation follows OpenSpec artifacts
|
||||
- **WHEN** this OpenSpec change is implemented
|
||||
- **THEN** the repository SHALL contain Python source code, automated tests, validation commands, and project-local pet assets aligned with the planning artifacts
|
||||
|
||||
### Requirement: Local microphone and speaker transport
|
||||
The system SHALL use the local microphone as the first-version input Transport and the local system speaker as the first-version output Transport.
|
||||
|
||||
#### Scenario: Microphone input is available
|
||||
- **WHEN** the configured microphone is available and permitted
|
||||
- **THEN** the Transport SHALL provide streaming audio frames to the wakeword and VAD stages
|
||||
|
||||
#### Scenario: Speaker output is available
|
||||
- **WHEN** TTS returns a valid audio segment and the configured speaker is available
|
||||
- **THEN** the Transport SHALL play the audio segment through the local speaker
|
||||
|
||||
#### Scenario: Audio device is unavailable
|
||||
- **WHEN** the microphone or speaker is missing, denied, or unsupported
|
||||
- **THEN** the system SHALL expose a recoverable Transport error with a stable error code
|
||||
|
||||
### Requirement: Wake word detection
|
||||
The system SHALL listen locally for the Chinese wake word “小杰小杰” before accepting user speech for a conversation turn.
|
||||
|
||||
#### Scenario: Wake word is detected
|
||||
- **WHEN** the user says “小杰小杰” and the wakeword provider returns confidence above the configured threshold
|
||||
- **THEN** the pipeline SHALL transition from wake listening to speech detection
|
||||
|
||||
#### Scenario: Wake word is not detected
|
||||
- **WHEN** background speech or noise does not match “小杰小杰”
|
||||
- **THEN** the pipeline SHALL remain in wake listening and SHALL NOT invoke STT, LLM, or TTS
|
||||
|
||||
#### Scenario: Wake model fails
|
||||
- **WHEN** the wakeword provider cannot load or process audio
|
||||
- **THEN** the system SHALL report a wakeword error and SHALL NOT crash the desktop pet process
|
||||
|
||||
### Requirement: VAD speech endpoint detection
|
||||
After wakeword detection, the system SHALL use VAD to identify when the user starts and stops speaking.
|
||||
|
||||
#### Scenario: User begins speaking
|
||||
- **WHEN** VAD detects continuous speech above the configured start threshold
|
||||
- **THEN** the pipeline SHALL begin recording the current utterance
|
||||
|
||||
#### Scenario: User stops speaking
|
||||
- **WHEN** VAD detects continuous silence above the configured end threshold
|
||||
- **THEN** the pipeline SHALL close the current audio segment and send it to STT
|
||||
|
||||
#### Scenario: User says nothing after wakeword
|
||||
- **WHEN** no speech is detected before the configured no-speech timeout
|
||||
- **THEN** the pipeline SHALL return to wake listening without invoking STT or LLM
|
||||
|
||||
#### Scenario: Recording exceeds maximum duration
|
||||
- **WHEN** speech continues beyond the configured maximum utterance duration
|
||||
- **THEN** the pipeline SHALL end the segment, mark the end reason, and continue to STT with the captured audio
|
||||
|
||||
### Requirement: Local STT transcription
|
||||
The system SHALL transcribe the captured user utterance through a local STT provider, with `sherpa-onnx` as the recommended first implementation candidate.
|
||||
|
||||
#### Scenario: STT succeeds
|
||||
- **WHEN** STT returns non-empty text for a captured audio segment
|
||||
- **THEN** the pipeline SHALL add the trimmed text as a user message to the conversation context
|
||||
|
||||
#### Scenario: STT returns empty text
|
||||
- **WHEN** STT returns empty text, punctuation-only text, or an invalid transcript
|
||||
- **THEN** the pipeline SHALL skip LLM invocation and return to wake listening with a recoverable status
|
||||
|
||||
#### Scenario: STT provider fails
|
||||
- **WHEN** the STT provider raises an error or cannot load its local model
|
||||
- **THEN** the system SHALL emit an STT error code and SHALL recover to a state where future wake attempts are possible
|
||||
|
||||
### Requirement: Conversation context management
|
||||
The system SHALL maintain an in-memory conversation context for the active desktop pet session.
|
||||
|
||||
#### Scenario: User transcript is accepted
|
||||
- **WHEN** a valid STT transcript is produced
|
||||
- **THEN** the system SHALL append it to the context as a user message before invoking the LLM
|
||||
|
||||
#### Scenario: Assistant reply completes
|
||||
- **WHEN** the LLM and TTS stages complete a reply
|
||||
- **THEN** the system SHALL append the assistant text to the context
|
||||
|
||||
#### Scenario: Context exceeds configured budget
|
||||
- **WHEN** the context exceeds the configured message, character, or token budget
|
||||
- **THEN** the system SHALL preserve the system prompt and most recent conversation turns while removing older ordinary messages
|
||||
|
||||
### Requirement: Cloud LLM streaming reply
|
||||
The system SHALL use a cloud LLM provider for reply generation, with OpenAI Responses API streaming as the recommended first implementation.
|
||||
|
||||
#### Scenario: LLM stream begins
|
||||
- **WHEN** the LLM provider receives valid context messages and configuration
|
||||
- **THEN** it SHALL stream reply deltas rather than waiting for the full reply before producing output
|
||||
|
||||
#### Scenario: LLM model is configured
|
||||
- **WHEN** the system starts
|
||||
- **THEN** the LLM model name SHALL be read from configuration and SHALL NOT be hard-coded in pipeline logic
|
||||
|
||||
#### Scenario: LLM request fails
|
||||
- **WHEN** the LLM provider times out, is rate limited, lacks credentials, or encounters a network error
|
||||
- **THEN** the pipeline SHALL emit a structured LLM error and transition through error recovery to wake listening
|
||||
|
||||
### Requirement: Local TTS synthesis and playback
|
||||
The system SHALL synthesize assistant replies through a local TTS provider, with `sherpa-onnx` as the recommended first implementation candidate.
|
||||
|
||||
#### Scenario: Sentence is ready for speech
|
||||
- **WHEN** the LLM stream produces a complete sentence or configured text chunk
|
||||
- **THEN** the TTS provider SHALL synthesize that text into an audio segment
|
||||
|
||||
#### Scenario: First TTS segment is ready
|
||||
- **WHEN** the first synthesized audio segment is available
|
||||
- **THEN** the output Transport SHALL begin playback without waiting for every remaining LLM delta
|
||||
|
||||
#### Scenario: TTS fails
|
||||
- **WHEN** TTS returns empty audio or raises an error
|
||||
- **THEN** the pipeline SHALL emit a structured TTS error and SHALL recover without terminating the application
|
||||
|
||||
### Requirement: Pipeline state machine
|
||||
The system SHALL expose deterministic pipeline states for `idle`, `wake_listening`, `speech_detecting`, `recording`, `transcribing`, `thinking`, `speaking`, `interrupted`, and `error_recovering`.
|
||||
|
||||
#### Scenario: Normal conversation turn
|
||||
- **WHEN** wakeword, VAD, STT, LLM, TTS, and playback all succeed
|
||||
- **THEN** the state sequence SHALL progress through wake listening, speech detection, recording, transcribing, thinking, speaking, and back to wake listening
|
||||
|
||||
#### Scenario: Recoverable provider error
|
||||
- **WHEN** any provider fails during a conversation turn
|
||||
- **THEN** the state machine SHALL enter error recovery and then return to wake listening after cleanup
|
||||
|
||||
#### Scenario: UI subscribes to state
|
||||
- **WHEN** the pipeline state changes
|
||||
- **THEN** the desktop pet UI SHALL be able to update its visible state without directly invoking audio or model providers
|
||||
|
||||
### Requirement: Desktop pet visual states
|
||||
The system SHALL define desktop pet visual states for idle, listening, recording, thinking, speaking, and error feedback.
|
||||
|
||||
#### Scenario: Wakeword is detected
|
||||
- **WHEN** the pipeline enters speech detection or recording
|
||||
- **THEN** the desktop pet SHALL show a listening or recording visual state
|
||||
|
||||
#### Scenario: LLM is generating
|
||||
- **WHEN** the pipeline enters thinking
|
||||
- **THEN** the desktop pet SHALL show a thinking visual state
|
||||
|
||||
#### Scenario: TTS is playing
|
||||
- **WHEN** the pipeline enters speaking
|
||||
- **THEN** the desktop pet SHALL show a speaking visual state
|
||||
|
||||
#### Scenario: Error occurs
|
||||
- **WHEN** the pipeline enters error recovery
|
||||
- **THEN** the desktop pet SHALL show a concise visible error state and then return to idle or wake listening when recovered
|
||||
|
||||
### Requirement: Generated pet asset specification
|
||||
The system SHALL specify that production desktop pet images are generated with the image generation skill when usable, or with a reproducible project-local fallback when image generation output is unavailable or fails validation, and saved inside the project as transparent PNG assets.
|
||||
|
||||
#### Scenario: Image generation output is usable
|
||||
- **WHEN** pet images are generated with `imagegen` and pass role and transparency validation
|
||||
- **THEN** the images SHALL be processed for transparency, validated, and saved to a project asset directory
|
||||
|
||||
#### Scenario: Image generation output is unavailable or invalid
|
||||
- **WHEN** image generation output is unavailable, inaccessible as a project file, or fails role/transparency validation
|
||||
- **THEN** the system SHALL use a reproducible project-local fallback to create transparent PNG pet state assets and SHALL validate them before use
|
||||
|
||||
#### Scenario: Planning phase is executed
|
||||
- **WHEN** this OpenSpec change is implemented
|
||||
- **THEN** generated or generated-derived pet assets SHALL be saved inside the project and SHALL NOT be referenced from a temporary generation directory
|
||||
|
||||
### Requirement: Audio feedback suppression
|
||||
The system SHALL prevent the desktop pet's own TTS playback from being treated as a new user wakeword or speech input.
|
||||
|
||||
#### Scenario: TTS playback is active
|
||||
- **WHEN** the output Transport is playing synthesized assistant speech
|
||||
- **THEN** wakeword detection and VAD input processing SHALL be suppressed or ignored for that playback window by default
|
||||
|
||||
#### Scenario: User speaks during playback
|
||||
- **WHEN** the user speaks while TTS playback is active in the first version
|
||||
- **THEN** the system SHALL ignore that speech unless a future interruption feature is explicitly specified
|
||||
|
||||
### Requirement: Structured errors and logging
|
||||
The system SHALL represent stage failures with structured error codes and log state transitions, latency metrics, and provider names.
|
||||
|
||||
#### Scenario: Stage transition occurs
|
||||
- **WHEN** the pipeline changes state
|
||||
- **THEN** the system SHALL log the turn ID, old state, new state, timestamp, and triggering event
|
||||
|
||||
#### Scenario: Provider call completes
|
||||
- **WHEN** a provider call succeeds or fails
|
||||
- **THEN** the system SHALL log provider name, stage, duration, and error code if present
|
||||
|
||||
#### Scenario: Sensitive data is present
|
||||
- **WHEN** logs are written
|
||||
- **THEN** the system SHALL NOT log API keys, authorization headers, or raw credentials
|
||||
|
||||
### Requirement: Performance targets
|
||||
The system SHALL define measurable first-version latency and reliability targets for wakeword response, endpoint detection, LLM first output, and speech playback.
|
||||
|
||||
#### Scenario: Wakeword latency is measured
|
||||
- **WHEN** wakeword detection succeeds in a normal local environment
|
||||
- **THEN** the system SHALL target visible listening feedback within 800 ms
|
||||
|
||||
#### Scenario: User stops speaking
|
||||
- **WHEN** the user stops speaking after a normal utterance
|
||||
- **THEN** VAD endpoint detection SHALL target transition to transcription within 900 ms
|
||||
|
||||
#### Scenario: LLM begins responding
|
||||
- **WHEN** a valid transcript has been sent to the LLM provider
|
||||
- **THEN** the system SHALL target first playable reply text within 3.5 seconds
|
||||
|
||||
#### Scenario: Normal reply is spoken
|
||||
- **WHEN** a normal Chinese question under 10 seconds is processed
|
||||
- **THEN** the system SHALL target the start of spoken playback within 5 seconds P95 after the user stops speaking
|
||||
|
||||
### Requirement: Security and privacy
|
||||
The system SHALL protect credentials and minimize audio persistence.
|
||||
|
||||
#### Scenario: API key is required
|
||||
- **WHEN** the LLM provider needs an OpenAI API key
|
||||
- **THEN** the key SHALL be loaded from environment or local uncommitted configuration and SHALL NOT be committed
|
||||
|
||||
#### Scenario: Audio is processed
|
||||
- **WHEN** user speech is captured for STT
|
||||
- **THEN** raw audio SHALL NOT be persisted by default
|
||||
|
||||
#### Scenario: Cloud LLM is invoked
|
||||
- **WHEN** the system sends data to the cloud LLM
|
||||
- **THEN** it SHALL send transcript text and necessary conversation context, not raw microphone audio
|
||||
|
||||
### Requirement: Testability
|
||||
The system SHALL be designed so each stage can be tested with mock providers and file-based audio fixtures.
|
||||
|
||||
#### Scenario: Pipeline is unit tested
|
||||
- **WHEN** mock providers produce deterministic events
|
||||
- **THEN** tests SHALL verify state transitions, context updates, and error recovery
|
||||
|
||||
#### Scenario: Audio fixture is replayed
|
||||
- **WHEN** a file-based audio fixture containing “小杰小杰” and user speech is replayed in a test Transport
|
||||
- **THEN** the pipeline SHALL be testable without using a live microphone
|
||||
|
||||
#### Scenario: OpenSpec planning validation runs
|
||||
- **WHEN** this planning change is complete
|
||||
- **THEN** `openspec validate add-voice-pet-pipeline --strict` and `openspec validate --all --strict` SHALL pass
|
||||
|
||||
### Requirement: Module-level git commit gates
|
||||
The system implementation process SHALL require an immediate Git commit after each major module or milestone is completed, where a major module means a top-level task group from `tasks.md` or a phase milestone from `proposal.md`.
|
||||
|
||||
#### Scenario: Major module is completed
|
||||
- **WHEN** an implementer completes a top-level task group or milestone phase
|
||||
- **THEN** the implementer SHALL run the applicable build, tests, lint/type checks, and OpenSpec validation before committing
|
||||
|
||||
#### Scenario: Build validation fails before commit
|
||||
- **WHEN** the applicable build, tests, lint/type checks, or OpenSpec validation fail
|
||||
- **THEN** the implementer SHALL NOT create the module commit until the failure is fixed and validation passes
|
||||
|
||||
#### Scenario: Module commit is created
|
||||
- **WHEN** validation passes for the completed module
|
||||
- **THEN** the implementer SHALL immediately create a Git commit using the Chinese message format “[模块名]:完成[具体功能描述],包含[关键变更]”
|
||||
|
||||
#### Scenario: Intermediate work remains after module completion
|
||||
- **WHEN** module-related changes remain unstaged or uncommitted after the module is completed
|
||||
- **THEN** the implementer SHALL include those changes in the module commit or explicitly separate unrelated external changes before continuing to the next module
|
||||
|
||||
Reference in New Issue
Block a user