[测试、性能、安全、验收与归档]:完成全量测试验收与 OpenSpec 归档,包含 CLI 验收、NewAPI smoke、安全扫描和主 spec 更新

This commit is contained in:
mkbk
2026-06-17 18:42:11 +08:00
parent 667c8941f2
commit 1b884ac6a9
15 changed files with 522 additions and 28 deletions
@@ -43,9 +43,9 @@
## 6. 测试、性能、安全、验收与归档
- [ ] 6.1 完成自动化测试套件;前置条件:所有模块已实现;验收标准:覆盖核心模型、配置、Transport、Wake/VAD/STT、Context、LLM/TTS、Pipeline、UI、资产;测试要点:`python3.11 -m unittest discover -s tests` 通过;优先级:P0;预计:60 分钟。
- [ ] 6.2 完成端到端验收命令;前置条件:Pipeline 可用;验收标准:fixture 音频链路可从“小杰小杰”走到播放输出和上下文更新;测试要点:CLI acceptance 命令退出码为 0;优先级:P0;预计:45 分钟。
- [ ] 6.3 完成 OpenAI/NewAPI smoke 验收;前置条件:提供临时 API key 环境变量;验收标准:调用配置的 base URL 获得非空 LLM 回复,失败时输出结构化诊断且不泄露 key;测试要点:真实 smoke 或可解释的网络失败证据;优先级:P0;预计:45 分钟。
- [ ] 6.4 完成安全检查;前置条件:全部代码已实现;验收标准:仓库中不包含 API key,日志脱敏,默认不持久化原始音频;测试要点:secret grep、配置测试和日志脱敏测试通过;优先级:P0;预计:30 分钟。
- [ ] 6.5 完成 OpenSpec archive;前置条件:所有任务完成并验证通过;验收标准:`openspec archive add-voice-pet-pipeline` 后主 spec 更新,change 进入 archive;测试要点:`openspec validate --all --strict` 通过;优先级:P0;预计:30 分钟。
- [ ] 6.6 完成最终验收提交;前置条件:6.1 至 6.5 已完成;验收标准:完整测试、OpenSpec strict 校验、git status 审计通过后立即 commit;测试要点:提交信息使用“`[测试、性能、安全、验收与归档]:完成[具体功能描述],包含[关键变更]`”格式;优先级:P0;预计:20 分钟。
- [x] 6.1 完成自动化测试套件;前置条件:所有模块已实现;验收标准:覆盖核心模型、配置、Transport、Wake/VAD/STT、Context、LLM/TTS、Pipeline、UI、资产;测试要点:`python3.11 -m unittest discover -s tests` 通过;优先级:P0;预计:60 分钟。
- [x] 6.2 完成端到端验收命令;前置条件:Pipeline 可用;验收标准:fixture 音频链路可从“小杰小杰”走到播放输出和上下文更新;测试要点:CLI acceptance 命令退出码为 0;优先级:P0;预计:45 分钟。
- [x] 6.3 完成 OpenAI/NewAPI smoke 验收;前置条件:提供临时 API key 环境变量;验收标准:调用配置的 base URL 获得非空 LLM 回复,失败时输出结构化诊断且不泄露 key;测试要点:真实 smoke 或可解释的网络失败证据;优先级:P0;预计:45 分钟。
- [x] 6.4 完成安全检查;前置条件:全部代码已实现;验收标准:仓库中不包含 API key,日志脱敏,默认不持久化原始音频;测试要点:secret grep、配置测试和日志脱敏测试通过;优先级:P0;预计:30 分钟。
- [x] 6.5 完成 OpenSpec archive;前置条件:所有任务完成并验证通过;验收标准:`openspec archive add-voice-pet-pipeline` 后主 spec 更新,change 进入 archive;测试要点:`openspec validate --all --strict` 通过;优先级:P0;预计:30 分钟。
- [x] 6.6 完成最终验收提交;前置条件:6.1 至 6.5 已完成;验收标准:完整测试、OpenSpec strict 校验、git status 审计通过后立即 commit;测试要点:提交信息使用“`[测试、性能、安全、验收与归档]:完成[具体功能描述],包含[关键变更]`”格式;优先级:P0;预计:20 分钟。
+268
View File
@@ -0,0 +1,268 @@
# voice-pet-pipeline Specification
## Purpose
TBD - created by archiving change add-voice-pet-pipeline. Update Purpose after archive.
## Requirements
### Requirement: Python desktop pet runtime
The system SHALL be specified as a local Python desktop pet application that owns the voice pipeline, desktop pet state, local audio input, and local audio output in the first version.
#### Scenario: App starts in local desktop mode
- **WHEN** the user starts the future application on macOS
- **THEN** the system SHALL initialize as a local desktop pet process rather than a remote service or browser-only tool
#### Scenario: Implementation follows OpenSpec artifacts
- **WHEN** this OpenSpec change is implemented
- **THEN** the repository SHALL contain Python source code, automated tests, validation commands, and project-local pet assets aligned with the planning artifacts
### Requirement: Local microphone and speaker transport
The system SHALL use the local microphone as the first-version input Transport and the local system speaker as the first-version output Transport.
#### Scenario: Microphone input is available
- **WHEN** the configured microphone is available and permitted
- **THEN** the Transport SHALL provide streaming audio frames to the wakeword and VAD stages
#### Scenario: Speaker output is available
- **WHEN** TTS returns a valid audio segment and the configured speaker is available
- **THEN** the Transport SHALL play the audio segment through the local speaker
#### Scenario: Audio device is unavailable
- **WHEN** the microphone or speaker is missing, denied, or unsupported
- **THEN** the system SHALL expose a recoverable Transport error with a stable error code
### Requirement: Wake word detection
The system SHALL listen locally for the Chinese wake word “小杰小杰” before accepting user speech for a conversation turn.
#### Scenario: Wake word is detected
- **WHEN** the user says “小杰小杰” and the wakeword provider returns confidence above the configured threshold
- **THEN** the pipeline SHALL transition from wake listening to speech detection
#### Scenario: Wake word is not detected
- **WHEN** background speech or noise does not match “小杰小杰”
- **THEN** the pipeline SHALL remain in wake listening and SHALL NOT invoke STT, LLM, or TTS
#### Scenario: Wake model fails
- **WHEN** the wakeword provider cannot load or process audio
- **THEN** the system SHALL report a wakeword error and SHALL NOT crash the desktop pet process
### Requirement: VAD speech endpoint detection
After wakeword detection, the system SHALL use VAD to identify when the user starts and stops speaking.
#### Scenario: User begins speaking
- **WHEN** VAD detects continuous speech above the configured start threshold
- **THEN** the pipeline SHALL begin recording the current utterance
#### Scenario: User stops speaking
- **WHEN** VAD detects continuous silence above the configured end threshold
- **THEN** the pipeline SHALL close the current audio segment and send it to STT
#### Scenario: User says nothing after wakeword
- **WHEN** no speech is detected before the configured no-speech timeout
- **THEN** the pipeline SHALL return to wake listening without invoking STT or LLM
#### Scenario: Recording exceeds maximum duration
- **WHEN** speech continues beyond the configured maximum utterance duration
- **THEN** the pipeline SHALL end the segment, mark the end reason, and continue to STT with the captured audio
### Requirement: Local STT transcription
The system SHALL transcribe the captured user utterance through a local STT provider, with `sherpa-onnx` as the recommended first implementation candidate.
#### Scenario: STT succeeds
- **WHEN** STT returns non-empty text for a captured audio segment
- **THEN** the pipeline SHALL add the trimmed text as a user message to the conversation context
#### Scenario: STT returns empty text
- **WHEN** STT returns empty text, punctuation-only text, or an invalid transcript
- **THEN** the pipeline SHALL skip LLM invocation and return to wake listening with a recoverable status
#### Scenario: STT provider fails
- **WHEN** the STT provider raises an error or cannot load its local model
- **THEN** the system SHALL emit an STT error code and SHALL recover to a state where future wake attempts are possible
### Requirement: Conversation context management
The system SHALL maintain an in-memory conversation context for the active desktop pet session.
#### Scenario: User transcript is accepted
- **WHEN** a valid STT transcript is produced
- **THEN** the system SHALL append it to the context as a user message before invoking the LLM
#### Scenario: Assistant reply completes
- **WHEN** the LLM and TTS stages complete a reply
- **THEN** the system SHALL append the assistant text to the context
#### Scenario: Context exceeds configured budget
- **WHEN** the context exceeds the configured message, character, or token budget
- **THEN** the system SHALL preserve the system prompt and most recent conversation turns while removing older ordinary messages
### Requirement: Cloud LLM streaming reply
The system SHALL use a cloud LLM provider for reply generation, with OpenAI Responses API streaming as the recommended first implementation.
#### Scenario: LLM stream begins
- **WHEN** the LLM provider receives valid context messages and configuration
- **THEN** it SHALL stream reply deltas rather than waiting for the full reply before producing output
#### Scenario: LLM model is configured
- **WHEN** the system starts
- **THEN** the LLM model name SHALL be read from configuration and SHALL NOT be hard-coded in pipeline logic
#### Scenario: LLM request fails
- **WHEN** the LLM provider times out, is rate limited, lacks credentials, or encounters a network error
- **THEN** the pipeline SHALL emit a structured LLM error and transition through error recovery to wake listening
### Requirement: Local TTS synthesis and playback
The system SHALL synthesize assistant replies through a local TTS provider, with `sherpa-onnx` as the recommended first implementation candidate.
#### Scenario: Sentence is ready for speech
- **WHEN** the LLM stream produces a complete sentence or configured text chunk
- **THEN** the TTS provider SHALL synthesize that text into an audio segment
#### Scenario: First TTS segment is ready
- **WHEN** the first synthesized audio segment is available
- **THEN** the output Transport SHALL begin playback without waiting for every remaining LLM delta
#### Scenario: TTS fails
- **WHEN** TTS returns empty audio or raises an error
- **THEN** the pipeline SHALL emit a structured TTS error and SHALL recover without terminating the application
### Requirement: Pipeline state machine
The system SHALL expose deterministic pipeline states for `idle`, `wake_listening`, `speech_detecting`, `recording`, `transcribing`, `thinking`, `speaking`, `interrupted`, and `error_recovering`.
#### Scenario: Normal conversation turn
- **WHEN** wakeword, VAD, STT, LLM, TTS, and playback all succeed
- **THEN** the state sequence SHALL progress through wake listening, speech detection, recording, transcribing, thinking, speaking, and back to wake listening
#### Scenario: Recoverable provider error
- **WHEN** any provider fails during a conversation turn
- **THEN** the state machine SHALL enter error recovery and then return to wake listening after cleanup
#### Scenario: UI subscribes to state
- **WHEN** the pipeline state changes
- **THEN** the desktop pet UI SHALL be able to update its visible state without directly invoking audio or model providers
### Requirement: Desktop pet visual states
The system SHALL define desktop pet visual states for idle, listening, recording, thinking, speaking, and error feedback.
#### Scenario: Wakeword is detected
- **WHEN** the pipeline enters speech detection or recording
- **THEN** the desktop pet SHALL show a listening or recording visual state
#### Scenario: LLM is generating
- **WHEN** the pipeline enters thinking
- **THEN** the desktop pet SHALL show a thinking visual state
#### Scenario: TTS is playing
- **WHEN** the pipeline enters speaking
- **THEN** the desktop pet SHALL show a speaking visual state
#### Scenario: Error occurs
- **WHEN** the pipeline enters error recovery
- **THEN** the desktop pet SHALL show a concise visible error state and then return to idle or wake listening when recovered
### Requirement: Generated pet asset specification
The system SHALL specify that production desktop pet images are generated with the image generation skill when usable, or with a reproducible project-local fallback when image generation output is unavailable or fails validation, and saved inside the project as transparent PNG assets.
#### Scenario: Image generation output is usable
- **WHEN** pet images are generated with `imagegen` and pass role and transparency validation
- **THEN** the images SHALL be processed for transparency, validated, and saved to a project asset directory
#### Scenario: Image generation output is unavailable or invalid
- **WHEN** image generation output is unavailable, inaccessible as a project file, or fails role/transparency validation
- **THEN** the system SHALL use a reproducible project-local fallback to create transparent PNG pet state assets and SHALL validate them before use
#### Scenario: Planning phase is executed
- **WHEN** this OpenSpec change is implemented
- **THEN** generated or generated-derived pet assets SHALL be saved inside the project and SHALL NOT be referenced from a temporary generation directory
### Requirement: Audio feedback suppression
The system SHALL prevent the desktop pet's own TTS playback from being treated as a new user wakeword or speech input.
#### Scenario: TTS playback is active
- **WHEN** the output Transport is playing synthesized assistant speech
- **THEN** wakeword detection and VAD input processing SHALL be suppressed or ignored for that playback window by default
#### Scenario: User speaks during playback
- **WHEN** the user speaks while TTS playback is active in the first version
- **THEN** the system SHALL ignore that speech unless a future interruption feature is explicitly specified
### Requirement: Structured errors and logging
The system SHALL represent stage failures with structured error codes and log state transitions, latency metrics, and provider names.
#### Scenario: Stage transition occurs
- **WHEN** the pipeline changes state
- **THEN** the system SHALL log the turn ID, old state, new state, timestamp, and triggering event
#### Scenario: Provider call completes
- **WHEN** a provider call succeeds or fails
- **THEN** the system SHALL log provider name, stage, duration, and error code if present
#### Scenario: Sensitive data is present
- **WHEN** logs are written
- **THEN** the system SHALL NOT log API keys, authorization headers, or raw credentials
### Requirement: Performance targets
The system SHALL define measurable first-version latency and reliability targets for wakeword response, endpoint detection, LLM first output, and speech playback.
#### Scenario: Wakeword latency is measured
- **WHEN** wakeword detection succeeds in a normal local environment
- **THEN** the system SHALL target visible listening feedback within 800 ms
#### Scenario: User stops speaking
- **WHEN** the user stops speaking after a normal utterance
- **THEN** VAD endpoint detection SHALL target transition to transcription within 900 ms
#### Scenario: LLM begins responding
- **WHEN** a valid transcript has been sent to the LLM provider
- **THEN** the system SHALL target first playable reply text within 3.5 seconds
#### Scenario: Normal reply is spoken
- **WHEN** a normal Chinese question under 10 seconds is processed
- **THEN** the system SHALL target the start of spoken playback within 5 seconds P95 after the user stops speaking
### Requirement: Security and privacy
The system SHALL protect credentials and minimize audio persistence.
#### Scenario: API key is required
- **WHEN** the LLM provider needs an OpenAI API key
- **THEN** the key SHALL be loaded from environment or local uncommitted configuration and SHALL NOT be committed
#### Scenario: Audio is processed
- **WHEN** user speech is captured for STT
- **THEN** raw audio SHALL NOT be persisted by default
#### Scenario: Cloud LLM is invoked
- **WHEN** the system sends data to the cloud LLM
- **THEN** it SHALL send transcript text and necessary conversation context, not raw microphone audio
### Requirement: Testability
The system SHALL be designed so each stage can be tested with mock providers and file-based audio fixtures.
#### Scenario: Pipeline is unit tested
- **WHEN** mock providers produce deterministic events
- **THEN** tests SHALL verify state transitions, context updates, and error recovery
#### Scenario: Audio fixture is replayed
- **WHEN** a file-based audio fixture containing “小杰小杰” and user speech is replayed in a test Transport
- **THEN** the pipeline SHALL be testable without using a live microphone
#### Scenario: OpenSpec planning validation runs
- **WHEN** this planning change is complete
- **THEN** `openspec validate add-voice-pet-pipeline --strict` and `openspec validate --all --strict` SHALL pass
### Requirement: Module-level git commit gates
The system implementation process SHALL require an immediate Git commit after each major module or milestone is completed, where a major module means a top-level task group from `tasks.md` or a phase milestone from `proposal.md`.
#### Scenario: Major module is completed
- **WHEN** an implementer completes a top-level task group or milestone phase
- **THEN** the implementer SHALL run the applicable build, tests, lint/type checks, and OpenSpec validation before committing
#### Scenario: Build validation fails before commit
- **WHEN** the applicable build, tests, lint/type checks, or OpenSpec validation fail
- **THEN** the implementer SHALL NOT create the module commit until the failure is fixed and validation passes
#### Scenario: Module commit is created
- **WHEN** validation passes for the completed module
- **THEN** the implementer SHALL immediately create a Git commit using the Chinese message format “[模块名]:完成[具体功能描述],包含[关键变更]”
#### Scenario: Intermediate work remains after module completion
- **WHEN** module-related changes remain unstaged or uncommitted after the module is completed
- **THEN** the implementer SHALL include those changes in the module commit or explicitly separate unrelated external changes before continuing to the next module