4.4 KiB
4.4 KiB
独立本地唤醒与终端转写显示设计
Context
真实运行日志显示,当前 run-live 会先截取语音段并使用 STT 判断是否包含“小杰小杰”。该实现把 wake detection 与 user utterance transcription 绑定在一起,导致唤醒慢、唤醒词污染正式问题、终端输出难以区分 STT 和 LLM 阶段。本设计把 wake word detection 升级为本地 KWS 模型路径,并把正式问题转写作为独立可见事件输出。
Goals
- 唤醒词“小杰小杰”由本地模型检测。
- wake 阶段不调用云 ASR,也不调用正式 STT provider。
- wake 命中后才开始正式问题录音和 STT。
- 终端在 LLM 前显示正式问题转写文本。
- 模型下载和检查覆盖 wake KWS、VAD、STT。
- 自动化测试证明 wake/STT 分离、重复对话、上下文和错误恢复仍正常。
Non-Goals
- 不实现 GUI 桌宠窗口。
- 不实现跨进程长期记忆。
- 不在本阶段实现逐字 partial ASR 字幕;本阶段先保证正式问题 STT 完成后立即可见,且出现在 LLM 前。
- 不把连续麦克风流上传到云端。
Architecture
run-live
-> AppConfig(.env)
-> SoundDeviceAudioTransport
-> SherpaOnnxKeywordWakeWordProvider(local models/wake)
-> VadRecorder(user utterance, hybrid local+energy VAD by default)
-> CloudAsrSttProvider or SherpaOnnxSttProvider
-> TerminalRuntimeReporter.transcript()
-> ConversationContext
-> OpenAICompatibleLlmProvider
-> CloudTtsProvider or MacSayTtsProvider
-> speaker playback
Runtime Lifecycle
- Load config.
- Validate
OWNER_WAKE_PROVIDER=local_kws. - Load KWS model from
OWNER_SPEECH_MODELS_DIR. - Load VAD, STT, LLM, TTS.
- Open microphone stream.
- Wait for KWS wake event by feeding frames directly into wake provider.
- On wake hit, reset wake stream and VAD recorder.
- Record user utterance with VAD. The default provider is
hybrid: project-localsherpa-onnxVAD remains the primary detector, and an energy threshold fallback prevents low microphone gain from being treated as no speech. - Transcribe user utterance.
- Emit transcript to terminal.
- Append user text and call LLM.
- Synthesize/play reply.
- Append assistant reply and return to standby.
Interfaces
SherpaOnnxKeywordWakeWordProvider
class SherpaOnnxKeywordWakeWordProvider:
def __init__(self, models_dir, keyword, keywords_file=None, threshold=0.25, score=1.0, sherpa_module=None): ...
def load(self) -> None: ...
def detect(self, frame: AudioFrame) -> WakeEvent | None: ...
def reset(self) -> None: ...
RuntimeReporter
class RuntimeReporter(Protocol):
def status(self, state: str, message: str, *, turn_id: int | None = None) -> None: ...
def transcript(self, text: str, *, final: bool, turn_id: int | None = None) -> None: ...
def error(self, stage: str, code: str, message: str, *, turn_id: int | None = None) -> None: ...
Model Files
models/
manifest.json
wake/
sherpa-onnx-kws-zipformer-wenetspeech-3.3M-2024-01-01-mobile/
tokens.txt
encoder-epoch-12-avg-2-chunk-16-left-64.int8.onnx
decoder-epoch-12-avg-2-chunk-16-left-64.onnx
joiner-epoch-12-avg-2-chunk-16-left-64.int8.onnx
keywords.txt
vad/
silero_vad.onnx
stt/
sherpa-onnx-streaming-zipformer-zh-14M-2023-02-23/
Error Handling
- Missing KWS model:
WAKE_MODEL_MISSING. - KWS load failure:
WAKE_MODEL_LOAD_FAILED. - KWS runtime failure:
WAKE_MODEL_LOAD_FAILEDwith retryable true. - Empty user STT: existing
STT_EMPTY_TRANSCRIPT. - Invalid wake provider config:
CONFIG_MISSING_VALUE. - Real microphone VAD miss: default
OWNER_VAD_PROVIDER=hybridSHALL accept speech when either the local model or the energy fallback detects speech.
Testing Strategy
- Unit test fake wake provider detects wake without STT calls.
- Unit test repeated runtime runs two turns with exactly two STT calls.
- Unit test terminal reporter records transcript before LLM stage.
- Unit test LLM user content excludes wake keyword.
- Unit test KWS provider missing model raises structured error.
- Model-check test validates wake required files.
Migration
No database migration. Users should run:
python3.11 scripts/download_speech_models.py --dir models
.venv/bin/python -m owner_voice_pet model-check --models-dir models
Existing .env remains valid because new wake keys have defaults.