Files
Owner/openspec/changes/separate-wake-and-realtime-transcript/design.md
T

12 KiB
Raw Blame History

独立本地唤醒与终端转写显示设计

Context

真实运行日志显示,当前 run-live 会先截取语音段并使用 STT 判断是否包含“小杰小杰”。该实现把 wake detection 与 user utterance transcription 绑定在一起,导致唤醒慢、唤醒词污染正式问题、终端输出难以区分 STT 和 LLM 阶段。本设计把 wake word detection 升级为本地 KWS 模型路径,并把正式问题转写作为独立可见事件输出。

Goals

  1. 唤醒词“小杰小杰”由本地模型检测。
  2. wake 阶段不调用云 ASR,也不调用正式 STT provider。
  3. wake 命中后才开始正式问题录音和 STT。
  4. 终端在 LLM 前显示正式问题转写文本。
  5. 模型下载和检查覆盖 wake KWS、VAD、STT。
  6. 自动化测试证明 wake/STT 分离、重复对话、上下文和错误恢复仍正常。

Non-Goals

  1. 不实现 GUI 桌宠窗口。
  2. 不实现跨进程长期记忆。
  3. 不在本阶段实现逐字 partial ASR 字幕;本阶段先保证正式问题 STT 完成后立即可见,且出现在 LLM 前。
  4. 不把连续麦克风流上传到云端。

Architecture

run-live
  -> AppConfig(.env)
  -> VoiceAssistantPipeline
  -> PipelineEventBus
  -> TurnController
  -> SoundDeviceAudioTransport
  -> SherpaOnnxKeywordWakeWordProvider(local models/wake)
  -> AcknowledgeStage(local "我在")
  -> CaptureStage(raw mic frames)
  -> AudioPreprocessStage(local GTCRN denoise, capture only)
  -> PrimarySpeakerEndpoint/VAD(denoised frames)
  -> SherpaOnnxSttProvider(local CTC partial/final by default)
  -> DialogStage(session memory)
  -> ConversationContext
  -> OpenAICompatibleLlmProvider
  -> MacSayTtsProvider
  -> speaker playback

Runtime Lifecycle

  1. Load config.
  2. Validate OWNER_WAKE_PROVIDER=local_kws.
  3. Load KWS model from OWNER_SPEECH_MODELS_DIR.
  4. Load VAD, STT, LLM, TTS.
  5. Open microphone stream.
  6. Wait for KWS wake event by feeding frames directly into wake provider.
  7. On wake hit, reset wake stream and VAD recorder.
  8. Record user utterance with VAD. The default provider is hybrid: project-local sherpa-onnx VAD remains the primary detector, and an energy threshold fallback prevents low microphone gain from being treated as no speech.
  9. When OWNER_ENDPOINT_MODE=primary_speaker, build a temporary per-turn speaker profile from the first valid user speech frames and end the capture when the primary speaker is absent for OWNER_SPEAKER_ABSENT_MS.
  10. When OWNER_NOISE_FILTER_ENABLED=1, pass formal user utterance frames through the local GTCRN denoiser before VAD, realtime STT, and final STT segment assembly.
  11. Transcribe user utterance with the configured STT provider. The default is local CTC STT; cloud STT remains an explicit compatibility option.
  12. Emit transcript event to terminal.
  13. Append final user text and call LLM.
  14. Synthesize/play reply with local TTS by default.
  15. Append assistant reply and return to standby.

Interfaces

SherpaOnnxKeywordWakeWordProvider

class SherpaOnnxKeywordWakeWordProvider:
    def __init__(self, models_dir, keyword, keywords_file=None, threshold=0.25, score=1.0, sherpa_module=None): ...
    def load(self) -> None: ...
    def detect(self, frame: AudioFrame) -> WakeEvent | None: ...
    def reset(self) -> None: ...

RuntimeReporter

class RuntimeReporter(Protocol):
    def status(self, state: str, message: str, *, turn_id: int | None = None) -> None: ...
    def transcript(self, text: str, *, final: bool, turn_id: int | None = None) -> None: ...
    def error(self, stage: str, code: str, message: str, *, turn_id: int | None = None) -> None: ...

Pipeline events

class PipelineEventBus:
    def subscribe(self, listener): ...
    def emit(self, event_type, *, turn_id=None, state=None, message="", payload=None): ...

Required event types are pipeline_started, wake_listening, wake_detected, ack_started, capture_started, speech_started, speech_ended, stt_started, transcript_final, llm_started, tts_started, playback_finished, standby_resumed, and stage_error.

AudioPreprocessor

class AudioPreprocessor:
    def load(self) -> None: ...
    def reset(self) -> None: ...
    def process_frame(self, frame: AudioFrame) -> AudioFrame: ...
    def flush(self) -> list[AudioFrame]: ...

NoopAudioPreprocessor preserves tests and disabled configurations. SherpaOnnxDenoiserPreprocessor uses sherpa_onnx.OnlineSpeechDenoiser with an OfflineSpeechDenoiserGtcrnModelConfig pointing to models/denoise/gtcrn_simple.onnx. It converts int16 PCM to float32, calls run(samples, sample_rate), converts returned DenoisedAudio.samples back to int16 PCM, and adds diagnostic metadata without storing audio.

The first version applies preprocessing only in capture. Wake listening remains raw by default because wake KWS models can be sensitive to spectral changes introduced by denoising. OWNER_WAKE_DENOISE_ENABLED=1 reserves an explicit future switch for wake preprocessing.

Local CTC STT

SherpaOnnxSttProvider reads providers.stt.type from models/manifest.json:

  1. sherpa-onnx-streaming-transducer: existing tokens/encoder/decoder/joiner files and OnlineRecognizer.from_transducer.
  2. sherpa-onnx-streaming-zipformer2-ctc: new default tokens/model files and OnlineRecognizer.from_zipformer2_ctc.

Realtime partial and final transcript use the same local recognizer when OWNER_SPEECH_PROVIDER=local. Partial output is filtered before emitting transcript_partial: texts with fewer than two meaningful characters, duplicate texts, and short regressions from the last displayed text are ignored. The final transcript remains the only user text appended to ConversationContext.

TurnController

class TurnController:
    def run_turn(self, turn_id: int) -> TurnResult: ...

The controller owns one turn state machine and delegates work to provider-backed stages. It does not write terminal text directly; it emits events only.

Primary speaker endpoint

The first implementation is a per-turn heuristic endpoint, not persistent voiceprint recognition. It extracts local PCM features from speech frames and compares future frames against the profile. If profile creation fails because the user speech is too short or too quiet, capture falls back to existing VAD silence endpoint.

Low Latency Capture Revision

真人运行反馈显示,当前 capture 仍可能把用户第一句话开头吞掉或需要用户重复提问才能结束。本修正把低延迟 capture 作为 pipeline 内部约束:

  1. OWNER_POST_PLAYBACK_DRAIN_MS 默认改为 0。ACK 播放结束后仅清掉播放期间积压在输入队列中的帧,不再主动等待并丢弃后续音频。
  2. SoundDeviceAudioTransport.read_frames() 在拿到首帧后立即返回队列中所有可用帧,避免真实麦克风回调积压时 pipeline 逐帧追赶。
  3. OWNER_SPEAKER_PROFILE_MIN_MS 控制临时主说话人画像最低就绪语音长度,默认 120 ms,不再复用 OWNER_VAD_MIN_DURATION_MS
  4. 主说话人画像就绪后,OWNER_SPEAKER_ABSENT_MS 是结束正式问题采集的主条件;主说话人连续缺席达到该值后直接进入 STT,不再额外等待普通 VAD 最小时长。
  5. 普通 VAD 静音仍作为画像不足或音色特征不可用时的兜底,最大录音时长仍作为最终保护。

Realtime Transcript Revision

录音期间新增 partial transcript 通道,用于解决“说话时看不到文字结果”的体验问题:

  1. Pipeline 事件新增 transcript_partial。终端 reporter 将其显示为 实时转写:<文本>,最终结果仍显示为 转写结果:<文本>
  2. partial transcript 只用于用户反馈,不写入 ConversationContext,不触发 LLMLLM 仍只消费 final STT 结果。
  3. 当 final ASR/TTS 走 cloud 时,partial transcript 使用本地 sherpa-onnx streaming STT,避免对云端 ASR 高频请求。
  4. VoiceAssistantPipelinespeech_started 后把 capture 阶段已开始录音的帧 feed 给 realtime STT session;当 partial 文本变化时才 emit,避免刷屏。
  5. OWNER_REALTIME_TRANSCRIPT_ENABLED=1 默认启用;设置为 0 可临时回退到只显示 final transcript。

Model Files

models/
  manifest.json
  wake/
    sherpa-onnx-kws-zipformer-wenetspeech-3.3M-2024-01-01-mobile/
      tokens.txt
      encoder-epoch-12-avg-2-chunk-16-left-64.int8.onnx
      decoder-epoch-12-avg-2-chunk-16-left-64.onnx
      joiner-epoch-12-avg-2-chunk-16-left-64.int8.onnx
    keywords.txt
  vad/
    silero_vad.onnx
  stt/
    sherpa-onnx-streaming-zipformer-ctc-zh-int8-2025-06-30/
      tokens.txt
      model.int8.onnx
  denoise/
    gtcrn_simple.onnx

Error Handling

  1. Missing KWS model: WAKE_MODEL_MISSING.
  2. KWS load failure: WAKE_MODEL_LOAD_FAILED.
  3. KWS runtime failure: WAKE_MODEL_LOAD_FAILED with retryable true.
  4. Empty user STT: existing STT_EMPTY_TRANSCRIPT.
  5. Invalid wake provider config: CONFIG_MISSING_VALUE.
  6. Real microphone VAD miss: default OWNER_VAD_PROVIDER=hybrid SHALL accept speech when either the local model or the energy fallback detects speech.
  7. Primary speaker endpoint profile failure: fallback to VAD silence endpoint.
  8. Pipeline stage failure: emit stage_error, recover to standby, and keep the process alive unless startup dependencies are missing.
  9. Playback drain misconfiguration: negative OWNER_POST_PLAYBACK_DRAIN_MS remains invalid; non-zero values are treated as explicit user tuning rather than default behavior.
  10. Speaker profile threshold misconfiguration: non-positive OWNER_SPEAKER_PROFILE_MIN_MS fails config validation.
  11. Realtime STT failure: startup model缺失按 model-check 暴露;capture 中 partial session 失败不得把 partial 文本写入上下文。
  12. Denoiser missing model: model-check reports denoise/gtcrn_simple.onnx and live startup fails before microphone listening.
  13. Denoiser runtime failure: the current turn emits stage_error and recovers to standby without invoking final STT or LLM.
  14. CTC manifest mismatch: missing model or tokens files report STT_MODEL_MISSING; transducer manifests remain supported for existing local setups.
  15. Partial text noise: single-character or transient partial output is ignored rather than displayed as terminal feedback.

Testing Strategy

  1. Unit test fake wake provider detects wake without STT calls.
  2. Unit test repeated runtime runs two turns with exactly two STT calls.
  3. Unit test terminal reporter records transcript before LLM stage.
  4. Unit test LLM user content excludes wake keyword.
  5. Unit test KWS provider missing model raises structured error.
  6. Model-check test validates wake required files.
  7. Pipeline event order test validates successful two-turn event sequence.
  8. Primary speaker endpoint test validates that background noise or a later repeated utterance does not extend the current turn after the main speaker disappears.
  9. First utterance preservation test validates that frames immediately after ACK are not discarded by post-playback drain.
  10. Low-latency endpoint test validates that a short first question ends by primary speaker absence without waiting for a repeated second question.
  11. Transport batching test validates that queued SoundDevice frames are returned together.
  12. Partial transcript event test validates that realtime text appears after speech start and before final transcript.
  13. Context isolation test validates partial transcript does not enter LLM messages.
  14. Denoiser manifest/model-check test validates the required GTCRN file is present.
  15. Fake denoiser test validates VAD, partial STT, and final STT receive denoised frames from the same processed stream.
  16. Local speech provider test validates OWNER_SPEECH_PROVIDER=local does not construct cloud ASR/TTS providers.
  17. Partial filter test validates single-character and short transient results are not emitted.
  18. CTC STT loading test validates manifest type selects from_zipformer2_ctc and does not require transducer encoder/decoder/joiner files.

Migration

No database migration. Users should run:

python3.11 scripts/download_speech_models.py --dir models
.venv/bin/python -m owner_voice_pet model-check --models-dir models

Existing .env remains valid because new wake keys have defaults. For the local voice-chain revision, users should also ensure:

OWNER_SPEECH_PROVIDER=local
OWNER_NOISE_FILTER_ENABLED=1
OWNER_WAKE_DENOISE_ENABLED=0
OWNER_NOISE_FILTER_PROVIDER=sherpa_onnx_gtcrn