[真实流程计时]:完成完整链路耗时输出,包含命令时间、LLM发送时间和阶段响应耗时

This commit is contained in:
mkbk
2026-06-18 00:45:10 +08:00
parent b9b501cf74
commit 39da9dd09f
7 changed files with 351 additions and 61 deletions
+2
View File
@@ -122,6 +122,8 @@ python3.11 scripts/download_speech_models.py --dir models
这个命令不用真人对着麦克风说话,会用 macOS `say/afconvert` 生成“小杰小杰”和两轮问题音频,再驱动真实 `VoiceAssistantPipeline`:本地 KWS、本地 VAD、本地降噪、本地实时字幕、本地 final STT、云端 LLM、本地 TTS、Transport 播放、恢复待机都会跑到。默认会真实播放 ACK 和回复;如果只想检查链路但不发声,可以加 `--no-playback`。输出 JSON 中 `success=true` 表示完整链路通过,并会检查第二轮 LLM 请求是否携带第一轮临时历史。 这个命令不用真人对着麦克风说话,会用 macOS `say/afconvert` 生成“小杰小杰”和两轮问题音频,再驱动真实 `VoiceAssistantPipeline`:本地 KWS、本地 VAD、本地降噪、本地实时字幕、本地 final STT、云端 LLM、本地 TTS、Transport 播放、恢复待机都会跑到。默认会真实播放 ACK 和回复;如果只想检查链路但不发声,可以加 `--no-playback`。输出 JSON 中 `success=true` 表示完整链路通过,并会检查第二轮 LLM 请求是否携带第一轮临时历史。
输出里的 `timing` 会给出命令 `started_at``finished_at`、总 `duration_ms`、各 phase 耗时,以及按 turn/stage 聚合的 `stage_timings`,例如 wake 等待、ACK、等待说话、录音、final STT、LLM 到首个 TTS、TTS 播放、整轮总耗时。`llm_request_timings` 会记录每次云端 LLM 请求真正发出去的 `sent_at`、首个回复文本耗时 `first_delta_ms` 和请求总耗时 `duration_ms`
单轮验收: 单轮验收:
```bash ```bash
@@ -266,6 +266,23 @@ Playback drain semantics are intentionally split by transport type:
`_drain_input_after_playback()` therefore always calls `flush_input()` once after actual playback. If `OWNER_POST_PLAYBACK_DRAIN_MS=0`, it returns immediately and does not perform any timed read/drop loop. If the value is positive, the positive window is treated as an explicit user tuning and the runtime reads and discards only that configured duration before a final flush. When `OWNER_WAKE_ACK_TEXT` is empty and no ACK playback occurs, the runtime skips playback drain entirely. `_drain_input_after_playback()` therefore always calls `flush_input()` once after actual playback. If `OWNER_POST_PLAYBACK_DRAIN_MS=0`, it returns immediately and does not perform any timed read/drop loop. If the value is positive, the positive window is treated as an explicit user tuning and the runtime reads and discards only that configured duration before a final flush. When `OWNER_WAKE_ACK_TEXT` is empty and no ACK playback occurs, the runtime skips playback drain entirely.
## Real Provider Timing Observability
`owner_voice_pet real-live-check` is also the repeatable latency diagnostic entry for the full chain. The command records timing at three levels:
1. Command lifecycle: `started_at`, `finished_at`, epoch millisecond timestamps, and total `duration_ms`.
2. Preparation phases: `prepare_config`, `generate_fixture_audio`, `build_pipeline`, and `pipeline_run`, each with start offset and duration.
3. Pipeline events: every emitted stage event is recorded with turn id, ISO timestamp, command-relative offset, and message. Derived `stage_timings` pair stable stage boundaries such as `wake_listening -> wake_detected`, `capture_started -> speech_started`, `speech_started -> speech_ended`, `stt_started -> transcript_final`, `llm_started -> tts_started`, `tts_started -> playback_finished`, and `wake_listening -> standby_resumed`.
The real LLM provider is wrapped by `RecordingLlmProvider`. For every cloud LLM call it records:
1. `sent_at` and `sent_offset_ms`, captured immediately before delegating to the OpenAI-compatible provider.
2. `first_delta_at` and `first_delta_ms`, captured when the first non-empty text delta is yielded.
3. `finished_at` and `duration_ms`, captured when the stream ends or unwinds through an error.
4. `message_count` and `last_user_preview`, enough to correlate the request with a turn without serializing credentials or full request payloads.
Timing output is diagnostic metadata only. It SHALL NOT include `OWNER_LLM_API_KEY`, raw audio, denoised audio, full provider request headers, or any other secret-bearing configuration.
## Migration ## Migration
No database migration. Users should run: No database migration. Users should run:
@@ -105,6 +105,8 @@
40. 系统 SHALL 提供 `owner_voice_pet real-live-check` 命令,把真实 Provider 完整链路自测固化为可重复入口,而不是依赖一次性内联脚本。 40. 系统 SHALL 提供 `owner_voice_pet real-live-check` 命令,把真实 Provider 完整链路自测固化为可重复入口,而不是依赖一次性内联脚本。
41. `real-live-check` SHALL 默认生成两轮音频并执行真实本地 KWS/VAD/降噪/STT、本地 TTS、Transport 播放和云端 LLM;输出 JSON 必须包含完成轮数、final transcripts、事件计数、播放段数、flush 次数、临时上下文检查和错误列表。 41. `real-live-check` SHALL 默认生成两轮音频并执行真实本地 KWS/VAD/降噪/STT、本地 TTS、Transport 播放和云端 LLM;输出 JSON 必须包含完成轮数、final transcripts、事件计数、播放段数、flush 次数、临时上下文检查和错误列表。
42. `real-live-check --no-playback` SHALL 仍执行本地 TTS 合成和 pipeline 播放事件记录,但不实际向扬声器发声,便于自动化测试和无声环境排查。 42. `real-live-check --no-playback` SHALL 仍执行本地 TTS 合成和 pipeline 播放事件记录,但不实际向扬声器发声,便于自动化测试和无声环境排查。
43. `real-live-check` SHALL 输出命令级开始/结束时间、总耗时、阶段耗时、pipeline 事件时间线和按 turn/stage 聚合的响应耗时。
44. `real-live-check` SHALL 输出每次云端 LLM 请求真正发出的 `sent_at`、首个文本回复耗时和请求总耗时,且不得泄露 API key。
### 非功能需求 ### 非功能需求
@@ -442,6 +444,13 @@ standby
- [ ] 13.3 执行真实命令验收和全量门禁;前置条件:13.2 完成;验收标准:`real-live-check --turns 2`、compileall、unittest、simulate-live、security-check、model-check、device-check、OpenSpec strict、git diff check 通过;优先级:P0;预计:60 分钟。 - [ ] 13.3 执行真实命令验收和全量门禁;前置条件:13.2 完成;验收标准:`real-live-check --turns 2`、compileall、unittest、simulate-live、security-check、model-check、device-check、OpenSpec strict、git diff check 通过;优先级:P0;预计:60 分钟。
- [ ] 13.4 提交“真实流程命令”模块;前置条件:13.3 通过;验收标准:中文 commit 信息为 `[真实流程命令]:完成完整链路自测入口,包含真实Provider CLI、播放验收和文档测试`,提交后 `git status --short` 为空;优先级:P0;预计:10 分钟。 - [ ] 13.4 提交“真实流程命令”模块;前置条件:13.3 通过;验收标准:中文 commit 信息为 `[真实流程命令]:完成完整链路自测入口,包含真实Provider CLI、播放验收和文档测试`,提交后 `git status --short` 为空;优先级:P0;预计:10 分钟。
### 14. 真实流程命令计时输出
- [ ] 14.1 更新 OpenSpec 和 README 描述真实流程计时字段;前置条件:`real-live-check` 已可运行;验收标准:proposal/design/spec/tasks/README 明确 `started_at``finished_at``duration_ms``stage_timings``llm_request_timings.sent_at`;测试要点:OpenSpec strict;优先级:P0;预计:30 分钟。
- [ ] 14.2 实现命令级和事件级计时;前置条件:14.1 完成;验收标准:`real-live-check` JSON 包含命令总耗时、生成音频/build/pipeline phase 耗时、每轮 wake/ACK/capture/STT/LLM/TTS/turn_total 耗时;测试要点:stage timing 单测;优先级:P0;预计:45 分钟。
- [ ] 14.3 实现 LLM 请求发送时间和响应耗时;前置条件:14.2 完成;验收标准:每次 LLM 调用记录 `sent_at``first_delta_ms``duration_ms`、最后 user 预览,且不包含 key;测试要点:RecordingLlmProvider 单测;优先级:P0;预计:35 分钟。
- [ ] 14.4 执行真实命令验收和全量门禁并提交;前置条件:14.1 至 14.3 完成;验收标准:`real-live-check --turns 2 --no-playback` 和默认播放版本均输出 timing 且成功,全量门禁通过后中文提交;优先级:P0;预计:60 分钟。
## Spec Deltas ## Spec Deltas
### 新增能力 ### 新增能力
@@ -464,6 +473,7 @@ standby
12. `Simulated microphone live acceptance`:新增不依赖真人麦克风的 live pipeline 模拟验收,覆盖两轮重复对话、短噪声过滤、背景噪声端点、fixture 写入/回放和空帧防卡死。 12. `Simulated microphone live acceptance`:新增不依赖真人麦克风的 live pipeline 模拟验收,覆盖两轮重复对话、短噪声过滤、背景噪声端点、fixture 写入/回放和空帧防卡死。
13. `Real provider live fixture acceptance`:新增真实 provider 脚本化闭环验收和 ACK/TTS 播放后 0ms flush 语义,确保真实本地模型、云端 LLM、本地 TTS/播放在同一 pipeline 中可连续运行。 13. `Real provider live fixture acceptance`:新增真实 provider 脚本化闭环验收和 ACK/TTS 播放后 0ms flush 语义,确保真实本地模型、云端 LLM、本地 TTS/播放在同一 pipeline 中可连续运行。
14. `Real live check CLI`:新增 `real-live-check` 可重复命令,把完整真实 Provider 链路验收从一次性脚本升级为仓库内正式工具。 14. `Real live check CLI`:新增 `real-live-check` 可重复命令,把完整真实 Provider 链路验收从一次性脚本升级为仓库内正式工具。
15. `Real live check timing`:新增真实流程命令计时输出,覆盖命令发起时间、LLM 请求发送时间、各阶段响应耗时和 pipeline 事件时间线。
### 删除项 ### 删除项
@@ -487,6 +497,7 @@ standby
10. M10:模拟麦克风自动验收完成后提交,作为真人验收前的可重复自测入口;仍不归档变更,直到真实 `run-live` 行为由用户确认。 10. M10:模拟麦克风自动验收完成后提交,作为真人验收前的可重复自测入口;仍不归档变更,直到真实 `run-live` 行为由用户确认。
11. M11:真实 provider 脚本化闭环和 ACK 缓冲修正完成后提交;仍保留用户真人 `run-live` 体验验收,不在用户确认前归档。 11. M11:真实 provider 脚本化闭环和 ACK 缓冲修正完成后提交;仍保留用户真人 `run-live` 体验验收,不在用户确认前归档。
12. M12`real-live-check` 正式 CLI、文档、测试和真实命令验收完成后提交;仍保留物理麦克风 `run-live` 真人体验验收。 12. M12`real-live-check` 正式 CLI、文档、测试和真实命令验收完成后提交;仍保留物理麦克风 `run-live` 真人体验验收。
13. M13`real-live-check` 计时输出完成后提交,使后续排查可以直接看到每段耗时。
估时: 估时:
@@ -161,6 +161,13 @@ The system SHALL support scripted full-chain acceptance that drives the live pip
- **THEN** the command SHALL still synthesize local TTS output and emit pipeline playback events - **THEN** the command SHALL still synthesize local TTS output and emit pipeline playback events
- **AND** it SHALL skip actual speaker output for automated or quiet environments - **AND** it SHALL skip actual speaker output for automated or quiet environments
#### Scenario: Real provider check reports timing diagnostics
- **WHEN** the developer runs `.venv/bin/python -m owner_voice_pet real-live-check --turns 2`
- **THEN** the JSON output SHALL include `timing.started_at`, `timing.finished_at`, `timing.duration_ms`, `timing.phases`, `timing.events`, and `timing.stage_timings`
- **AND** `timing.stage_timings` SHALL include turn-scoped durations for wake wait, acknowledgement, wait-for-speech, capture, final STT, LLM-to-first-TTS, TTS playback, and full turn when the corresponding events are emitted
- **AND** the JSON output SHALL include `llm_request_timings` with one item per LLM call containing `sent_at`, `sent_offset_ms`, `first_delta_ms`, `finished_at`, `duration_ms`, `message_count`, and a short last-user preview
- **AND** the timing output SHALL NOT contain API keys, raw audio payloads, denoised audio payloads, or full secret-bearing provider request headers
### Requirement: Wake acknowledgement before recording ### Requirement: Wake acknowledgement before recording
The live runtime SHALL provide an audible local acknowledgement after local wake detection and before it starts recording the user's formal question. The live runtime SHALL provide an audible local acknowledgement after local wake detection and before it starts recording the user's formal question.
@@ -104,3 +104,10 @@
- [x] 13.2 实现 `owner_voice_pet real-live-check`;前置条件:13.1 完成;验收标准:命令从 `.env` 读取配置,生成 wake/question PCM,运行真实 KWS/VAD/denoise/STT/LLM/TTS/pipeline/playback,并输出结构化 JSON;测试要点:CLI wiring 测试、transport 单测;优先级:P0;预计:60 分钟。 - [x] 13.2 实现 `owner_voice_pet real-live-check`;前置条件:13.1 完成;验收标准:命令从 `.env` 读取配置,生成 wake/question PCM,运行真实 KWS/VAD/denoise/STT/LLM/TTS/pipeline/playback,并输出结构化 JSON;测试要点:CLI wiring 测试、transport 单测;优先级:P0;预计:60 分钟。
- [x] 13.3 执行真实命令验收和全量门禁;前置条件:13.2 完成;验收标准:`real-live-check --turns 2`、compileall、unittest、simulate-live、security-check、model-check、device-check、OpenSpec strict、git diff check 通过;优先级:P0;预计:60 分钟。 - [x] 13.3 执行真实命令验收和全量门禁;前置条件:13.2 完成;验收标准:`real-live-check --turns 2`、compileall、unittest、simulate-live、security-check、model-check、device-check、OpenSpec strict、git diff check 通过;优先级:P0;预计:60 分钟。
- [x] 13.4 提交“真实流程命令”模块;前置条件:13.3 通过;验收标准:中文 commit 信息为 `[真实流程命令]:完成完整链路自测入口,包含真实Provider CLI、播放验收和文档测试`,提交后 `git status --short` 为空;优先级:P0;预计:10 分钟。 - [x] 13.4 提交“真实流程命令”模块;前置条件:13.3 通过;验收标准:中文 commit 信息为 `[真实流程命令]:完成完整链路自测入口,包含真实Provider CLI、播放验收和文档测试`,提交后 `git status --short` 为空;优先级:P0;预计:10 分钟。
## 14. 真实流程命令计时输出
- [x] 14.1 更新 OpenSpec 和 README 描述真实流程计时字段;前置条件:`real-live-check` 已可运行;验收标准:proposal/design/spec/tasks/README 明确 `started_at``finished_at``duration_ms``stage_timings``llm_request_timings.sent_at`;测试要点:OpenSpec strict;优先级:P0;预计:30 分钟。
- [x] 14.2 实现命令级和事件级计时;前置条件:14.1 完成;验收标准:`real-live-check` JSON 包含命令总耗时、生成音频/build/pipeline phase 耗时、每轮 wake/ACK/capture/STT/LLM/TTS/turn_total 耗时;测试要点:stage timing 单测;优先级:P0;预计:45 分钟。
- [x] 14.3 实现 LLM 请求发送时间和响应耗时;前置条件:14.2 完成;验收标准:每次 LLM 调用记录 `sent_at``first_delta_ms``duration_ms`、最后 user 预览,且不包含 key;测试要点:RecordingLlmProvider 单测;优先级:P0;预计:35 分钟。
- [x] 14.4 执行真实命令验收和全量门禁并提交;前置条件:14.1 至 14.3 完成;验收标准:`real-live-check --turns 2 --no-playback` 和默认播放版本均输出 timing 且成功,全量门禁通过后中文提交;优先级:P0;预计:60 分钟。
+197 -4
View File
@@ -1,5 +1,8 @@
from __future__ import annotations from __future__ import annotations
import time
from contextlib import contextmanager
from datetime import datetime
import subprocess import subprocess
import tempfile import tempfile
import wave import wave
@@ -29,6 +32,85 @@ DEFAULT_QUESTIONS = [
"我叫阿明,请你记住我的名字。", "我叫阿明,请你记住我的名字。",
"我叫什么名字?", "我叫什么名字?",
] ]
STAGE_PAIRS = [
("wake_wait", "wake_listening", "wake_detected"),
("acknowledgement", "ack_started", "question_prompt"),
("wait_for_speech", "capture_started", "speech_started"),
("capture", "speech_started", "speech_ended"),
("final_stt", "stt_started", "transcript_final"),
("llm_to_first_tts", "llm_started", "tts_started"),
("turn_total", "wake_listening", "standby_resumed"),
]
def _now_iso() -> str:
return datetime.now().astimezone().isoformat(timespec="milliseconds")
def _epoch_ms() -> int:
return int(time.time() * 1000)
def _elapsed_ms(start: float, end: float | None = None) -> int:
return int(round(((end if end is not None else time.perf_counter()) - start) * 1000))
class TimingRecorder:
def __init__(self) -> None:
self.started_at = _now_iso()
self.started_epoch_ms = _epoch_ms()
self._started_mono = time.perf_counter()
self.finished_at: str | None = None
self.finished_epoch_ms: int | None = None
self.duration_ms: int | None = None
self.phases: list[dict[str, object]] = []
self.events: list[dict[str, object]] = []
@contextmanager
def phase(self, name: str) -> Iterable[None]:
started_mono = time.perf_counter()
record: dict[str, object] = {
"name": name,
"started_at": _now_iso(),
"start_offset_ms": _elapsed_ms(self._started_mono, started_mono),
}
try:
yield
finally:
finished_mono = time.perf_counter()
record["finished_at"] = _now_iso()
record["duration_ms"] = _elapsed_ms(started_mono, finished_mono)
self.phases.append(record)
def record_event(self, event_type: str, *, turn_id: int | None, message: str) -> None:
self.events.append(
{
"type": event_type,
"turn": turn_id,
"at": _now_iso(),
"offset_ms": _elapsed_ms(self._started_mono),
"message": message,
}
)
def finish(self) -> None:
self.finished_at = _now_iso()
self.finished_epoch_ms = _epoch_ms()
self.duration_ms = _elapsed_ms(self._started_mono)
def to_json(self) -> dict[str, object]:
if self.duration_ms is None:
self.finish()
return {
"started_at": self.started_at,
"started_epoch_ms": self.started_epoch_ms,
"finished_at": self.finished_at,
"finished_epoch_ms": self.finished_epoch_ms,
"duration_ms": self.duration_ms,
"phases": self.phases,
"events": self.events,
"stage_timings": _stage_timings_from_events(self.events),
}
@dataclass(slots=True) @dataclass(slots=True)
@@ -50,13 +132,37 @@ class RealLiveCheckReporter:
class RecordingLlmProvider: class RecordingLlmProvider:
def __init__(self, delegate: OpenAICompatibleLlmProvider) -> None: def __init__(self, delegate: OpenAICompatibleLlmProvider, timing: TimingRecorder) -> None:
self.delegate = delegate self.delegate = delegate
self.timing = timing
self.calls: list[list[Message]] = [] self.calls: list[list[Message]] = []
self.request_timings: list[dict[str, object]] = []
def stream_reply(self, messages: Sequence[Message]) -> Iterable[ReplyDelta]: def stream_reply(self, messages: Sequence[Message]) -> Iterable[ReplyDelta]:
self.calls.append(list(messages)) self.calls.append(list(messages))
yield from self.delegate.stream_reply(messages) sent_mono = time.perf_counter()
request_timing: dict[str, object] = {
"call_index": len(self.calls),
"sent_at": _now_iso(),
"sent_offset_ms": _elapsed_ms(self.timing._started_mono, sent_mono),
"message_count": len(messages),
"last_user_preview": _last_user_preview(messages),
}
first_delta_mono: float | None = None
try:
for delta in self.delegate.stream_reply(messages):
if first_delta_mono is None and delta.text_delta:
first_delta_mono = time.perf_counter()
request_timing["first_delta_at"] = _now_iso()
request_timing["first_delta_ms"] = _elapsed_ms(sent_mono, first_delta_mono)
yield delta
finally:
finished_mono = time.perf_counter()
request_timing["finished_at"] = _now_iso()
request_timing["duration_ms"] = _elapsed_ms(sent_mono, finished_mono)
if first_delta_mono is None:
request_timing["first_delta_ms"] = None
self.request_timings.append(request_timing)
class FixtureLiveAudioTransport: class FixtureLiveAudioTransport:
@@ -113,8 +219,10 @@ def run_real_live_check(
questions: list[str] | None = None, questions: list[str] | None = None,
play_audio: bool = True, play_audio: bool = True,
) -> dict[str, object]: ) -> dict[str, object]:
timing = TimingRecorder()
if turns <= 0: if turns <= 0:
raise ValueError("turns must be positive") raise ValueError("turns must be positive")
with timing.phase("prepare_config"):
selected_questions = _question_list(turns, questions) selected_questions = _question_list(turns, questions)
config = replace( config = replace(
config, config,
@@ -125,12 +233,15 @@ def run_real_live_check(
post_playback_drain_ms=0, post_playback_drain_ms=0,
) )
config.require_llm_credentials() config.require_llm_credentials()
with timing.phase("generate_fixture_audio"):
frames = _generated_live_frames(config, turns=turns, voice=voice, wake_text=wake_text, questions=selected_questions) frames = _generated_live_frames(config, turns=turns, voice=voice, wake_text=wake_text, questions=selected_questions)
with timing.phase("build_pipeline"):
reporter = RealLiveCheckReporter() reporter = RealLiveCheckReporter()
event_bus = PipelineEventBus() event_bus = PipelineEventBus()
event_bus.subscribe(lambda event: timing.record_event(event.type, turn_id=event.turn_id, message=event.message))
stt = SherpaOnnxSttProvider(str(config.speech_models_dir)) stt = SherpaOnnxSttProvider(str(config.speech_models_dir))
tts = MacSayTtsProvider() tts = MacSayTtsProvider()
llm = RecordingLlmProvider(OpenAICompatibleLlmProvider(config, timeout_s=60)) llm = RecordingLlmProvider(OpenAICompatibleLlmProvider(config, timeout_s=60), timing)
transport = FixtureLiveAudioTransport( transport = FixtureLiveAudioTransport(
frames, frames,
play_audio=play_audio, play_audio=play_audio,
@@ -171,8 +282,11 @@ def run_real_live_check(
reporter=reporter, reporter=reporter,
event_bus=event_bus, event_bus=event_bus,
) )
with timing.phase("pipeline_run"):
summary = pipeline.run(max_turns=turns) summary = pipeline.run(max_turns=turns)
timing.finish()
event_counts = Counter(event.type for event in event_bus.events) event_counts = Counter(event.type for event in event_bus.events)
timing_data = timing.to_json()
context_messages = [{"role": item.role, "content": item.content[:120]} for item in pipeline.context.messages()] context_messages = [{"role": item.role, "content": item.content[:120]} for item in pipeline.context.messages()]
checks = { checks = {
"completed_turns": summary.completed_turns == turns, "completed_turns": summary.completed_turns == turns,
@@ -184,8 +298,11 @@ def run_real_live_check(
"tts_per_turn": event_counts["tts_started"] >= turns, "tts_per_turn": event_counts["tts_started"] >= turns,
"playback_per_turn": event_counts["playback_finished"] >= turns, "playback_per_turn": event_counts["playback_finished"] >= turns,
"standby_per_turn": event_counts["standby_resumed"] >= turns, "standby_per_turn": event_counts["standby_resumed"] >= turns,
"temporary_context_in_second_llm": _second_call_contains_first_turn(llm.calls), "temporary_context_in_second_llm": turns < 2 or _second_call_contains_first_turn(llm.calls),
"real_output_playback": (not play_audio) or len(transport.played_segments) >= turns * 2, "real_output_playback": (not play_audio) or len(transport.played_segments) >= turns * 2,
"timing_present": bool(timing_data.get("duration_ms") and timing_data.get("stage_timings")),
"llm_request_timing_present": len(llm.request_timings) >= turns
and all("sent_at" in item and "duration_ms" in item for item in llm.request_timings),
} }
return { return {
"success": all(checks.values()) and not reporter.errors, "success": all(checks.values()) and not reporter.errors,
@@ -202,6 +319,8 @@ def run_real_live_check(
"errors": reporter.errors, "errors": reporter.errors,
"play_audio": play_audio, "play_audio": play_audio,
"voice": voice, "voice": voice,
"timing": timing_data,
"llm_request_timings": llm.request_timings,
} }
@@ -337,3 +456,77 @@ def _second_call_contains_first_turn(calls: list[list[Message]]) -> bool:
return False return False
second_contents = "\n".join(message.content for message in calls[1]) second_contents = "\n".join(message.content for message in calls[1])
return "我叫阿明" in second_contents and any(message.role == "assistant" for message in calls[1]) return "我叫阿明" in second_contents and any(message.role == "assistant" for message in calls[1])
def _last_user_preview(messages: Sequence[Message]) -> str:
for message in reversed(messages):
if message.role == "user":
return message.content[:80]
return ""
def _stage_timings_from_events(events: list[dict[str, object]]) -> list[dict[str, object]]:
timings: list[dict[str, object]] = []
turns = sorted({int(event["turn"]) for event in events if event.get("turn") is not None})
for turn_id in turns:
turn_events = [event for event in events if event.get("turn") == turn_id]
for stage_name, start_type, end_type in STAGE_PAIRS:
timing = _pair_timing(turn_events, stage_name, start_type, end_type)
if timing is not None:
timings.append(timing)
timings.extend(_repeated_pair_timings(turn_events, "tts_playback", "tts_started", "playback_finished"))
return timings
def _pair_timing(
events: list[dict[str, object]],
name: str,
start_type: str,
end_type: str,
) -> dict[str, object] | None:
start = next((event for event in events if event.get("type") == start_type), None)
if start is None:
return None
end = next(
(event for event in events if event.get("type") == end_type and int(event["offset_ms"]) >= int(start["offset_ms"])),
None,
)
if end is None:
return None
return _timing_record(name, start, end)
def _repeated_pair_timings(
events: list[dict[str, object]],
name: str,
start_type: str,
end_type: str,
) -> list[dict[str, object]]:
timings: list[dict[str, object]] = []
pending: dict[str, object] | None = None
segment_index = 1
for event in events:
if event.get("type") == start_type:
pending = event
continue
if pending is not None and event.get("type") == end_type:
timing = _timing_record(name, pending, event)
timing["segment_index"] = segment_index
timings.append(timing)
segment_index += 1
pending = None
return timings
def _timing_record(name: str, start: dict[str, object], end: dict[str, object]) -> dict[str, object]:
return {
"turn": start.get("turn"),
"name": name,
"started_at": start.get("at"),
"finished_at": end.get("at"),
"start_offset_ms": start.get("offset_ms"),
"end_offset_ms": end.get("offset_ms"),
"duration_ms": int(end["offset_ms"]) - int(start["offset_ms"]),
"start_event": start.get("type"),
"end_event": end.get("type"),
}
+55 -2
View File
@@ -2,8 +2,13 @@ from __future__ import annotations
import unittest import unittest
from owner_voice_pet.models import AudioFrame, AudioSegment, PlaybackResult, TransportHealth from owner_voice_pet.models import AudioFrame, AudioSegment, Message, PlaybackResult, ReplyDelta, TransportHealth
from owner_voice_pet.real_live_check import FixtureLiveAudioTransport from owner_voice_pet.real_live_check import (
FixtureLiveAudioTransport,
RecordingLlmProvider,
TimingRecorder,
_stage_timings_from_events,
)
def frame(idx: int) -> AudioFrame: def frame(idx: int) -> AudioFrame:
@@ -42,6 +47,54 @@ class RealLiveCheckTests(unittest.TestCase):
self.assertEqual(output.played, [segment]) self.assertEqual(output.played, [segment])
self.assertEqual(transport.played_segments, [segment]) self.assertEqual(transport.played_segments, [segment])
def test_stage_timings_are_built_from_pipeline_events(self) -> None:
events = [
{"type": "wake_listening", "turn": 1, "at": "t0", "offset_ms": 0, "message": ""},
{"type": "wake_detected", "turn": 1, "at": "t1", "offset_ms": 120, "message": ""},
{"type": "ack_started", "turn": 1, "at": "t2", "offset_ms": 130, "message": ""},
{"type": "question_prompt", "turn": 1, "at": "t3", "offset_ms": 250, "message": ""},
{"type": "capture_started", "turn": 1, "at": "t4", "offset_ms": 260, "message": ""},
{"type": "speech_started", "turn": 1, "at": "t5", "offset_ms": 300, "message": ""},
{"type": "speech_ended", "turn": 1, "at": "t6", "offset_ms": 700, "message": ""},
{"type": "stt_started", "turn": 1, "at": "t7", "offset_ms": 720, "message": ""},
{"type": "transcript_final", "turn": 1, "at": "t8", "offset_ms": 900, "message": ""},
{"type": "llm_started", "turn": 1, "at": "t9", "offset_ms": 910, "message": ""},
{"type": "tts_started", "turn": 1, "at": "t10", "offset_ms": 1300, "message": ""},
{"type": "playback_finished", "turn": 1, "at": "t11", "offset_ms": 1500, "message": ""},
{"type": "standby_resumed", "turn": 1, "at": "t12", "offset_ms": 1520, "message": ""},
]
timings = _stage_timings_from_events(events)
by_name = {str(item["name"]): item for item in timings if "segment_index" not in item}
self.assertEqual(by_name["wake_wait"]["duration_ms"], 120)
self.assertEqual(by_name["acknowledgement"]["duration_ms"], 120)
self.assertEqual(by_name["capture"]["duration_ms"], 400)
self.assertEqual(by_name["final_stt"]["duration_ms"], 180)
self.assertEqual(by_name["llm_to_first_tts"]["duration_ms"], 390)
self.assertEqual(by_name["turn_total"]["duration_ms"], 1520)
tts = [item for item in timings if item["name"] == "tts_playback"]
self.assertEqual(tts[0]["duration_ms"], 200)
def test_recording_llm_provider_records_sent_time_and_response_duration(self) -> None:
class FakeLlm:
def stream_reply(self, messages):
yield ReplyDelta("你好", finish_reason="stop")
timing = TimingRecorder()
provider = RecordingLlmProvider(FakeLlm(), timing) # type: ignore[arg-type]
messages = [Message("user", "现在几点", 0.0)]
deltas = list(provider.stream_reply(messages))
self.assertEqual(deltas[0].text_delta, "你好")
self.assertEqual(len(provider.request_timings), 1)
request = provider.request_timings[0]
self.assertIn("sent_at", request)
self.assertIn("finished_at", request)
self.assertIsInstance(request["duration_ms"], int)
self.assertEqual(request["last_user_preview"], "现在几点")
if __name__ == "__main__": if __name__ == "__main__":
unittest.main() unittest.main()