The 940ms Voice Agent: Where Every Millisecond Goes
A conversation has a social latency budget, not a technical one — go past a second and people think the thing is broken. How a voice agent went from 1,530ms to 940ms to first audio, why every millisecond came from overlapping stages rather than speeding them up, the speculative optimisation that cost 38% more tokens for 60ms, and what a user hears when a dependency vanishes mid-sentence.
Ask a person a question and they start answering in about 200 milliseconds. Go past roughly a second and something changes in the room — they think you did not hear them, or that the thing is broken, and they start talking again over the top of your answer. Voice does not have a latency budget in the way an API does. It has a social one, and you can feel yourself violating it.
My first working voice agent took 1,530 ms to start speaking. It was correct, well grounded, and unpleasant to talk to. This is how it got to 940, what each change actually bought, the optimisation that looked brilliant and was not, and what happens when a dependency disappears mid-sentence.
Why this is hard mode
Nobody notices 400 ms in an API
In a request/response system, 400 ms of extra work is a number in a dashboard. In a conversation it is the pause where somebody says "hello?". And the components are stacked in series, so every stage's latency is additive and the slowest one sets the floor for the whole experience.
Worse, the pipeline has five moving parts and four of them are somebody else's service:
Phase one
Get it flowing, optimise nothing
The temptation is to start fast. Do not — you will optimise a pipeline whose shape is about to change. Phase one is one WebSocket carrying typed events in both directions and a working conversation at whatever speed it happens to run.
async def handle(ws: WebSocket) -> None:
session = Session(ws)
async for frame in ws.audio_frames(): # 20 ms PCM
result = await asr.feed(frame)
if result.is_partial:
await session.emit("partial_transcript", text=result.text)
continue
await session.emit("final_transcript", text=result.text)
chunks = await retriever.search(result.text)
async for token in llm.stream(result.text, chunks):
await session.emit("token", text=token)
if sentence_boundary(token):
await tts.speak(session.take_sentence()) # speak as we go
Getting that to work end to end took longer than I expected and produced no impressive number at all. It is still the right order: you cannot decompose a latency budget for a pipeline that does not yet run.
Phase two
You cannot optimise "one and a half seconds"
The single most useful thing I did was stop measuring one number. Total latency is not actionable; a per-stage breakdown, recorded on every request, tells you precisely which service to go after.
That last point is the lesson of the whole phase. I spent the first week trying to make individual stages faster and got almost nowhere; every stage was already close to its floor. The 590 ms came from changing what waits for what.
The three changes that worked
| Change | Saved | What it costs |
|---|---|---|
| Retrieve on the partial transcript instead of waiting for ASR to finalise | 260 ms | ~30% of retrievals are thrown away when the final transcript differs |
| Sentence-level TTS — speak sentence one while the model writes sentence two | 150 ms | Slightly worse prosody across sentence boundaries |
| Keep the reranker warm and shrink the candidate set from 30 to 20 | 110 ms | A resident process; recall 0.89 → 0.88 |
| Binary audio frames instead of base64 JSON on the socket | 70 ms | A less pleasant protocol to debug by eye |
The one that did not work
I was very pleased with this idea for about a day. If retrieval can start on a partial transcript, why not the LLM? Start generating speculatively at the partial, and if the final transcript matches, you have already paid the time to first token.
60 ms off p50 time to first audio.
Real, measurable, and it did work when the speculation was correct.
+38% tokens, because roughly three speculations in ten were discarded and paid for anyway.
And a race condition where a discarded generation's tokens occasionally reached the client, which is a much worse bug than a slow answer.
Sixty milliseconds for a 38% cost increase and a correctness hazard. I reverted it. Retrieval speculation survives the same trade because a wasted vector search costs a few milliseconds of CPU, not tokens — the asymmetry is the whole reason one idea is good and its obvious sibling is not.
Phase three
What happens when a dependency vanishes mid-sentence
Four of the five stages are network calls to services I do not control. The question is not whether one fails — it is what the person hears when it does. There is exactly one unacceptable answer, and it is silence.
| Failure | Detected by | Degradation | What the user gets |
|---|---|---|---|
| ASR unavailable | Socket close or 3 s without a partial | Switch to typed input, announce it | "I'm having trouble hearing — you can type instead" |
| Cloud LLM timeout | 4 s deadline on first token | Fall back to the local 3B model | A shorter answer, 260 ms later |
| TTS unavailable | 5 s timeout | Return the answer as text on screen | The answer, silently |
| Retrieval empty | No chunk over the score threshold | Refuse and offer to rephrase | "I don't have anything on that" |
async def generate(query: str, chunks: list[Chunk]) -> AsyncIterator[str]:
try:
async with asyncio.timeout(4.0):
first = await anext(cloud.stream(query, chunks))
yield first
async for token in cloud.stream_rest():
yield token
return
except (TimeoutError, ProviderError) as e:
log.warning("llm.fallback", reason=type(e).__name__)
# The local model from project two. Worse answers, but it is on this box
# and it cannot fail for the same reason the cloud just did.
await emit("degraded", stage="llm")
async for token in local.stream(query, chunks):
yield token
Note the degraded event. The client uses it to show a small indicator, and that is deliberate: a system quietly giving worse answers is how you lose trust slowly. Saying so costs nothing and is the difference between a fallback and a lie.
Replay mode
The last thing I built and the first thing I would build next time. 74 recorded sessions — raw audio in, every event, every timing — that can be fed back through the pipeline deterministically.
Without it, debugging voice means talking to your laptop and trying to reproduce a thing that happened once, and every attempt is a slightly different input. With it, "it did something weird yesterday" becomes a fixture. It also means a latency optimisation can be verified against 74 real conversations instead of whatever sentence you happen to say while testing, which is invariably the clearest sentence you know how to say.
Honesty
What 940 ms does not include
- It is a p50 on my network. p95 is 1.8 s and that is the number a user forms an opinion from. Quote the p95.
- The VAD's 240 ms is a choice, not a cost. It is the silence I wait for before deciding somebody has stopped talking. Shorten it and you interrupt people; lengthen it and you feel slow. It is a UX dial disguised as a latency number.
- Barge-in is not solved here. Talking over the agent while it speaks works, but the cancellation is coarse and occasionally clips a syllable.
- One language, one accent, quiet room. ASR latency and accuracy both move considerably outside that, and I have not measured how much.
- No concurrency testing. Every number here is a single session on an idle machine. The warm reranker that saves 110 ms is a shared resource, and I do not know what it does at fifty sessions.
The checklist
- A typed event protocol first, before any optimisation.
- Per-stage latency on every request, not a single total.
- Report p95, not p50, when describing how it feels.
- Overlap stages before trying to speed them up — the wins are in the waiting.
- Speculate only where waste is cheap: retrieval yes, generation no.
- Warm anything with a cold-start penalty, and measure the first request after idle.
- A timeout on every external stage, with a named fallback.
- Announce degradation rather than silently serving worse answers.
- Never fail silent — say something, even if it is "one moment".
- Replay mode, built early, from real recorded sessions.
That closes the series
Six posts, four projects, one system. A retrieval pipeline that refuses when it cannot ground an answer; an agent loop with guardrails on every side; an evaluation set that blocks a bad merge; a local model that catches a cloud timeout; a tracing layer that makes silent failure visible; and this, the streaming shell that puts a voice on all of it.
If there is one thing I would hand to somebody starting: the projects that look weakest in a portfolio — the eval set, the tracing layer — are the two that did the most work. Everything else was easier to build and much easier to be wrong about.
- WebSocket events
- VAD
- partial transcripts
- TTFT
- sentence TTS
- warm pools
- deadlines
- replay mode
The full twelve-week write-up with every number in one place is here. If you are building a voice agent and have solved barge-in properly, I would genuinely like to know how — that is the piece I am least happy with. The comments are open.
Comments (0)
No comments yet
Be the first to share a thought on this article.
Join the conversation