Skip to content
A AhsanLab.Tech
AI 15 min read · April 12, 2026

The 940ms Voice Agent: Where Every Millisecond Goes

A conversation has a social latency budget, not a technical one — go past a second and people think the thing is broken. How a voice agent went from 1,530ms to 940ms to first audio, why every millisecond came from overlapping stages rather than speeding them up, the speculative optimisation that cost 38% more tokens for 60ms, and what a user hears when a dependency vanishes mid-sentence.

A Ahsan Habib Save

Ask a person a question and they start answering in about 200 milliseconds. Go past roughly a second and something changes in the room — they think you did not hear them, or that the thing is broken, and they start talking again over the top of your answer. Voice does not have a latency budget in the way an API does. It has a social one, and you can feel yourself violating it.

My first working voice agent took 1,530 ms to start speaking. It was correct, well grounded, and unpleasant to talk to. This is how it got to 940, what each change actually bought, the optimisation that looked brilliant and was not, and what happens when a dependency disappears mid-sentence.

940 msto first audio, p50from 1,530 ms
1.8 sp95the number that matters
3tested degradation pathsnone of them silence

Why this is hard mode

Nobody notices 400 ms in an API

In a request/response system, 400 ms of extra work is a number in a dashboard. In a conversation it is the pause where somebody says "hello?". And the components are stacked in series, so every stage's latency is additive and the slowest one sets the floor for the whole experience.

Worse, the pipeline has five moving parts and four of them are somebody else's service:

THE PIPELINE · ONE WEBSOCKET, FIVE STAGES MIC 20 ms frames VAD + ASR partials, then final RETRIEVE hybrid + rerank LLM streamed tokens TTS per sentence FOUR OF THE FIVE ARE SOMEBODY ELSE'S SERVICE · EACH NEEDS A TIMEOUT AND A FALLBACK events: audio_frame → partial_transcript → final_transcript → retrieval_done → token → sentence_ready → audio_chunk → turn_complete
FIGURE 1 — STRUCTURE THE EVENTS FIRST · Every optimisation later in this post is a rearrangement of when these events fire, so the protocol is the thing to get right in phase one.

Phase one

Get it flowing, optimise nothing

The temptation is to start fast. Do not — you will optimise a pipeline whose shape is about to change. Phase one is one WebSocket carrying typed events in both directions and a working conversation at whatever speed it happens to run.

voice/session.py — typed events, no cleverness yet
async def handle(ws: WebSocket) -> None:
    session = Session(ws)

    async for frame in ws.audio_frames():          # 20 ms PCM
        result = await asr.feed(frame)

        if result.is_partial:
            await session.emit("partial_transcript", text=result.text)
            continue

        await session.emit("final_transcript", text=result.text)

        chunks = await retriever.search(result.text)
        async for token in llm.stream(result.text, chunks):
            await session.emit("token", text=token)
            if sentence_boundary(token):
                await tts.speak(session.take_sentence())   # speak as we go

Getting that to work end to end took longer than I expected and produced no impressive number at all. It is still the right order: you cannot decompose a latency budget for a pipeline that does not yet run.

Phase two

You cannot optimise "one and a half seconds"

The single most useful thing I did was stop measuring one number. Total latency is not actionable; a per-stage breakdown, recorded on every request, tells you precisely which service to go after.

TIME TO FIRST AUDIO · BEFORE AND AFTER 0 500 ms 1000 ms 1500 ms BEFORE 1,530 ms VAD 240 ASR 310 retrieval 290 LLM TTFT 420 TTS 180 AFTER 940 ms VAD 240 ASR 250 LLM TTFT 420 retrieval 180 · runs UNDER the ASR tail, off the critical path 940 ms 1,530 ms the LLM slice is IDENTICAL in both rows · every millisecond came from overlap, not from making a stage faster
FIGURE 2 — THE BUDGET · The LLM's time to first token is the largest single slice and I never reduced it. The wins came from what runs alongside it.

That last point is the lesson of the whole phase. I spent the first week trying to make individual stages faster and got almost nowhere; every stage was already close to its floor. The 590 ms came from changing what waits for what.

The three changes that worked

ChangeSavedWhat it costs
Retrieve on the partial transcript instead of waiting for ASR to finalise 260 ms ~30% of retrievals are thrown away when the final transcript differs
Sentence-level TTS — speak sentence one while the model writes sentence two 150 ms Slightly worse prosody across sentence boundaries
Keep the reranker warm and shrink the candidate set from 30 to 20 110 ms A resident process; recall 0.89 → 0.88
Binary audio frames instead of base64 JSON on the socket 70 ms A less pleasant protocol to debug by eye
The one worth stealing Warming the reranker. A cold cross-encoder added roughly 600 ms to the first request after any idle period — which is precisely the request a person is most likely to be forming their opinion on. Every latency number you measure in a tight benchmark loop is a warm number, and no real user ever gets one.

The one that did not work

I was very pleased with this idea for about a day. If retrieval can start on a partial transcript, why not the LLM? Start generating speculatively at the partial, and if the final transcript matches, you have already paid the time to first token.

What it bought

60 ms off p50 time to first audio.

Real, measurable, and it did work when the speculation was correct.

What it cost

+38% tokens, because roughly three speculations in ten were discarded and paid for anyway.

And a race condition where a discarded generation's tokens occasionally reached the client, which is a much worse bug than a slow answer.

Sixty milliseconds for a 38% cost increase and a correctness hazard. I reverted it. Retrieval speculation survives the same trade because a wasted vector search costs a few milliseconds of CPU, not tokens — the asymmetry is the whole reason one idea is good and its obvious sibling is not.

Phase three

What happens when a dependency vanishes mid-sentence

Four of the five stages are network calls to services I do not control. The question is not whether one fails — it is what the person hears when it does. There is exactly one unacceptable answer, and it is silence.

FailureDetected byDegradationWhat the user gets
ASR unavailable Socket close or 3 s without a partial Switch to typed input, announce it "I'm having trouble hearing — you can type instead"
Cloud LLM timeout 4 s deadline on first token Fall back to the local 3B model A shorter answer, 260 ms later
TTS unavailable 5 s timeout Return the answer as text on screen The answer, silently
Retrieval empty No chunk over the score threshold Refuse and offer to rephrase "I don't have anything on that"
voice/resilience.py — a deadline, a fallback, and an announcement
async def generate(query: str, chunks: list[Chunk]) -> AsyncIterator[str]:
    try:
        async with asyncio.timeout(4.0):
            first = await anext(cloud.stream(query, chunks))
            yield first
            async for token in cloud.stream_rest():
                yield token
        return
    except (TimeoutError, ProviderError) as e:
        log.warning("llm.fallback", reason=type(e).__name__)

    # The local model from project two. Worse answers, but it is on this box
    # and it cannot fail for the same reason the cloud just did.
    await emit("degraded", stage="llm")
    async for token in local.stream(query, chunks):
        yield token

Note the degraded event. The client uses it to show a small indicator, and that is deliberate: a system quietly giving worse answers is how you lose trust slowly. Saying so costs nothing and is the difference between a fallback and a lie.

Every timeout in this pipeline is a promise about the worst case. A stage without one is a stage that can hang the conversation forever, and "forever" is about four seconds before somebody hangs up.

Replay mode

The last thing I built and the first thing I would build next time. 74 recorded sessions — raw audio in, every event, every timing — that can be fed back through the pipeline deterministically.

Without it, debugging voice means talking to your laptop and trying to reproduce a thing that happened once, and every attempt is a slightly different input. With it, "it did something weird yesterday" becomes a fixture. It also means a latency optimisation can be verified against 74 real conversations instead of whatever sentence you happen to say while testing, which is invariably the clearest sentence you know how to say.

Honesty

What 940 ms does not include

  • It is a p50 on my network. p95 is 1.8 s and that is the number a user forms an opinion from. Quote the p95.
  • The VAD's 240 ms is a choice, not a cost. It is the silence I wait for before deciding somebody has stopped talking. Shorten it and you interrupt people; lengthen it and you feel slow. It is a UX dial disguised as a latency number.
  • Barge-in is not solved here. Talking over the agent while it speaks works, but the cancellation is coarse and occasionally clips a syllable.
  • One language, one accent, quiet room. ASR latency and accuracy both move considerably outside that, and I have not measured how much.
  • No concurrency testing. Every number here is a single session on an idle machine. The warm reranker that saves 110 ms is a shared resource, and I do not know what it does at fifty sessions.

The checklist

  • A typed event protocol first, before any optimisation.
  • Per-stage latency on every request, not a single total.
  • Report p95, not p50, when describing how it feels.
  • Overlap stages before trying to speed them up — the wins are in the waiting.
  • Speculate only where waste is cheap: retrieval yes, generation no.
  • Warm anything with a cold-start penalty, and measure the first request after idle.
  • A timeout on every external stage, with a named fallback.
  • Announce degradation rather than silently serving worse answers.
  • Never fail silent — say something, even if it is "one moment".
  • Replay mode, built early, from real recorded sessions.

That closes the series

Six posts, four projects, one system. A retrieval pipeline that refuses when it cannot ground an answer; an agent loop with guardrails on every side; an evaluation set that blocks a bad merge; a local model that catches a cloud timeout; a tracing layer that makes silent failure visible; and this, the streaming shell that puts a voice on all of it.

If there is one thing I would hand to somebody starting: the projects that look weakest in a portfolio — the eval set, the tracing layer — are the two that did the most work. Everything else was easier to build and much easier to be wrong about.

  • WebSocket events
  • VAD
  • partial transcripts
  • TTFT
  • sentence TTS
  • warm pools
  • deadlines
  • replay mode

The full twelve-week write-up with every number in one place is here. If you are building a voice agent and have solved barge-in properly, I would genuinely like to know how — that is the piece I am least happy with. The comments are open.

#Real-Time #Performance #AI Tools #LLM #Agents #Agentic AI

Comments (0)

No comments yet

Be the first to share a thought on this article.

Join the conversation

Comments are moderated before they appear.

Keep reading

Related articles

AI 13 min

Connecting OpenAI to a Real Backend: Lessons From SEO Automation

The naive version worked perfectly on fifty pages and failed four ways on twelve thousand — none of them about prompt quality. Wiring a model into a real backend is a batch-processing problem with an expensive, rate-limited, non-deterministic function in the middle, and it needs everything batch processing always needed.

May 28, 2026
AI 14 min

Your AI System Is Failing Right Now and Nothing Is Alerting

On a Tuesday my RAG system got quietly worse. Every request returned 200, no alert fired, and the answers stayed fluent and confident while being subtly about the wrong thing. Why AI systems fail silently, what a trace must actually record, why the mean is a lie, and the forty-minute root cause that would have taken days.

Mar 27, 2026