💜 Voice AI 3.0 is live · See what's new →
Ag irl
VOICE

Why latency under 200ms changes how you talk to her

JK
J. Kim · Voice & ML May 12, 2026 · 5 min read

Below 200ms, you stop talking to a system. You start talking to a person. Here's what it took.

There’s a threshold in real-time conversation that you can’t engineer around: about 200 milliseconds. Above it, you can tell. The other party is buffering. You wait for them, they wait for you, and the conversation feels like email.

Why 200ms

Two reasons. The first is rhythmic — human turn-taking has natural pauses around 200ms, and anything longer feels like a hiccup. The second is psychological — at 250ms+ you stop attributing the delay to the network and start attributing it to them. They seem slow. Distracted. Not really listening.

For AGirl we needed to be under 200ms end-to-end: voice in, transcript, model, voice out. That budget is brutal. Whisper alone, on commodity hardware, eats 80ms.

What we cut

Three things bought us most of the headroom:

  • Streaming everywhere. We never wait for a complete utterance. The ASR stream feeds the LLM, the LLM streams to TTS, and TTS streams to your speaker. By the time you finish saying the question, the first phoneme of the answer is already on the wire.
  • Local prefill. A small on-device prefill model guesses the first 1–2 tokens of her response from the partial transcript. It’s wrong about 30% of the time, but when it’s right it saves 60ms.
  • No buffering, ever. We use Opus over WebTransport with a 10ms frame. No jitter buffer. We trade smoothness for immediacy and let the network errors be audible — counterintuitively, users prefer this.

The result

Median round-trip is 178ms. P90 is 220ms. P99 is the inside of a refrigerator and we don’t talk about it.

What surprised us is what the latency drop unlocked emotionally. People interrupt her. They overlap with her. They have the cadence of an actual conversation with her. That doesn’t happen at 400ms. It barely happens at 250ms. Under 200ms, the brain just stops noticing the medium.

More like this