Technical

Let Me Finish: How Voice AI Learns to Be Interrupted

Barge-in lets someone interrupt an AI assistant mid-sentence. A plain-language guide to echo cancellation, adaptive noise floors, and playback timing.

The 40-second answer nobody wanted

You ask a voice assistant a question. It starts a long, thorough, well-structured answer. Four seconds in, you realise it misunderstood you. You say "no, sorry —" and it keeps going. You say it louder. It keeps going. You wait out the remaining thirty seconds of an answer to a question you did not ask, then try again.

Everybody has had this experience, and it is the single largest gap between talking to a machine and talking to a person. Human conversation is full of interruption; we cut in constantly, and the other person stops. The technical name for supporting this is barge-in, and it is much harder than "stop when you hear the user talk."

Why the obvious approach fails

Here is the naive version: while the assistant is speaking, keep the microphone on. If the mic picks up sound above some threshold, assume it is the person and stop talking.

Try it and the assistant interrupts itself within half a second, every single time.

The reason is the room. If the assistant's voice is coming out of a speaker and the microphone is in the same room — which describes every laptop, phone, and tablet ever made — then the microphone hears the assistant. Loudly. That returning sound is called echo, and to a threshold detector it is indistinguishable from a person speaking. The assistant hears itself, concludes it has been interrupted, stops mid-word, listens, hears the tail of its own sentence, gets confused, and produces a response to its own voice.

So the first job is not detecting speech. It is telling the person's voice apart from the assistant's own.

Layer one: cancel the echo

Browsers ship with acoustic echo cancellation built in. The browser knows what audio it just played, so it can subtract a version of that signal from what the microphone picks up. When it works well, the assistant's own voice largely disappears from the mic feed and what is left is the room and whoever is in it.

Two related settings matter here, and one of them is counterintuitive:

  • Noise suppression: on. Removes steady background noise — fans, traffic, air conditioning. Straightforwardly good.
  • Automatic gain control: off. This one surprises people. AGC continuously adjusts microphone sensitivity to keep the input at a comfortable level. During a quiet moment it turns the gain up, which amplifies the room and the residual echo along with it. Everything downstream that reasons about loudness gets its ground shifted underneath it. Turning AGC off costs a little consistency in transcription volume and buys a stable signal to make decisions from.

Echo cancellation is never perfect. Cheap speakers, hard rooms, high volume, and Bluetooth latency all leave residue. So you cannot rely on it alone — you need a second layer that assumes some of the assistant's voice is still getting through.

Layer two: a noise floor that moves

The second layer is a threshold, but not a fixed one. A fixed threshold cannot work: the number that correctly separates speech from echo in a quiet home office is far too sensitive in a café and far too deaf in a car.

Instead, the assistant measures the echo it is actually getting right now and sets the bar above it. Concretely: while the assistant is speaking, it tracks a running average of the sound coming back through the mic, and it requires anything claiming to be a person to be meaningfully louder than that — a multiple of the observed echo, with a floor so that a perfectly silent room does not produce an absurdly low bar. In a quiet room the threshold settles low and a normal speaking voice clears it easily. In a noisy one it rises and only a deliberate interruption gets through.

The detail that makes or breaks it

The running average must only update on sound that failed the test.

If loud sound counted toward the average, a person speaking would raise the very bar they need to clear. Their first word bumps the average, the second word now faces a higher bar, and the harder they try the more the system ignores them — a maddening experience with no visible cause. Excluding qualifying sound from the average means real speech cannot raise its own bar, and the average keeps describing what it is supposed to describe: the echo, not the interruption.

Require it to last

A single loud frame is a door, a cough, a keyboard. Requiring the level to stay above the bar for a short sustained stretch removes essentially all of that, at the cost of a few tens of milliseconds of responsiveness that nobody perceives.

Problem Naive approach What actually works
Assistant hears itself Mute the mic while speaking (kills barge-in) Browser echo cancellation plus an echo-aware threshold
Room noise varies wildly One tuned threshold Threshold derived from the echo observed right now
Loud speech raises its own bar Update the average on everything Update only on sound that did not qualify
Doors, coughs, keyboards Trigger instantly Require the level to persist briefly
Mic reopens too early Reopen when sending finishes Reopen when playback actually finishes

Layer three: the gap nobody expects

This one caught us out, and it is invisible until you look for it.

The server finishes sending audio well before the speaker finishes playing it. Audio is streamed and buffered; several seconds of speech can be fully transmitted while the person is still hearing the middle of it. If you decide "the assistant has stopped talking, open the microphone" at the moment sending completes, you open the mic into the tail of your own output. The assistant then hears the last two seconds of its own answer, treats it as a new utterance, transcribes it, and replies to itself. The transcript ends up containing the assistant's own words attributed to the visitor, which is a genuinely disorienting thing to read.

The fix is to have the client — the only party that knows when the last sample actually left the speaker — say so explicitly, and treat that message as the authoritative signal to arm the microphone. Every microphone frame in the audible window is discarded before it reaches speech recognition, not filtered out afterwards. Filtering afterwards means depending on the content of transcripts to decide what was real, which is a losing game: block phrases and you will eventually block a visitor who said the same thing.

For older clients that cannot report playback completion, a duration estimate derived from the audio size serves as a fallback. It deliberately errs long — a slightly late microphone is a small annoyance, an early one poisons the conversation.

What good barge-in feels like

When all three layers are working, the experience is unremarkable, which is the point. You start talking, the assistant stops within a couple hundred milliseconds, and it responds to what you said rather than to the half-sentence it was in the middle of. You never think about it.

When one layer is missing, it is very obvious. No echo cancellation and the assistant stutters and interrupts itself. Fixed threshold and it works perfectly in the demo room and nowhere else. No playback boundary and it periodically answers itself. All three failures read to a user as "this thing is broken," and none of them look like the underlying cause — which is why the place to hunt for them is the set of spoken sessions that ended early, not the aggregate. Tagging voice and typed conversations separately is what makes that set findable.

The companion problem is the silence before the assistant speaks — what to show during the second or two of processing, and why fake indicators make it worse. That is covered in honest voice status. And if you are still deciding whether voice belongs in your product at all, voice AI vs text chat lays out where each mode wins.

Barge-in is on by default in browser voice conversations with hiroi. The best way to evaluate it is the rude way: ask a question that will produce a long answer, then cut in three seconds later and see what happens.

Try hiroi free.

Put an AI agent on your site for chat and voice — no credit card required.