The silence problem
Ask a voice assistant a question out loud, then count. One. Two. Nothing yet. Did it hear you? Did the mic cut out? Should you say it again?
That gap is the single most common way a voice interface loses someone. Cross-language research on conversational turn-taking puts the average gap between one person finishing and the next starting at roughly 200 milliseconds. It is astonishingly consistent, and it is the timing your brain is calibrated to. A machine cannot hit it. Speech recognition, then a language model, then speech synthesis — a good pipeline lands somewhere between one and three seconds, and longer if the assistant has to look something up.
You cannot close that gap by making the model faster. Two seconds of dead air is still two seconds. What you can do is make the silence legible. If the person knows why nothing is happening, the wait stops feeling like a failure.
Why a fake indicator is worse than none
Most chat products solve this with a three-dot typing indicator. Plenty of them show it the moment you hit send, on a timer, whether or not a request has actually been made. It is decoration wearing the clothes of a status signal.
That works exactly once. Trust in an interface is mostly a running tally of whether its signals predicted anything. The first time someone watches the dots animate for eight seconds and then get an error, the dots stop meaning "a response is coming" and start meaning "something is on screen." After that, the indicator is noise, and you have spent the one piece of screen real estate you had for telling the truth.
Voice makes this worse, not better. In text you can at least see that your message was sent. In voice the only evidence you spoke at all is whatever the assistant does next. If that evidence is fabricated, you have removed the person's ability to tell a working system from a broken one.
Four states worth distinguishing
A voice turn is not one operation. It is a chain, and the parts fail differently. If your interface collapses them into "loading," you have thrown away the diagnostic information the person needs.
| State | What it actually means | When it starts |
|---|---|---|
| Listening | The mic is open and the assistant is waiting for you to finish | Between turns |
| Processing | Your speech was captured and is being transcribed | The moment an utterance is accepted, not when sound is detected |
| Thinking | The transcript is with the model, which may be looking things up | After transcription returns |
| Speaking | Audio is actually playing | At the first synthesized sentence |
The distinction between processing and thinking looks pedantic until something goes wrong. If a session sticks on "processing," speech recognition is struggling — probably a noisy room, a quiet speaker, or a bad connection, and repeating yourself louder may genuinely help. If it sticks on "thinking," recognition worked and the model or a lookup is slow — repeating yourself will make it worse. Two different problems, two different reactions from the person, and a single spinner tells them neither.
The rule that makes it honest
The rule is short: a state changes when the underlying thing changes, and never before.
That sounds obvious. It is not, because there is always a convenient earlier moment.
The hard case is "speaking"
A voice turn starts with an internal event that says, in effect, the assistant now has the floor — the client uses it to reset the audio pipeline and suppress the mic. It is tempting to paint the UI from that event, because it arrives first and it is easy to reach.
But there is often a second or more between the turn starting and any audio existing. Painting "speaking" at turn start means the interface claims the assistant is talking while the room is silent — which is the exact lie we were trying to remove, just relocated. So the visible state moves to "speaking" when the first sentence of audio is ready, and the internal turn event stays internal, doing the plumbing it was always doing.
Debounce the short states
A cough, a door, a chair scrape — anything the mic accepts and then discards — would otherwise flash the whole status stack for 150 milliseconds, which reads as a glitch rather than as information. A couple hundred milliseconds of delay before "processing" appears means real turns still look instantaneous to human perception, and noise turns never light anything up.
Show the work, not just the wait
Once states are honest, the next question is what happens during the long ones. "Thinking" covers a lot of ground: it might be a one-shot answer, or it might be a knowledge-base search followed by a form being prepared.
Every tool the assistant runs emits a paired start and end, and the interface renders each one as a small chip in the transcript and as a line above the status pill: searching your documents, then a tick when it finishes. This costs almost nothing and changes the character of the wait completely. Four seconds of unexplained silence feels broken. Four seconds where you watched a document search start and finish feels like work.
The same discipline applies to failures. When a provider call fails, the honest move is to surface it as an error and stop, rather than let the state sit on "thinking" until the person gives up. A visible failure is recoverable — the person can rephrase or come back. An invisible one just looks like your product does not work.
What not to put in the label
There is a tempting mistake at the end of all this: having built accurate state, you want to write it everywhere.
We did the opposite. The pill next to the orb carries a status dot and the assistant's name — nothing else. State reads through color and motion, not words. The reason is that voice interfaces are used by people who are, by definition, not staring at the screen; a rotating word label demands a read, and a color change does not. The detailed states live in the conversation panel where someone who is looking will find them, and the ambient signal stays ambient.
Getting this split right is the difference between an interface that informs and one that nags. Same information, very different feel. If you are weighing voice against text for your own use case, the reasoning in voice AI vs text chat covers where each mode earns its place.
Why this is a product decision, not a polish task
It is easy to file honest status under "UI polish" and let it slide behind features. That is a mistake, and the tell is what people do when status is wrong: they repeat themselves. A double-ask corrupts the transcript, doubles the recognition work, and produces an answer to a question that was already being answered. The cost of a dishonest indicator is not aesthetic — it is a worse conversation, and it shows up in your logs as spoken sessions that end early. Separating those out is worth doing on its own, which is the point of tagging voice and typed conversations.
The related half of this problem is what happens when someone decides not to wait and starts talking over the assistant. That is a harder engineering problem than it sounds, and it is covered in natural interruption in voice AI.
If you want to watch the states move, start a voice conversation in the live preview in your dashboard at hiroi.ai and ask something that forces a document lookup. It is a small thing to watch, and a hard thing to fake.