Product

Voice or Typed: What You Learn When Your Transcripts Know the Difference

Tagging conversations as voice or text changes what your analytics tell you. How spoken and typed sessions differ in intent, length, and what they need.

An average that describes nobody

Open your conversation log. You have a list of transcripts, a length, a timestamp, and a rough sense of whether each one went well. You compute an average message length, an average session length, a completion rate.

If some of those conversations were spoken and some were typed, every one of those averages is a blend of two populations that behave nothing alike. The mean sits in the gap between them and describes no actual visitor. You then tune your assistant against that mean, which is how you end up with an assistant that is slightly wrong for everybody.

This is not a subtle effect. Speaking runs at roughly 125 to 150 words per minute. Typing runs at 38 to 40 on a keyboard and closer to 20 to 25 with thumbs on a phone. Before anyone has said anything interesting, the two channels differ by a factor of four in how much input arrives per unit of patience.

Tag the mode, per message

The fix is to record how each message arrived and let the transcript carry it. In practice you want three labels, not two:

  • Voice — every visitor message in the session was spoken
  • Text — every visitor message was typed
  • Mixed — the session crossed over

Mixed is the interesting one, and it is the reason you tag per message rather than per session. People switch, and where they switch is diagnostic. Someone who talks for six turns and then types their email address is telling you that spelling an address out loud is unpleasant. Someone who types three questions and then starts talking has usually just gotten comfortable. Someone who switches to typing immediately after a misrecognition is telling you your recognition failed and they gave up on it rather than complaining.

None of that is visible if the mode is a session-level flag.

What the two channels actually look like

Once you can filter, the differences show up fast and they are consistent enough to design around.

Spoken Typed
Message shape One long run-on containing the whole situation Short fragments across several turns
Turns per session Fewer More
Question type Open-ended, situational, "here is what is going on" Specific, lookup-flavoured, "what is the price of X"
Tolerance for long answers Low — listening tops out near 150 wpm High — reading runs near 250 wpm
Precise data Avoided Willingly typed
Where it happens Mobile, hands-busy, after hours Desktop, working hours

The practical consequences are not symmetrical. A voice visitor who gets a 200-word answer sits through 80 seconds of talking and will interrupt or leave. A text visitor who gets a 30-word answer often has to ask three follow-ups to get what they needed. The same "good" answer length is wrong in both directions depending on channel — which is why an assistant that speaks its typed replies, or writes the way it talks, feels subtly broken in a way that is hard to name until you separate the two.

The other consistent finding: spoken sessions front-load intent. Because talking is cheap, people describe their whole situation in the first utterance instead of rationing it across turns. That is genuinely useful — the first spoken message in a session is frequently the single most informative thing in your entire log, and it is buried in an average with a typed first message that says "hi."

What to actually look at

A few filters earn their keep immediately.

Voice sessions that ended without resolution

Usually one of three things: an answer too long to sit through, a recognition failure the person did not bother to report, or a request for structured data that is painful to speak. All three are fixable, and none of them show up in a blended view.

Typed sessions on mobile

If people are typing on phones when voice was right there, either the entry point is not obvious or something about the voice experience put them off. Read a few in full.

The switch point in mixed sessions

Cluster on what the assistant said immediately before the visitor changed channel. There is almost always a pattern, and it is almost always something you can fix in one edit.

Voice sessions by hour

Spoken traffic skews to evenings and commutes far more than typed traffic does. If your follow-up path assumes someone is at a desk, this is where you find out that assumption is wrong.

What we deliberately do not store

Tagging the mode raises an obvious question: are you keeping the audio?

No. We store the transcript and the mode label, not the recording. This is a deliberate constraint rather than a storage decision.

A voice recording is biometric data in a growing number of jurisdictions — Illinois's BIPA is the well-known example, and it is not the only one. Retaining voiceprints pulls a small business into a consent-and-retention regime it almost certainly has not budgeted for, in exchange for data that is worse than the transcript for every analytical purpose anybody actually has. Nobody reviews audio at volume. They read transcripts.

So the rule is: the mode is metadata, the transcript is the record, and the audio ends when the turn does. If you are going to be asked "do you record calls?" — and you will be — the cleanest answer is the one where you never started.

Where this feeds back

The reason to separate voice from text is not the report. It is that it changes what you fix.

If your voice completion rate is well below your text completion rate, the problem is almost never the model. It is response length, or status feedback during the wait, or a data-collection step that assumes a keyboard. Each of those has a specific remedy: shorter spoken answers, honest state signals during the gap — the reasoning in honest voice status — and a form panel instead of an assistant asking someone to spell an address, which is the case for in-chat forms.

If your text sessions run long with many turns and low satisfaction, the problem is usually retrieval or answer completeness, not modality at all.

Same assistant, same prompt, two entirely different diagnoses. You cannot reach either one from a blended average. For the broader set of metrics worth tracking alongside the mode split, using agent analytics to improve your agent covers the rest of the picture.

Conversations in the hiroi.ai dashboard carry the mode on every message, with a filter for spoken sessions. The first hour you spend reading only the voice ones tends to be worth more than the previous month of aggregate charts.

Try hiroi free.

Put an AI agent on your site for chat and voice — no credit card required.