Picking a voice feels like a branding decision and mostly is not. The voice a visitor remembers is the one that pronounced your product name correctly, paused in the right places, and did not sound like it was reading a bulleted list aloud — because it was.
Here is how to set voice up properly: deciding whether you want it, choosing the voice, teaching it your vocabulary, and testing it before anyone else hears it.
First: do you want voice on this site?
Voice in Hiroi runs in the visitor's browser, over their microphone. No phone number, no telephony, nothing to provision. That makes it cheap to try — but not right for every site.
| Voice earns its place when | Text-only is the better call when |
|---|---|
| Hands or eyes are busy — kitchens, workshops, cars, shop floors | Visitors are at a desk in a shared office and will not talk out loud |
| The answer is a short explanation, better heard than read | The answer is a code, a SKU or a URL that must be copied exactly |
| Accessibility matters and reading is the barrier | Your audience is largely in quiet or public spaces |
| The device is a kiosk or tablet where typing is awkward | Conversations routinely involve sensitive details said aloud |
You do not have to choose exclusively. Voice is an option in the widget, not a mode you force on anyone — a visitor can type on a busy train and speak at home. Enabling it adds a route; it does not remove one. Voice AI vs Text Chat makes the fuller case.
Turning voice on
Open your assistant and select the Voice tab. A spoken conversation needs both switches:
- Enable Voice Input — lets visitors speak instead of typing. With it off, the widget is text-only: the call button disappears from the chat header and holding the orb does nothing.
- Enable Text-to-Speech — reads the assistant's replies aloud. Until this is on, the voice picker is replaced by a prompt to turn it on.
The voice connection is refused when text-to-speech is off, even if voice input is on. If you flipped one switch and nothing happened, that is why. Save with Save (⌘S / Ctrl+S), like every other tab.
Choosing the voice
Under Text-to-Speech, two providers sit in a segmented control.
Azure Neural is the default catalogue, available on every account. It covers US, UK, Australian and Indian English plus Spanish, French, German, Portuguese, Italian, Japanese, Korean, Chinese, Swedish, Turkish and more.
ElevenLabs is marked Premium and is a paid-account feature — free accounts are pinned to Azure. It also consumes noticeably more credits per character, which matters on a high-traffic site; see the pricing page.
Three controls narrow the list: Search matches on name, locale, accent and style, Locale chips filter the Azure catalogue by accent, and Gender chips apply to both providers. Azure names also carry their tier — Premium HD voices are more natural and expressive, Standard voices cost less to speak. On a heavy-traffic site, a Standard voice that suits your brand beats a Premium HD one that does not.
Every row has a play button for a short sample. Click a name to select it; the Selected field below always shows what will be saved.
How to actually choose
Sampling voices in isolation is misleading — they all sound fine reading one cheerful sentence. Do this instead:
- Shortlist three voices from the sample button. No more.
- Save one and open Test Widget on the Deploy tab.
- Ask a real question — one whose answer contains your product name, a number and an abbreviation.
- Listen for the three things that actually break: mispronounced proper nouns, numbers read as digits when they should be words, and pace too fast to follow.
- Repeat with the other two. Six minutes, and it is the only test that predicts what visitors hear.
Match the accent to your audience, not your own ear. A UK audience notices an American voice immediately, and the reverse is equally true.
Teach it your vocabulary
Speech recognition mis-hears proper nouns more than anything else — product names, brand names, menu items, local place names. Speech vocabulary (key terms), under Voice Input, is where you fix that: anything listed is passed to the recogniser as a hint so those words are favoured.
- Separate terms with commas, semicolons or new lines.
- Your assistant's name is included automatically — do not repeat it.
- Duplicates are removed, terms longer than 60 characters are dropped, and the first 100 are used.
Good entries are the words a stranger would spell wrong: your product names, your practitioners' surnames, your model numbers, the street your shop is on. This one field fixes more voice complaints than anything else on the tab.
Voice Language is a recognition hint, not a translation setting
Voice Language tells the recogniser what to expect. Leave it on Auto-detect (default) if your visitors are multilingual — it then leans on the language the visitor last spoke and your assistant's primary language rather than assuming English. Set an explicit language when every visitor speaks the same one and you want the fewest possible mis-hearings.
It does not change what language the assistant replies in. That is Response Language Mode, on the General tab.
Accessibility settings: leave them on
Three controls under Accessibility decide how much of the conversation is also shown as text.
| Control | Default | What it does |
|---|---|---|
| Show Speech Transcription | On | Briefly echoes what the assistant heard the visitor say |
| Show AI Response Text | On | Shows the spoken reply as text as well as audio |
| Transcript Display Duration | 2.5 s | How long a spoken message stays visible before fading (0.5–30 s) |
The two toggles are what make a voice conversation usable for someone who cannot hear the reply, and what lets a visitor check a number was heard correctly. Turn them off only if the text genuinely gets in the way.
Under Voice Input, Enable Waveform shows the animated ring on the orb while either party speaks. Visitors who have asked their system for reduced motion get a still waveform either way.
Write your instructions for the ear
Chat replies and spoken replies read from the same Instructions field on the General tab. Formatting that looks tidy on screen — bullets, bold headings, tables — becomes an unlistenable monotone read aloud. One line usually fixes it: "Answer in short spoken sentences. Use a list only when the visitor asks for one." Set Up Your Assistant's Persona and Tone covers the rest of that field.
Long answers are worse aloud than on screen too. If you have not already capped reply length, voice is the reason to.
Testing before anyone hears it
Voice needs two things the text widget does not: a secure page and microphone permission. Both are easy to miss.
- Save the assistant — Preview and Test Widget both load the last saved version.
- Open Test Widget from the Deploy tab and grant the microphone prompt.
- Press and hold the orb for about half a second to start a voice session. The assistant greets you, and the orb moves through its listening, processing, thinking and speaking states as the turn progresses.
- Talk over it mid-sentence. It should stop, the same way a person would.
- Ask something whose answer contains a proper noun you added to the vocabulary field, and check it comes back correctly in the transcription line.
- Tap the orb again to end the session.
If the voice option is not there at all, check both switches on the Voice tab. If it is there but nothing happens, the browser has denied microphone permission for that site, or the page is not served over HTTPS.
Finally, look at the orb's state colours on the Appearance tab. Listening and Speaking carry the whole "is it my turn?" signal, so pick two a visitor can tell apart at a glance — covered in Make the Widget Match Your Brand.
Voice draws on your credit balance on top of the normal chat turn, since it adds transcription and speech synthesis. Check the pricing page before turning it on for a high-traffic page.
Try it on an assistant at hiroi.ai — the test widget is enough to hear the difference between three voices in about six minutes.