Voice
The Voice tab controls whether visitors can speak to your assistant, which voice it speaks back in, and how much of the conversation is shown as text. Open an assistant from Assistants, then select Voice.

The tab has four sections: Voice Input, Text-to-Speech, Voice Language and Accessibility. Changes are saved with Save (⌘S / Ctrl-S) like every other tab.
Note
A spoken conversation needs both Enable Voice Input and Enable Text-to-Speech turned on. The voice connection is refused when text-to-speech is off, even if voice input is enabled.
Voice Input
Enable Voice Input lets visitors speak instead of typing. With it off, the widget is text-only: the call button disappears from the chat header and holding the orb does nothing.
Enable Waveform
Shows the animated waveform ring inside the orb while the visitor and the assistant speak. Visitors who have asked their system for reduced motion get a still waveform either way.
Note
This switch only appears in Widget display mode with Show waves off — the one configuration where the orb itself is the surface. In Minimal, and in Widget with the wave skin on, there is no orb canvas for it to govern, so the control is hidden. Minimal's waves are Waves at the Bottom on the Appearance tab.
Auto-greeting
There is no auto-greet switch. The assistant always speaks first when a voice session starts — a silent orb waiting for the visitor reads as broken — so the behaviour is forced on for every assistant. No greeting is sent when the conversation is already underway on screen.
Speech vocabulary (key terms)
Speech recognition mis-hears proper nouns more than anything else — product names, brand names, menu items, local place names. Anything you list here is passed to the recogniser as a vocabulary hint so those words are favoured.
- Separate terms with commas, semicolons or new lines.
- Your assistant's name is included automatically; you do not need to repeat it.
- Duplicates are removed, terms longer than 60 characters are dropped, and the first 100 terms are used.
Good entries are the words a stranger would spell wrong: Hiroi, Cafecito, Dr. Okonkwo, NPS-40.
Text-to-Speech
Enable Text-to-Speech reads the assistant's replies aloud. Until it is on, the voice picker is replaced by a prompt to turn it on.
Choosing a provider
Two providers sit in a segmented control at the top of the picker.
- Azure Neural — the default catalogue, with the number of available voices next to the label. Covers English (US, UK, Australian, Indian), Spanish, French, German, Portuguese, Italian, Japanese, Korean, Chinese, Swedish, Turkish and more.
- ElevenLabs — marked Premium. If the ElevenLabs catalogue is unavailable it is marked Off and cannot be selected.
Until you pick one, Selected reads Platform default (varies by detected language). There is no single default voice: the platform resolves one per language, so a Spanish turn and an English turn in the same conversation are spoken by different voices. Choose a voice explicitly if you want one voice everywhere.
Selecting ElevenLabs shows a warning in the panel that ElevenLabs voices cost several times more credits than Azure ones. On the rates currently charged it is close to 6× — 35 credits per 1,000 characters against 6 for Azure standard. See Billing and credits.
Paid feature. Free accounts are limited to Azure voices — the service forces the Azure provider on a free account. Add credits to your account to unlock ElevenLabs.
Finding a voice
- Search matches on name, locale, accent and style.
- Locale chips filter the Azure catalogue: All, US English, British English, Australian English, Spanish (Mexico), French, German, Brazilian Portuguese, Japanese. The locale row is Azure-only.
- Gender chips — All, Female, Male — apply to both providers.
Azure voice names tell you the tier they belong to. Voices labelled Premium HD are Azure's Dragon HD tier — more natural and more expressive. Voices labelled Standard are regular neural voices and cost less to speak.
Each row in the list carries a play button. Press it to hear that voice say a short sample — "Hi, I'm your AI assistant. How can I help you today?" — before you commit to it. Click the voice's name to select it; a check mark marks the current choice, and the Selected field below the list always shows the voice that will be saved.
Voice Language
Voice Language is a hint for speech recognition, not a translation setting. Leave it on Auto-detect (default) for multilingual visitors: the recogniser then leans on the language the visitor last spoke and the assistant's own primary language rather than assuming English.
Setting an explicit language — English, Spanish, French, German, Italian, Portuguese, Japanese, Korean, Chinese, Arabic, Hindi, Polish, Dutch, Russian, Swedish or Turkish — pins recognition to that language and turns auto-detection off. Pin it when you know every caller speaks the same language and you want the fewest possible mis-hearings.
Accessibility
| Control | Default | What it does |
|---|---|---|
| Show Speech Transcription | On | Briefly echoes what the assistant heard the visitor say, in the voice status line. |
| Show AI Response Text | On | Shows the assistant's spoken reply as text as well as audio. |
| Transcript Display Duration | 2.5 s | How long a spoken message stays visible above the orb before fading. Accepts 0.5 to 30 seconds, in tenths. |
Leave both toggles on unless the text genuinely gets in the way — they are what makes a voice conversation usable for someone who cannot hear the reply.
How visitors use voice
In the Widget presentation the orb is the microphone control; in Minimal the ambient pill behaves the same way.
With open mic:
- Press and hold the orb (or the pill) for about half a second. Voice starts and the assistant greets the visitor.
- The visitor speaks; the orb shows listening, processing, thinking and speaking states as the turn progresses.
- Talking over the assistant interrupts it, the same way you would interrupt a person.
- Tapping the orb again stops the voice session. A plain tap while idle opens the text chat instead.
With push to talk:
- Hold the orb (or the pill). Capture starts after a short hold and runs for as long as it is held — there is no greeting, because the visitor is already talking.
- Releasing ends the turn and sends it.
- A quick tap opens the text chat rather than starting voice.
A visitor who is already typing can switch to voice from the call button in the chat header. That button is the keyboard-accessible route into voice — the press-and-hold shortcut is pointer-only, and screen readers are told to open chat and use the call button.
Voice usage draws on your credit balance: transcription is charged per turn and speech is charged per character, at the provider rate. Check Billing and credits before turning voice on for a high-traffic site.
Related
- Appearance — presentation, orb placement and colours.
- Deploying your assistant — putting the widget on your site.
- General settings — name, greeting and status.