Documentation

Limits and quotas

Every limit the product enforces, in one place: what your plan allows, how long each settings field can be, what the widget accepts from a visitor, and the rate limits on the endpoints your website calls.

Numbers marked (default) can be changed on self-hosted deployments through environment configuration. On hiroi.ai they are the values in force.

Plan limits

Free means the account has never purchased credits. Paid means it has — and it is permanent, so spending your balance back down does not move you back to the free column. See Billing and credits.

Limit Free Paid
Assistants 1 20
Knowledge base Not available 20 documents per assistant
Text and voice chat Yes Yes
Voice provider Azure only Any available provider
Contacts Not available Unlimited
Shared workspaces Not available 1 workspace, 10 members

When you hit the assistant limit as a paid account, the dashboard says so rather than offering an upgrade — 20 is the current ceiling for everyone.

Paid-only capabilities, gated by the same signal: the knowledge base, Session Signed widget authentication, hiding the "Powered by" branding, page intelligence, organizations and teams, conversation export, data retention control, advanced analytics, and a custom text-to-speech endpoint.

Note

Shared workspaces are not yet available on hiroi.ai — creating one, inviting people into it and deleting one are all turned off for launch. Every account has its personal workspace. The numbers above are the limits that apply once the feature opens.

Assistant settings

These are the caps on the fields in the assistant editor. Most fields stop you at the limit as you type, and the server rejects any save that exceeds one.

Field Tab Limit
Assistant name General 255 characters
Domain General 255 characters
Instructions General 10,000 characters
Goal General 2,000 characters
Email on completion General 255 characters
Memory General 1–50 messages (10 by default) — values above 10 currently have no effect
Suggested questions General 5 questions, 200 characters each
Terms heading General 100 characters
Terms text General 20,000 characters
Accept button label General 60 characters
Rate limit General 1–10,000 requests per hour (2,000 by default)
Widget Title Appearance 100 characters
Escalation channels Channels 5 channels, name 100 characters, description 500 characters
Allowed domains Deploy 50 domains, 255 characters each

The whole settings object for one assistant is capped at 500,000 bytes, which matters only if you have indexed a very large site map.

Uploads

Upload Limit
Knowledge base document 10 MB per file
Knowledge base file types PDF, DOCX, TXT, MD, HTML
Text indexed per document 500,000 characters, 2,000 chunks
Logo URL 2,048 characters, https only — there is no logo upload

File type is checked twice: once against the type your browser declares, and once against the file's actual contents. A file renamed to .pdf is rejected.

Text past the 500,000-character mark in a single document is not indexed. If you have a large manual, split it into several documents — see Knowledge and capabilities.

Conversations

Limits that apply to a visitor talking to your assistant.

Limit Value
Message length 4,000 characters
Conversation history sent per request 100 messages
Messages the assistant remembers 10 — the server caps history at 10 messages, so a Memory setting above 10 has no effect
Spoken reply length 2,000 characters per request, truncated beyond that
Speech clip uploaded for transcription 25 MB
Single spoken turn 30 seconds, then it is transcribed automatically

Voice sessions have their own ceilings, all enforced on the live connection:

Limit Value
Turns per voice session 200 (default)
Turns per minute 30 (default)
Concurrent voice connections per visitor IP 8 (default)
Single WebSocket message 1 MB

Reaching the per-session turn cap ends the session. Reaching the per-minute cap returns an error for that turn and the visitor can keep talking.

Page tools

Applies to tools your page registers with the widget's JavaScript API. See Page tools.

Limit Value
Registered tools offered to the assistant 50
Tool parameter schema 8,192 characters
Tool description 1,024 characters
Tool result returned to the assistant 4,096 characters
Any single string inside a tool result 2,000 characters
Page state passed to setContext 12,000 characters, of which 8,000 reach the assistant

Anything over a cap is truncated rather than rejected, so a large result still produces a usable answer — it just arrives shortened. Keep results small and structured.

Actions

Applies to every outbound Action, whichever trigger it uses. See Actions.

Limit Value
Actions per organization 50
Request timeout 5 seconds
Response body read 1 MB
Retries (event trigger) 3 attempts total, 2s then 4s apart
Retries (ai and form triggers) None — single attempt
Calls per action per conversation (ai trigger) 5

Form submissions from the widget are additionally bounded per field: field names up to 100 characters, text values up to 5,000 characters, and list values up to 50 items of 500 characters each. Values past those limits are dropped from the submission, except individual list items, which are truncated to 500 characters.

Widget session tokens

Applies to Session Signed authentication, where your backend mints a token for the browser. See Widget authentication.

Limit Value
Token lifetime (ttl_minutes) 15 minutes by default, clamped to 1–60
Context object 2 KB of JSON
user_id 255 characters
user_name 100 characters
user_email 255 characters

Expired tokens are rejected. Mint a fresh one for each page load rather than reusing one across a long session.

Rate limits

Two layers apply to the widget: a per-assistant hourly budget you control, and fixed per-endpoint limits that protect the service.

Your hourly budget

Rate Limit on the General tab sets how many requests one assistant may serve per hour — 2,000 by default, adjustable between 1 and 10,000. The counter starts on the first counted request and clears an hour later; requests over the budget are refused with a rate-limit error until the window rolls over.

Every authentication mode spends the same budget: Domain Safelist requests, server-to-server session creation, and the requests a browser makes with a session token all count against it. A session token is not a way around the limit. (Local and development deployments skip the count, so testing does not burn a real assistant's allowance.) The budget is per assistant rather than per visitor, so raise it before you put an assistant on a high-traffic page.

Per-endpoint limits

Fixed limits on the endpoints the widget calls from a visitor's browser. "Per assistant" means the limit is counted per assistant per visitor IP — each visitor gets their own allowance on that assistant; "per IP" means the limit is counted per visitor IP across all assistants.

Endpoint Limit Counted
/api/widget/init 300 per minute Per IP
/api/widget/chat 60 per minute Per assistant
/api/widget/chat/stream 60 per minute Per assistant
/api/widget/chat/tool-result 60 per minute Per assistant
/api/widget/stt 60 per minute Per assistant
/api/widget/tts 180 per minute Per assistant
/api/widget/feedback 60 per minute Per assistant
/api/widget/satisfaction 30 per minute Per assistant
/api/widget/end 30 per minute Per assistant
/api/widget/escalate 5 per minute Per assistant
/api/widget/user-token 10 per minute Per assistant
/api/widget/session/create 20 per minute Per IP
Widget form submission 10 per minute Per IP

Note

Escalation is deliberately the tightest of these. Five hand-offs per minute is far more than a real site produces, and it stops a loop in a custom integration from flooding your team's inbox.