Limits and quotas
Every limit the product enforces, in one place: what your plan allows, how long each settings field can be, what the widget accepts from a visitor, and the rate limits on the endpoints your website calls.
Numbers marked (default) can be changed on self-hosted deployments through environment configuration. On hiroi.ai they are the values in force.
Plan limits
Free means the account has never purchased credits. Paid means it has — and it is permanent, so spending your balance back down does not move you back to the free column. See Billing and credits.
| Limit | Free | Paid |
|---|---|---|
| Assistants | 1 | 20 |
| Knowledge base | Not available | 20 documents per assistant |
| Text and voice chat | Yes | Yes |
| Voice provider | Azure only | Any available provider |
| Contacts | Not available | Unlimited |
| Shared workspaces | Not available | 1 workspace, 10 members |
When you hit the assistant limit as a paid account, the dashboard says so rather than offering an upgrade — 20 is the current ceiling for everyone.
Paid-only capabilities, gated by the same signal: the knowledge base, Session Signed widget authentication, hiding the "Powered by" branding, page intelligence, organizations and teams, conversation export, data retention control, advanced analytics, and a custom text-to-speech endpoint.
Note
Shared workspaces are not yet available on hiroi.ai — creating one, inviting people into it and deleting one are all turned off for launch. Every account has its personal workspace. The numbers above are the limits that apply once the feature opens.
Assistant settings
These are the caps on the fields in the assistant editor. Most fields stop you at the limit as you type, and the server rejects any save that exceeds one.
| Field | Tab | Limit |
|---|---|---|
| Assistant name | General | 255 characters |
| Domain | General | 255 characters |
| Instructions | General | 10,000 characters |
| Goal | General | 2,000 characters |
| Email on completion | General | 255 characters |
| Memory | General | 1–50 messages (10 by default) — values above 10 currently have no effect |
| Suggested questions | General | 5 questions, 200 characters each |
| Terms heading | General | 100 characters |
| Terms text | General | 20,000 characters |
| Accept button label | General | 60 characters |
| Rate limit | General | 1–10,000 requests per hour (2,000 by default) |
| Widget Title | Appearance | 100 characters |
| Escalation channels | Channels | 5 channels, name 100 characters, description 500 characters |
| Allowed domains | Deploy | 50 domains, 255 characters each |
The whole settings object for one assistant is capped at 500,000 bytes, which matters only if you have indexed a very large site map.
Uploads
| Upload | Limit |
|---|---|
| Knowledge base document | 10 MB per file |
| Knowledge base file types | PDF, DOCX, TXT, MD, HTML |
| Text indexed per document | 500,000 characters, 2,000 chunks |
| Logo URL | 2,048 characters, https only — there is no logo upload |
File type is checked twice: once against the type your browser declares, and once against
the file's actual contents. A file renamed to .pdf is rejected.
Text past the 500,000-character mark in a single document is not indexed. If you have a large manual, split it into several documents — see Knowledge and capabilities.
Conversations
Limits that apply to a visitor talking to your assistant.
| Limit | Value |
|---|---|
| Message length | 4,000 characters |
| Conversation history sent per request | 100 messages |
| Messages the assistant remembers | 10 — the server caps history at 10 messages, so a Memory setting above 10 has no effect |
| Spoken reply length | 2,000 characters per request, truncated beyond that |
| Speech clip uploaded for transcription | 25 MB |
| Single spoken turn | 30 seconds, then it is transcribed automatically |
Voice sessions have their own ceilings, all enforced on the live connection:
| Limit | Value |
|---|---|
| Turns per voice session | 200 (default) |
| Turns per minute | 30 (default) |
| Concurrent voice connections per visitor IP | 8 (default) |
| Single WebSocket message | 1 MB |
Reaching the per-session turn cap ends the session. Reaching the per-minute cap returns an error for that turn and the visitor can keep talking.
Page tools
Applies to tools your page registers with the widget's JavaScript API. See Page tools.
| Limit | Value |
|---|---|
| Registered tools offered to the assistant | 50 |
| Tool parameter schema | 8,192 characters |
| Tool description | 1,024 characters |
| Tool result returned to the assistant | 4,096 characters |
| Any single string inside a tool result | 2,000 characters |
Page state passed to setContext |
12,000 characters, of which 8,000 reach the assistant |
Anything over a cap is truncated rather than rejected, so a large result still produces a usable answer — it just arrives shortened. Keep results small and structured.
Actions
Applies to every outbound Action, whichever trigger it uses. See Actions.
| Limit | Value |
|---|---|
| Actions per organization | 50 |
| Request timeout | 5 seconds |
| Response body read | 1 MB |
Retries (event trigger) |
3 attempts total, 2s then 4s apart |
Retries (ai and form triggers) |
None — single attempt |
Calls per action per conversation (ai trigger) |
5 |
Form submissions from the widget are additionally bounded per field: field names up to 100 characters, text values up to 5,000 characters, and list values up to 50 items of 500 characters each. Values past those limits are dropped from the submission, except individual list items, which are truncated to 500 characters.
Widget session tokens
Applies to Session Signed authentication, where your backend mints a token for the browser. See Widget authentication.
| Limit | Value |
|---|---|
Token lifetime (ttl_minutes) |
15 minutes by default, clamped to 1–60 |
| Context object | 2 KB of JSON |
user_id |
255 characters |
user_name |
100 characters |
user_email |
255 characters |
Expired tokens are rejected. Mint a fresh one for each page load rather than reusing one across a long session.
Rate limits
Two layers apply to the widget: a per-assistant hourly budget you control, and fixed per-endpoint limits that protect the service.
Your hourly budget
Rate Limit on the General tab sets how many requests one assistant may serve per hour — 2,000 by default, adjustable between 1 and 10,000. The counter starts on the first counted request and clears an hour later; requests over the budget are refused with a rate-limit error until the window rolls over.
Every authentication mode spends the same budget: Domain Safelist requests, server-to-server session creation, and the requests a browser makes with a session token all count against it. A session token is not a way around the limit. (Local and development deployments skip the count, so testing does not burn a real assistant's allowance.) The budget is per assistant rather than per visitor, so raise it before you put an assistant on a high-traffic page.
Per-endpoint limits
Fixed limits on the endpoints the widget calls from a visitor's browser. "Per assistant" means the limit is counted per assistant per visitor IP — each visitor gets their own allowance on that assistant; "per IP" means the limit is counted per visitor IP across all assistants.
| Endpoint | Limit | Counted |
|---|---|---|
/api/widget/init |
300 per minute | Per IP |
/api/widget/chat |
60 per minute | Per assistant |
/api/widget/chat/stream |
60 per minute | Per assistant |
/api/widget/chat/tool-result |
60 per minute | Per assistant |
/api/widget/stt |
60 per minute | Per assistant |
/api/widget/tts |
180 per minute | Per assistant |
/api/widget/feedback |
60 per minute | Per assistant |
/api/widget/satisfaction |
30 per minute | Per assistant |
/api/widget/end |
30 per minute | Per assistant |
/api/widget/escalate |
5 per minute | Per assistant |
/api/widget/user-token |
10 per minute | Per assistant |
/api/widget/session/create |
20 per minute | Per IP |
| Widget form submission | 10 per minute | Per IP |
Note
Escalation is deliberately the tightest of these. Five hand-offs per minute is far more than a real site produces, and it stops a loop in a custom integration from flooding your team's inbox.