General settings
The General tab is the first tab in the assistant editor. It holds the settings that decide who your assistant is and how it thinks: its name, its instructions, the goal it works toward, the model behind it, how much conversation it remembers, and what languages it answers in.
Open it from Assistants, click the assistant you want to edit, and stay on the General tab.

Before you start: how saving works
Nothing on this tab takes effect until you save.
- Edits are held locally. As soon as something changes, an Unsaved changes marker appears in the header next to the Save button.
- Click Save — or press
Cmd+S/Ctrl+S— to write everything on every tab in one request. - If you try to leave the page with unsaved changes, your browser warns you first.
- Preview opens the last-saved version of the assistant in a new tab. Save before previewing if you want to see your edits.
- A save can come back as Saved with warnings. The settings were stored, but something was adjusted or flagged — read the message before moving on.
Basic Information
| Control | What it does |
|---|---|
| Assistant Name | The name of the assistant inside the dashboard. It is also passed to the AI-generation helper as context. |
| Domain | The domain where this assistant will be embedded, e.g. example.com. Hidden for personal-assistant type assistants, which are not embedded on a website. |
| Assistant Status | Enables or disables the assistant on your website. The badge next to it reads Active or Disabled, and the editor header shows Live or Off. |
Turning Assistant Status off is the fastest way to take an assistant down without deleting it or removing the embed code from your site.
Domain is not the security boundary
The Domain field here is descriptive. The list of origins actually allowed to load the widget is configured separately, on the Deploy tab.
AI Behavior
Instructions
Instructions is the assistant's system prompt — the single most important setting on this tab. It defines how the assistant behaves. The field's own guidance says it best: be specific about tone, style, and what topics to focus on.
- The limit is 10,000 characters. A live counter under the field shows
used / 10,000, and turns amber once you pass 8,000 characters. - The server enforces the same 10,000-character limit, so a longer prompt is rejected on save.
- The platform's built-in default prompt is used only for an assistant that has never had Instructions set. Clearing the field on an assistant that already had a prompt leaves it with no instructions of its own, not the built-in default.
Writing instructions that work
A few habits that make a measurable difference:
Lead with the role and the audience. The first line sets the frame for everything after it: "You are the front-desk assistant for Bayside Dental. You talk to patients and prospective patients."
Describe behaviour, not personality adjectives. "Friendly" is vague. "Greet by name when you know it, answer in two or three sentences, and offer to book an appointment once the question is answered" is something the model can actually follow.
Say what is out of scope. Most bad answers come from questions the assistant should have deflected. Name them: "Do not quote prices, discuss insurance coverage, or give medical advice. For any of those, offer to take a message."
Set the length and format. Chat replies and spoken replies both read from this field, so avoid formatting that only makes sense on screen if voice is enabled. "Keep answers under 60 words unless asked for detail" is a good default.
Put facts in the knowledge base, not the prompt. Opening hours, prices, and policies change. Uploading documents on the Intelligence tab keeps the prompt short and the facts current — an instruction file that has drifted out of date is worse than no instruction at all.
Edit against real transcripts. After the assistant has been live for a day, read the conversations, find the answer you disliked, and add one line to the prompt that would have prevented it. Repeat. That loop beats writing 2,000 characters up front.
Generating instructions with AI
The AI button in the section header opens a small panel that writes a first draft for you.
- Click AI.
- Describe the assistant in the field — up to 500 characters, e.g. "friendly support assistant for a SaaS product". Your assistant's name is sent along as context.
- Click Generate (or press
Enter). Cancel closes the panel.
The generated text replaces whatever is currently in Instructions, so copy anything you want to keep first. Generation is rate-limited to 10 requests per hour per account; past that you'll see Rate limit reached — try again later. The result is a draft — read it, cut what doesn't apply, and add your specifics.
Outcomes
This section is what turns conversations into a measurable result.
Goal
Goal is a plain-English description of what success looks like, up to 2,000 characters. For example: "You succeed when the customer books a demo or shares their email."
It does two things:
- It steers the assistant. The goal is appended to the system prompt for every turn as an objective the assistant should work toward.
- It enables resolution grading. When a conversation or call ends, the transcript is graded
against this exact text by an LLM judge, which records a status (
achieved/not_met) and a grade from 1 to 5 with a short written reason.
No goal means no grading
Grading is skipped entirely when Goal is empty. That is why the resolution and self-served figures in Analytics stay blank for assistants that have never had a goal set — ungraded conversations simply aren't counted. Setting a goal here is what switches those metrics on, going forward.
Write goals that a reader could verify from a transcript. "Be helpful" cannot be graded. "The visitor gets a working answer to their billing question, or is handed to a human" can.
Email on completion
Email on completion is optional. When an address is set, it receives a one-line outcome summary after each ended conversation — the goal, the status, the grade, and the judge's reasoning.
Systems that need the full payload should subscribe to the conversation.resolved event (and
call.resolved for phone calls) as an Action instead of relying on email.
AI Model
Model
Model picks the LLM behind every channel — chat, voice, and phone calls. More capable models cost more credits per message.
| Option | Rate |
|---|---|
| Default (fastest, lowest cost) | Whatever the default model's own rate is — see below. |
| Claude Opus 4.6 — Most capable | 2.5× |
| Claude Sonnet 4.6 — Balanced | 1.5× |
| Claude Haiku 4.5 — Fast, affordable | 0.5× |
| Nova Lite — Fastest, lowest cost | 0.05× |
| Nova Micro — Cheapest | 0.05× |
| gpt-oss 120B — Open weights, very fast | 0.1× |
| GPT-5.4 Mini — Newest, fast | 0.5× |
| GPT-4.1 — Most capable | 2.5× |
| GPT-4.1 Mini — Fast, affordable | 0.5× |
| GPT-4.1 Nano — Cheapest | 0.1× |
| GPT-4o — Balanced | 3.0× |
Credits are charged on total tokens (input plus output) multiplied by the model's rate, with a minimum of one credit for any non-zero usage. The practical consequence: a chatty, high-volume support assistant on Nova Lite costs a fraction of the same traffic on GPT-4o.
Default is not a rate of its own
Default leaves the setting empty and runs whichever model the deployment is configured to use. You are billed at that model's rate from the table above — never at a flat 1×. On the hiroi.ai cloud the configured default is currently GPT-4.1 Nano, so Default and GPT-4.1 Nano bill identically today. Self-hosted deployments set their own default, and it can be changed, so treat the row above as "whatever is running", not as a fixed price.
Temperature
Temperature is a slider from Precise (0) to Creative (1) in steps of 0.1. It controls response variety — lower values give consistent, predictable answers; higher values give more creative, varied ones.
Support and booking assistants generally want the low end. Leave it alone unless answers feel either robotically identical or unpredictably off-script.
Admin only in enterprise deployments
In enterprise (self-hosted) deployments, Model and Temperature are cost-affecting settings restricted to org admins. Non-admins see an Admin only badge on this section and both controls are disabled. On hiroi.ai they are available to everyone.
Chat
| Control | Default | What it does |
|---|---|---|
| Typing Indicator | On | Shows the animation in the widget while the AI is composing a reply. |
| Memory | 10 messages | How many of the most recent messages are sent back to the AI as context. Range 1–50. |
| Suggested Questions | none | Clickable suggestions shown after the welcome message. |
Memory is a direct cost/quality trade. More messages means the assistant keeps track of what was said earlier in a long conversation, but every one of those messages is re-sent — and re-charged — on each turn. Ten is a sane default; raise it only for assistants that hold long, winding conversations.
Suggested Questions accepts up to 5 entries of up to 200 characters each. Use + Add suggestion to add one and the × next to a row to remove it; once you have five, the button is replaced by a Maximum 5 suggestions note. Good suggestions are the questions people actually ask in the first ten seconds — they set expectations about what the assistant can do far better than a welcome message does.
Language
Response Language Mode
| Mode | Behaviour |
|---|---|
| Auto-detect & respond in kind | The assistant detects the visitor's language and replies in that language. This is the default. |
| Fixed language only | The assistant always replies in Primary Language, regardless of the language it was written to in. |
| No restriction | The assistant replies in whatever language it detects, with no restrictions. |
Detection runs on each visitor message and the resolved language is stored on the conversation. In auto mode the signals are ranked: a confident detection from the current message wins, then the language the conversation has already settled into, then the browser locale. That ordering is what stops a short reply like "ok gracias" from flipping an established Spanish conversation back to English.
Primary Language
Primary Language is shown for the Auto-detect and Fixed modes.
- In Fixed mode it is the language the bot will always respond in.
- In Auto-detect mode it is the fallback language when detection is uncertain.
The picker offers English (the default), Spanish, French, German, Italian, Portuguese, Japanese, Korean, Chinese, Arabic, Hindi, Polish, Dutch, Russian, Swedish and Turkish.
Greetings by Language
Greetings by Language holds custom welcome messages shown to visitors based on their browser language. With none configured, the panel says so and the default welcome message is used everywhere.
To add one, pick a language from the select and click + Add greeting, then type the greeting text. The × next to a row removes it.
The widget matches the visitor's browser locale first (es-MX), then the base language (es), and
falls back to the standard welcome message when neither is configured. Only the greeting is swapped
— the rest of the conversation still follows the mode above.
Terms & Conditions
Require visitors to accept terms gates the assistant behind a terms panel: visitors must accept before they can chat or speak with it. Acceptance is remembered per device, and editing the text re-prompts everyone — the stored acceptance is bound to a hash of the terms text, so any change invalidates it.
Turning the toggle on reveals three fields:
| Field | Limit | Notes |
|---|---|---|
| Heading | 100 characters | Shown at the top of the panel. Defaults to "Terms & Conditions" when blank. |
| Terms text | 20,000 characters | Plain text. Line breaks are preserved; Markdown and HTML are not rendered. A live counter shows how much you've used. |
| Accept button label | 60 characters | Defaults to "I Accept". A live preview of the button label is shown underneath. |
If the toggle is on and Terms text is empty, the editor shows a warning and the save is rejected — the gate refuses to go live with nothing to accept.
The gate applies to both text chat and voice: a visitor who hasn't accepted gets a terms_required
error instead of a reply.
Access & Limits
Rate Limit (requests/hour) caps how many widget requests this assistant will serve per hour. Accepted range is 1–10000; the default is 2,000.
The counter is kept on the assistant. The first request starts a one-hour window; when that hour elapses the counter resets to zero and a new window begins. Once the cap is hit, further widget requests are refused with a rate-limit error until the window rolls over. It is a blunt instrument — it protects a single assistant from runaway traffic and from credit burn caused by abuse, and it applies to the assistant as a whole rather than to any one visitor.
Raise it if a busy, legitimately popular assistant starts refusing visitors. Lower it if you are embedding on a public page and want a hard ceiling on spend.