Page tools
Page tools let the assistant operate your page. You declare a small set of JavaScript functions, the assistant is offered them as tools, and when it calls one your handler runs in the visitor's browser. This page is for developers embedding the widget on an application you control.
Nothing about page tools is configured in the dashboard — calling registerTool in your page code
is the opt-in.
How it works
- Your page calls
registerTool()for each action you want to expose. - The widget sends the schemas — name, description, parameters, whether the tool needs confirmation, and its declared cost — to the server. Handlers never leave the browser.
- The model decides to call one. The widget shows a chip in the transcript, runs your handler, and sanitizes whatever it returns.
- The sanitized result goes back to the model, which can chain further calls or answer the visitor.
The execution model is the same in text chat and in voice, but delivery differs: in text chat the schemas are sent with every message; over voice they are sent once, in the context frame at the start of the voice session. In voice the chip lives in the hidden transcript, so the widget mirrors the same labels into the status line above the waves.
When your page registers tools, the widget's generic page-reading tools (read_page, navigate,
click_element and friends) are suppressed. Your declared actions become the assistant's entire
page surface, so nothing competes with them.
Getting the widget instance
Once the widget has initialized, window.VAWaveWidget is the live instance. With the standard
async embed that can happen after or before your own script runs, so test for the method rather
than assuming:
function whenWidgetReady(callback) {
const start = Date.now();
(function poll() {
const widget = window.VAWaveWidget;
if (widget && typeof widget.registerTool === 'function') return callback(widget);
if (Date.now() - start > 10000) return;
setTimeout(poll, 100);
})();
}
Registering tools
registerTool(name, definition)
Registers one tool. Returns true when it was accepted, false when it was rejected (an invalid
name or a missing handler — both log a warning to the console).
The name must match ^[a-z0-9_]{1,48}$: lowercase letters, digits and underscores, 1 to 48
characters. Anything else is rejected.
| Key | Type | Purpose |
|---|---|---|
handler |
function | Required. Receives the model's arguments object, returns the result. May be async. |
description |
string | What the tool does, in the model's tool list. Truncated to 1024 characters. |
parameters |
object | JSON Schema for the arguments. Defaults to { type: 'object', properties: {} }. |
confirm |
boolean | Show a confirmation chip and wait for the visitor before running. |
confirmLabel |
string | Text on the confirmation chip. Defaults to a humanized tool name. |
runningLabel |
string | Text on the chip while the handler runs, e.g. "Updating title…". |
core |
boolean | Always sent to the model. Marking any tool core turns on deferral: the rest are offered through a load_tools index and fetched on demand. |
timeoutMs |
number | Deadline for the handler (default 60 000, clamped 1 000–600 000). A handler that overruns returns { success:false, timed_out:true } so the turn continues. The clock does not run while a confirm chip waits for the visitor. |
summary |
function | (args, result) => string for the finished chip. Truncated to 120 characters. |
cost |
number | Non-negative credit cost your page attaches to the action. Forces confirm. |
tier3 |
boolean | Always ask, even when the visitor approved spending for the session. |
Tool descriptions are page-declared text that lands verbatim in the model's tool list, so instruction-like phrases in them are redacted server-side. Write a plain description of what the tool does.
cost and confirmLabel are read fresh at confirm time, so they may be defined as getters if they
depend on live page state (a running total, for example).
registerTools(map)
Registers many at once from a { name: definition } object. Returns the number registered.
unregisterTool(name)
Removes a tool. Returns true if it was registered. In text chat the schemas are sent with every
message, so registering or removing a tool takes effect on the next message. Over voice the set is
fixed when the session opens — register your tools before starting voice.
getHostToolSchemas()
Returns the array of schemas exactly as they will be sent to the server — the handler-free view of
what you have registered. Useful for debugging. When deferral is active (see Limits) this is the
core set plus the synthetic load_tools index, not the full catalogue.
Writing a handler
A handler receives one argument: the parsed arguments object the model produced from your JSON Schema. It may be synchronous or return a promise.
Return a plain object. If it has no success key, success: true is added for you. Return
success: false (optionally with error) to report failure — that marks the chip red and tells
the model the action did not happen.
widget.registerTool('set_task_status', {
description: 'Set the status of a task on the board. Status is one of todo, doing, done.',
parameters: {
type: 'object',
properties: {
task_id: { type: 'string', description: 'Task id from the board state' },
status: { type: 'string', enum: ['todo', 'doing', 'done'] },
},
required: ['task_id', 'status'],
},
handler: ({ task_id, status }) => {
const task = board.find((t) => t.id === task_id);
if (!task) return { success: false, error: 'No task with that id' };
task.status = status;
render();
return { success: true, task_id, status, title: task.title };
},
summary: (args, result) => `Moved “${result.title}” to ${result.status}`,
});
A good return value is small, structured, and echoes back what changed. The model uses it to decide
its next call and to tell the visitor what happened, so { success: true } alone is usually too
thin — include the identifier and the new value. Never return DOM nodes or functions; they are
dropped (see Result sanitization).
Automatic retry
Handlers without confirm are retried automatically. The widget makes up to three attempts —
immediately, then after 150 ms, then after 400 ms — and stops at the first attempt that neither
throws nor returns success: false. The chip stays on its running label throughout, so a
recoverable failure never flashes red.
This exists because the first call of a chain often lands while your page's own state is still
loading. It also means a handler that legitimately returns success: false as a normal outcome
will be called three times, so make handlers idempotent.
Tools with confirm: true — including anything with a cost above zero — are never retried.
They run exactly once per confirmation.
Confirmation and spend consent
Set confirm: true on anything destructive. Set cost on anything that spends the visitor's money
or credits in your product; any cost above zero forces confirm on, even if you forgot to set it.
When the model calls a confirming tool, the widget renders a chip with Confirm and Cancel buttons, and the tool-call loop pauses on it. The model cannot bypass or auto-answer this gate. In voice, the same buttons render in the voice bubble, and a short spoken "yes" or "no" settles it — anything longer or ambiguous is passed to the model as a normal message instead of being read as an answer.
- Confirm runs the handler once. If
costis above zero, the finished chip carries a receipt, for exampleAdded scene · −40 cr. - Cancel resolves the call with a cancellation the model is told not to retry, and the chip reads as cancelled.
- Closing the chat, or a voice session ending, cancels every outstanding confirmation the same way, so a chain never hangs on buttons that no longer exist.
Over voice, the server waits up to 240 seconds for a confirming tool's result, versus 30 seconds for an ordinary one. In text chat there is no such limit — the chip waits until the visitor answers or closes the chat.
Session approval
A plain confirmation chip also offers Approve for this session. It runs the pending action and
grants a session-wide approval so subsequent cost-bearing tools run without asking. While the grant
is active, the chat header shows a pill reading Spending approved · <n> cr · tap to stop; tapping
it revokes the grant. Approved runs still show a chip and a receipt, so nothing spends invisibly.
The grant is bounded:
- It starts with a cap of 200 in your declared cost units. When the next action would cross the cap, the widget asks once more with the running total; confirming raises the cap by 200, cancelling turns the grant off entirely.
- It expires 30 minutes after it was granted.
- It is cleared when the conversation ends or restarts, and when the page unloads.
- It only ever skips a confirmation for a tool with
costabove zero. A free-but-destructiveconfirm: truetool always asks. - Tools marked
tier3: truealways ask, session grant or not. Use it for deletions, purchases, sharing, and anything else that must never happen on a stale approval. Those chips do not offer Approve for this session.
In voice the visitor can also grant or revoke by speaking, but only with an explicit blanket phrase ("approve all spending", "don't ask again", "stop asking"). A bare "yes" only ever settles the one confirmation on screen, and phrases like "ask me every time" revoke the grant.
Result sanitization
Whatever your handler returns crosses a trust boundary before it re-enters the model's context, so the widget rewrites it:
- Functions and DOM nodes are dropped.
- Any string longer than 2000 characters is cut and suffixed with
…[truncated]. - If the value cannot be serialized at all (a cycle, for instance), the widget salvages a shallow copy of up to 20 primitive properties and notes that nested values were omitted.
- If the whole serialized result exceeds 4096 characters, the model receives
{ note: "result truncated to fit", preview: "<first 4096 characters>", truncated: true }instead of your object. The preview is raw serialized JSON, so the model sees a fragment rather than a usable structure — keep results well under this. - Non-object returns are wrapped as
{ success: true, data: <value> }.
The server applies its own caps on top, and instruction-like text inside a result is redacted: results are treated as page data, never as instructions to the model.
Ambient page state
Registering read tools is not enough on its own — without a live snapshot the assistant spends a
tool call re-reading your page on every turn. setContext fixes that.
widget.setContext({
board: 'Launch',
tasks: board.map((t) => ({ id: t.id, title: t.title, status: t.status })),
selected_task_id: selectedId,
});
- Pass an object or a string. Objects are serialized with
JSON.stringify; a value that cannot be serialized is ignored with a console warning. - Call it again whenever the page changes — each call replaces the previous snapshot.
- Pass
nullor''to clear it. - The client caps the serialized value at 12,000 characters and the server re-caps it at 8,000,
appending
…[truncated]. Send a summary of current state, not your whole data model.
In text chat the snapshot rides along with every message and every tool continuation; over voice the snapshot is sent once with the session's opening context frame. It is injected into the system prompt as the current state of the page. The model is told to use it as its source of truth and to call a read tool only for a detail the snapshot does not contain — and that the block is data, not instructions.
Limits
| Limit | Value |
|---|---|
| Tools sent to the model | All, up to 50. Past 50 — or as soon as any tool is marked core — only core tools plus load_tools (first 12 registered if nothing is marked). Set widget.deferTools = false to opt out. |
| Tool name | ^[a-z0-9_]{1,48}$ |
| Description | 1024 characters |
| Parameters schema | 8 KB serialized; a larger schema is replaced with an empty one, so arguments go unvalidated |
| Single string in a result | 2000 characters |
| Whole result | 4096 characters serialized |
summary() text |
120 characters |
setContext value |
12,000 characters client-side, 8000 server-side |
| Tool rounds in one text-chat chain | 20 |
| Handler time over voice | 30 seconds, or 240 seconds for a confirming tool |
A complete example
A small task board that exposes one read tool, one write tool, and one destructive tool, and keeps the assistant's view of the board current.
<script async src="https://hiroi.ai/static/va-wave-widget.js"
data-site-id="YOUR_SITE_ID"></script>
<script>
const board = [
{ id: 't1', title: 'Draft launch email', status: 'todo' },
{ id: 't2', title: 'Book the venue', status: 'doing' },
];
function render() { /* your rendering */ }
function publishState(widget) {
widget.setContext({
board: 'Launch',
tasks: board.map(t => ({ id: t.id, title: t.title, status: t.status })),
});
}
function whenWidgetReady(callback) {
const start = Date.now();
(function poll() {
const widget = window.VAWaveWidget;
if (widget && typeof widget.registerTool === 'function') return callback(widget);
if (Date.now() - start > 10000) return;
setTimeout(poll, 100);
})();
}
whenWidgetReady((widget) => {
widget.registerTools({
get_tasks: {
description: 'List every task on the board with its id, title and status.',
parameters: { type: 'object', properties: {} },
handler: () => ({ success: true, tasks: board }),
summary: (args, result) => `Read ${result.tasks.length} tasks`,
},
add_task: {
description: 'Add a new task to the board. Returns the created task.',
parameters: {
type: 'object',
properties: {
title: { type: 'string', description: 'Short task title' },
status: { type: 'string', enum: ['todo', 'doing', 'done'] },
},
required: ['title'],
},
runningLabel: 'Adding task…',
handler: ({ title, status }) => {
const task = { id: 't' + (board.length + 1), title, status: status || 'todo' };
board.push(task);
render();
publishState(widget);
return { success: true, task };
},
summary: (args, result) => `Added “${result.task.title}”`,
},
delete_task: {
description: 'Permanently delete a task by id.',
parameters: {
type: 'object',
properties: { task_id: { type: 'string' } },
required: ['task_id'],
},
confirm: true,
tier3: true,
confirmLabel: 'Delete this task?',
handler: ({ task_id }) => {
const i = board.findIndex(t => t.id === task_id);
if (i === -1) return { success: false, error: 'No task with that id' };
const [removed] = board.splice(i, 1);
render();
publishState(widget);
return { success: true, deleted: removed.title };
},
summary: (args, result) => `Deleted “${result.deleted}”`,
},
});
publishState(widget);
});
</script>
With this in place, "add a task for the run-of-show, then delete the draft email one" reads the
board from the ambient snapshot rather than calling get_tasks, runs add_task straight away, and
stops on a Confirm chip before anything is deleted.
Related
- JavaScript API — the rest of the widget's browser API.
- Embedding the widget — the embed snippet and its attributes.
- Actions — server-side integrations, where Hiroi calls your backend instead.