The bug your test suite cannot see
You edit one line of your assistant's system prompt to fix how it handles pricing questions. Every test passes. You ship.
Three weeks later you notice the assistant has stopped asking for an email address before promising a follow-up. It never threw an error. Every endpoint returned 200. The response was well-formed, fluent, and helpful-sounding. It was just wrong in a way no assertion in your codebase was looking for.
This is the defining problem of testing AI products. Your existing tests verify that the machinery works — the request routes, the retrieval query runs, the response serialises. None of them verify the only thing your users experience, which is what the assistant actually said.
Two kinds of test, doing different jobs
| Unit and integration tests | Behavioural evals | |
|---|---|---|
| What they check | The code did what the code should do | The assistant said what it should say |
| Determinism | Same input, same output | Same input, varying output |
| Cost per run | Effectively free | Real money — every case is a live model call |
| Speed | Seconds | Minutes |
| When they run | Every commit | Nightly, and before prompt changes |
| What they catch | Broken plumbing | Broken behaviour |
You need both, and the mistake is trying to make one do the other's job. People attempt to unit-test behaviour with string matching against a canned response, which tests nothing except that the model is deterministic, which it is not. Or they skip behaviour testing entirely on the grounds that it is unreliable — and then ship prompt regressions for months without noticing.
The two run on different schedules for a good reason. Evals cost money and take minutes, so putting them on every commit is both slow and expensive. Nightly is the right default, with a manual run before any change to prompts, tools, or model version.
What a case looks like
A behavioural eval case is a scenario plus a set of expectations about the reply.
The scenario is a realistic conversation — not a single clever prompt, but the two or three turns of context that produce the situation you care about. The expectations come in a few flavours, and choosing the right one matters more than writing more cases:
- Must contain — the reply has to mention something specific. Good for factual grounding: an answer about business hours should include the actual hours.
- Must abstain — the reply must decline rather than invent. This is the workhorse for anything the assistant should not know.
- Must call a tool — a question that needs a document search should trigger one. Behaviour includes what the assistant did, not only what it said.
- Judge rubric — a second model grades the reply against written criteria. Slower and fuzzier, but the only option for "did this sound like a professional, not a salesperson."
Cases are worth tiering by priority. A small set of high-priority cases covering the behaviours that would embarrass you can run often; the long tail runs nightly. If your whole suite is one undifferentiated list, you will eventually stop running it.
Four traps we walked into
These cost us real time. They are all obvious in hindsight and none of them were obvious at the time.
Never assert "must not contain" on a phrase a correct answer would echo
Say you have a case checking that the assistant declines to give medical advice, written as: the reply must not contain "diagnosis." The correct refusal is "I cannot offer a diagnosis — you would want to speak to a clinician." It contains the word. The test now fails on the right answer, and the obvious fix — editing the prompt until the phrase disappears — makes the assistant worse, because it starts refusing vaguely instead of clearly. A correct refusal almost always names the thing it is refusing. Use an abstention check and a judge rubric, never a negative substring.
Normalise before you match
A model will write "don't" with a curly apostrophe, "U.S." with periods, and, in French, "2 400 $" with a space and a trailing symbol. All three are correct and all three fail a naive comparison. Your matching helper should fold punctuation variants, currency formats, and unicode quotes before comparing, or you will spend your evenings debugging tests that were right.
A blocked prompt is sometimes a pass
For adversarial cases, the safety layer in front of the model may reject the input before it ever reaches your code, which surfaces as an error rather than a graded reply. Treating that as a failure is backwards — the defence worked. Those cases need an explicit "a provider block is an acceptable outcome" flag, or your suite will punish you for being safe.
Graders drift away from prompts
This is the expensive one. If your prompt states a rule and a grader checks that rule, the two are a pair. Change the prompt without changing the grader and the eval keeps passing while enforcing a rule you no longer have — a test that is green and meaningless, which is worse than no test at all because it buys false confidence. Keep them in the same review, and treat "edited a prompt, did not touch a grader" as a smell.
The workflow that makes it useful
Evals earn their cost when they change how you edit prompts. The loop:
- Write the case before the fix. Something went wrong in a real conversation. Turn it into a case. Watch it fail.
- Change the prompt. One change.
- Run the affected suite. Not the whole thing — the subset that covers the behaviour you touched, plus the high-priority set.
- Read the failures, not just the count. A score is a summary. The actual reply text is where you find out that the model technically passed for the wrong reason.
- Let the full suite run overnight. The regressions you did not predict are the ones worth catching.
That third step is the one people skip and the one that changes the economics. If you run everything on every edit, evals become a thing you do twice and then avoid. If you can run twelve relevant cases in ninety seconds, they become the thing you do while editing.
Step four deserves emphasis. Behavioural evals produce a number, and the number is the least useful output. An assistant that scores 94% by giving confident wrong answers to the 6% is a worse product than one scoring 88% that abstains cleanly. Read the transcripts.
What this catches that nothing else does
Three classes of problem, all of which are easy to ship without noticing:
Instruction erosion. Prompts grow. Each addition slightly dilutes the ones before it, and eventually a rule that was reliably followed at 400 words is followed 70% of the time at 900. There is no error and no commit to blame, just a slow decline. Only repeated behavioural sampling finds it.
Silent tool loss. A tool description gets reworded, the model stops recognising when to use it, and the assistant starts answering from general knowledge instead of your documents. Everything still works. The answers are just no longer grounded — which is exactly the failure that makes answer citations worth having on the visitor side too.
Refusal collapse. The hardest to notice. Abstention is the first behaviour to break when a prompt is edited for helpfulness, and the resulting answers look better than the correct ones. They are fluent, confident, and made up. Without a case explicitly testing that the assistant declines what it does not know, this regression is invisible until a customer quotes it back to you.
If you are writing or rewriting the prompt those evals grade, system prompt engineering covers the construction side, and chatbot mistakes to avoid covers the behaviours worth writing cases against first.
Every assistant on hiroi.ai runs behind the same prompt and tool machinery this suite grades. The starting point for your own version is easier than it sounds: take the last three conversations that went wrong, and write one case each.