Technical

The Knowledge Base That Stays Accurate

Why an AI knowledge base rots, how chunking shapes the answers you get, and when to re-index so your assistant stops quoting the document you replaced.

Rot is the default

Setting up a knowledge base is a one-afternoon job — Building a Knowledge Base Your Assistant Can Actually Use walks through it. Keeping it true is the actual work, and almost nobody plans for it.

The failure is quiet. You upload your policy documents, the assistant answers well, everyone is pleased. Six months later your refund window changed from 30 days to 14, the website was updated, the policy PDF was not, and your assistant has been confidently telling customers 30 days ever since. Nothing broke. No error was logged. The answers stayed fluent and well-sourced right up until one of them became a commitment you did not intend to make.

A stale knowledge base is worse than an empty one. An assistant with no documents says "I do not have that information" and hands off to a person. An assistant with wrong documents answers wrong, with a citation, in your brand voice.

What earns a place in the index

The instinct is to upload everything. Resist it — retrieval quality goes down as you add marginal documents, because every additional document is another candidate competing for the handful of slots the model gets to read.

Start from your conversations, not your file server. Take the twenty questions people actually ask, and upload the documents that answer them. That is usually four or five files, not forty.

Content Upload? Why
Policies: refunds, shipping, warranty, privacy Yes High-stakes, frequently asked, factual
Product specs and compatibility Yes Exactly the lookup a search engine is bad at
Setup guides and troubleshooting steps Yes Deflects the most support volume
FAQ documents Yes, if maintained Otherwise the first thing to rot
Marketing brochures No Teaches the assistant to sell instead of answer
Slide decks No Chunk badly; the meaning was in the presenter
Superseded versions of anything No The single biggest source of wrong answers
Internal drafts, meeting notes No Not accurate, and not for strangers
Anything with customer data No Retrieval has no concept of who should see what
Live data: stock, order status, availability No Use an integration; documents cannot be current

Two of those rows deserve emphasis. Customer data: a knowledge base is a public surface, so if a document is retrievable, assume anything in it can be spoken aloud to a stranger who asks the right question. Redaction by "nobody would ask that" is not redaction. Live data: documents are a snapshot, and no amount of re-uploading makes them current enough to answer "is this in stock." That question needs a live integration or an honest "I cannot check that from here."

What chunking does to your answers

Documents are not searched whole. They are split into passages, each passage is indexed, and a question retrieves the handful of passages closest in meaning. The model sees those passages, not your file.

This has a consequence that changes how you should write: a passage has to make sense on its own. The model never sees the heading three pages up that established you were talking about international orders.

Practical implications:

  • Headings are structure, not decoration. Splitting follows document structure, so descriptive headings produce clean passages. "Returns — International Orders" is a good heading. "Section 4.2" tells the splitter nothing and tells the reader nothing.
  • One topic per section. A section covering shipping times and return windows will get retrieved for shipping questions and drag return information into the model's context, where it becomes a plausible source of an irrelevant sentence.
  • Say the subject in the section. If a paragraph says "the window is 14 days," add what window. Pronouns and implied subjects are fine for a human reading top to bottom and useless in a passage retrieved on its own.
  • Long tables split badly. A 40-row specification table will be cut somewhere, and the half without the header row is meaningless. Break large tables into sections with their own headings, or restate the key facts in prose alongside.
  • Dates in the text, not just in metadata. "Effective 1 March 2026" inside the document lets the model tell the reader how current the information is. A file modification date cannot do that job.

None of this requires rewriting your documentation. It mostly requires adding headings and finishing a few sentences.

Re-indexing without leaving landmines

When content changes, the correct move is to replace the document, not add the new version alongside the old.

This sounds obvious and is violated constantly, usually with good intentions: someone uploads refund-policy-2026.pdf and leaves refund-policy.pdf in place for reference. Both are now indexed. Both will be retrieved for refund questions. The model has no reliable way to tell which is current — filenames are not part of what it reads, and both documents assert their contents with equal confidence. You have not added a version, you have added a coin flip.

The same applies to overlapping documents that are not literally versions. If your FAQ answers the refund question in one sentence and your policy document answers it in a paragraph, and the two disagree slightly, you have built a machine for producing inconsistent answers. Pick one place for each fact.

A maintenance rhythm that actually happens

  1. When a policy changes, replace the document the same day. Attach this to whatever process already updates your website. If it is a separate task, it will not happen.
  2. Once a quarter, list your documents and ask what each is for. Anything you cannot justify in one sentence comes out. This takes fifteen minutes and is the highest-leverage maintenance you will do.
  3. Read the conversations where the assistant abstained or escalated. Those are your gaps, stated by real people in their own words. Repeated escalation on one topic means a missing document, not a difficult visitor — the loop described in human handoff done right.
  4. Spot-check five citations a month. Open the source the assistant cited and confirm the claim is actually in there. This is how you catch rot before a customer does.

Fewer documents, better answers

There is a real cap on how many documents an assistant can carry, and it varies by plan — see hiroi.ai/pricing for the current numbers. It tends to be treated as a constraint. In practice it is a useful forcing function, because the assistants that answer best are almost never the ones with the most uploaded.

The reason is mechanical. The model reads a small number of passages per question. If your best passage is competing with fifteen near-duplicates from an old brochure, a superseded FAQ, and three overlapping guides, it may not make the cut. Removing the noise does not just tidy the list — it directly changes which passage the model reads, and therefore what it says.

If you can only keep a handful of documents, keep the ones that answer questions people actually ask, in their current versions, written so each section stands alone.

The check that keeps you honest

Every claim your assistant makes should be traceable to a document you would be comfortable showing the person who received the answer. That is the whole standard.

It is also the reason to keep citations visible: a badge naming the source turns a maintenance problem into something a visitor can catch for you, and it makes your own spot checks a two-second job. The reasoning behind that is in show your sources, and the mechanics of how retrieval works underneath are in building a knowledge-base agent with RAG.

You can upload, replace, and test documents against real questions in the dashboard at hiroi.ai. Start with the five documents that answer your twenty most common questions — and put a recurring reminder in your calendar for the quarterly cull, because that is the step everybody skips.

Try hiroi free.

Put an AI agent on your site for chat and voice — no credit card required.