The knowledge base is the part of an assistant most people get wrong in the same direction: they upload everything they have, get vague answers, and conclude the retrieval is bad. Usually the retrieval is fine. It was handed 400 pages of overlapping, half-superseded prose and asked to find the current refund window in it.
This is a practical guide to what to upload, what to leave out, and how to shape a document so the passage that comes back is the one that answers the question. For the mechanics of how retrieval works underneath, read Build a Knowledge-Base Agent with RAG in Minutes.
Where it lives
Open your assistant and go to the Intelligence tab. Knowledge Base is the first card. Turn the switch on, then expand the card to manage documents.
Drop files onto the upload area or click it to browse. You can select several at once.
| Constraint | Value |
|---|---|
| File types | PDF, DOCX, TXT, MD, HTML |
| Size | 10 MB per file |
| Text indexed per document | 500,000 characters |
Files are checked against their real contents, not just the extension, so a .docx renamed to .pdf is rejected. Each assistant also has a cap on how many documents it can hold — a small number on a free account, considerably more once the account is paid. Current figures are on the pricing page.
Watch the status column
Every uploaded file appears in a table with a Status:
| Status | Meaning |
|---|---|
pending |
Uploaded, waiting to be processed |
processing |
Text is being extracted and indexed |
indexed |
Ready — the assistant can retrieve from it |
error |
Processing failed; the file is not searchable |
Only documents at indexed are searched. A row sitting at error is ignored entirely, silently, forever. Indexing runs in the background, so reopen the card or reload the page to see the status settle. Check it once after every upload — this is the single most common reason an assistant "ignores" a document you are certain you gave it.
There is no in-place edit. To update a document, delete the old row and upload the new version.
What to upload
Start from questions, not from documents. Open your support inbox or your last month of conversations, write down the twenty questions that come up most, and upload only what answers those. You can always add more.
Good candidates:
- Policy pages — returns, cancellations, warranty, delivery, eligibility.
- Product or service detail — specifications, what is included, what is not.
- Process explanations — how to book, how to claim, what to bring, what happens next.
- FAQs you already maintain, if the answers sit next to the questions.
- Onboarding material — the things a new customer asks in week one.
What to leave out
This matters more than what you put in, because irrelevant content does not sit quietly. It competes.
| Leave out | Why |
|---|---|
| Marketing copy and brochures | Adjective-dense, fact-light. Retrieves well, answers nothing. |
| Superseded versions of a live policy | The search cannot tell which of two contradictory passages is current. |
| Scanned PDFs and screenshots of text | A picture of a page has no extractable text and indexes as empty. |
| Slide decks exported as PDF | Layout is flattened; bullet fragments lose the sentence they belonged to. |
| Anything internal or confidential | The assistant answers visitors from these documents. Assume every sentence is quotable. |
| Prices and hours that change weekly | Fine to include — but own the update, or it will confidently state last quarter's figure. |
| Your entire website | Volume dilutes. Twenty focused pages beat two hundred vague ones. |
The confidentiality one deserves emphasis. There is no "internal only" flag on a document. If a passage is in the knowledge base, a visitor who asks the right question can be told what it says.
How to structure a document so it retrieves well
Retrieval works on passages, not whole files. Each document is split into overlapping chunks, and the search matches a chunk at a time — so a chunk has to make sense on its own, because it will be read on its own. Five rules do most of the work.
1. One topic per section, with a heading
Give every section a descriptive heading — "Return Policy for International Orders", not "Section 4.2". A chunk that begins with a real heading carries enough context to stand alone. A chunk that begins mid-paragraph does not.
2. Keep the question next to its answer
An FAQ where the answers live three pages from the questions retrieves badly, because the chunk holding the answer never mentions what was asked. Question, then answer, then the next question.
3. Spell out what a table or diagram says
Layout is flattened to text during processing. A tidy comparison table becomes a run of unlabelled cells; a flowchart becomes nothing at all. If a fact only exists in a diagram, write the sentence too.
4. Say the thing, don't point at it
"Contact support for pricing details" gives the assistant nothing. "Standard onboarding is included; expedited onboarding is quoted per project" gives it something to say. Same for cross-references: if the shipping policy depends on the returns policy, restate the dependency in both.
5. Split large manuals
A single document indexes up to 500,000 characters, and the tail past that is simply not indexed. Long before the cap, a 300-page manual is doing you no favours anyway — split it by topic so each document has a coherent subject.
What actually happens when a visitor asks something
Worth knowing, because it explains three behaviours people report as bugs. Each visitor message triggers a search over the indexed passages of that assistant's documents. The most relevant passages — up to five, and only those clearing a relevance threshold — are handed to the model alongside your instructions before it writes the reply.
- Nothing relevant, nothing added. If no passage clears the threshold, the assistant answers from its instructions alone. Uploading a document does not force it to be mentioned.
- Small talk skips the search. Greetings and thanks are not searched, so "hi" will never be grounded in your price list.
- Grounded replies are marked. When an answer used your documents, the visitor sees a Referenced knowledge base marker under that message. That marker is your fastest debugging tool: if it is missing on an answer you expected to be grounded, the retrieval did not fire, and the problem is the document, not the model.
Documents supply facts; behaviour still comes from Instructions on the General tab — see Set Up Your Assistant's Persona and Tone for that half. A knowledge base cannot make a badly briefed assistant answer well, only a well-briefed one accurate.
Testing your uploads
Test with the questions a stranger types, not the ones you had in mind while writing the document.
- Upload, and wait for every row to reach
indexed. - Open Test Widget on the Deploy tab.
- Ask ten questions your documents should answer — in the visitor's words. "can i send it back" rather than "what is the returns policy".
- Check the Referenced knowledge base marker appears where an answer should be grounded.
- For any vague answer, find the section that should have answered it and ask why a chunk of it would not stand alone.
Almost every failure at step 5 is one of the five structure rules above.
Keeping it accurate over time
A knowledge base rots quietly. The assistant never tells you a document is stale; it just keeps quoting it. The Knowledge Base That Stays Accurate goes deep on chunking, re-indexing and curation; for getting started, two habits are enough:
- Delete before you add. When a policy changes, remove the old document in the same sitting as you upload the new one. Two versions of the truth is the worst state to be in.
- Let the transcripts choose your next upload. The Needs review feed on the Analytics Quality tab collects conversations where a visitor disliked a reply or the assistant fell back on an "I don't know" answer. That feed is a list of missing documents, written by your customers — Reading Your Assistant Analytics covers how to work it.
Start with the twenty questions you already know you get. Upload the four documents that answer them, check the statuses, ask ten questions in a visitor's words, and expand from what next week's transcripts tell you.
You can set all of this up on an assistant at hiroi.ai.