Skip to content

Add knowledge

Teach your agent with a website crawl, single pages, files, text and Q&A pairs, and keep it up to date.

Your agent answers from its knowledge base, the content you give it. Good knowledge is the single biggest factor in good answers. Everything on this page happens on the agent's Knowledge tab.

Acme Assistant's knowledge: the www.acme.example crawl, product manuals and curated Q&A.

Source types

Click Add source and choose a type:

Crawl a website
Your whole site or help centre. Buddy reads your sitemap, follows links and re-crawls on a schedule.
Single page
One URL, such as your FAQ or pricing page.
Upload a document
PDF, Word, PowerPoint, Excel, TXT, Markdown, CSV or HTML files, such as a product manual.
Plain text
Anything that isn't online: policies, opening hours, internal notes you want customers to know.
Q&A pairs
Exact answers to specific questions. This is the most precise source type.

Fill in the dialog and click Add & index. New content is searchable as soon as its status changes to ready.

Pick a source type to add.

Crawl a website

  1. Choose Add source → Crawl a website

  2. Enter the start URL

    For example https://www.acme.example. To crawl only a section, start there, for example your help centre.

  3. Set Max pages

    The default is 100. The crawl never goes beyond your plan's page limit, shared across all your sources.

  4. Click Add & index

    The status moves from pending to processing to ready. Large sites can take a few minutes.

What the crawler does

  • Reads /sitemap.xml (including nested sitemaps), then follows links on the same domain, up to four links deep.
  • Respects your robots.txt.
  • Skips pages that aren't useful to customers, such as cart, checkout, login, sign-up, account and admin pages, and files that aren't web pages.
  • Handles JavaScript-heavy sites. If a page shows almost no text without JavaScript, Buddy renders it in a headless browser first. Those pages are marked JS-rendered.
  • Adds the pages it finds to your agent's site map, so it can take visitors there. See Page navigation.

Coverage and page caps

Each crawl shows how many pages were fetched out of how many were found, for example “85 of 120 found — capped”. If a crawl stopped at its cap, raise max on the source row and re-run it. You'll see a note when your plan's page or character limit stopped some content from being indexed.

Keep knowledge up to date

Website and single-page sources have a Schedule: Manual, Weekly or Daily. New website sources are re-crawled weekly by default. To refresh a source right away, click Re-index now on its row. For example, do this after you publish a new Acme Lock support article.

Re-crawls are efficient: Buddy fingerprints every page and only re-processes the pages that changed. Pages that disappeared from your site are removed from the knowledge base.

To refresh an uploaded file, delete the source and upload the new version.

Upload documents

Choose Upload a document. Accepted file types:

.pdf .docx .pptx .xlsx .txt .md .csv .html .htm, up to 25 MB per file.

  • Slides are indexed together with their speaker notes.
  • Spreadsheets and CSV files: each row becomes one searchable record, with the first row used as column headers. This works well for price lists or a table of Acme Home Hub compatible devices.
  • Scanned PDFs without a text layer can't be read yet, because there is no OCR. Upload a text-based PDF or paste the text as a Plain text source instead.
  • Older formats such as .doc and .ppt aren't supported. Save them as .docx or .pptx first.

Plain text and Q&A pairs

Plain text takes a Text field and an optional Name. Q&A pairs takes any number of pairs in the Pairs field, written as Q: / A: blocks separated by a blank line:

Q: Does Acme Cam work without Wi-Fi?
A: Acme Cam records to its microSD card when Wi-Fi is down and uploads the clips once it reconnects.

Q: How long is the Acme Lock warranty?
A: Acme Lock has a 2-year limited warranty from the date of purchase.

See what the agent knows

Click View indexed pages on any source to see every page it contains, how many characters and chunks each one has, and when it was indexed. Click a page to see the exact passages (chunks) the agent retrieves from it. Only the most relevant chunks are sent with each question.

A page with 0 chunks was fetched but had almost no usable text after navigation and boilerplate were removed, so the agent can't answer from it.

Indexed pages for the www.acme.example crawl. Click a page to see its chunks.

Turn sources on and off

Use the switch on a source's row to hide its content from the agent without deleting it, for example a seasonal promotion. Turn it back on to restore it. To remove a source and its content for good, use Delete source.

Answers you approve in the answer bank are stored in a special Corrections source. It's managed from the Answers tab, and the agent prefers it when two sources are equally relevant. See Answer bank and knowledge gaps.

Content flagged as prompt injection

Buddy checks every chunk it indexes for text that looks like instructions aimed at an AI. Examples are “ignore your previous instructions”, role hijacks, jailbreak phrases, hidden Unicode characters or instructions hidden in encoded text. Such text could make your agent misbehave, so flagged chunks are hidden from the agent automatically.

  1. Look for flagged for review on a source row, or the warning at the top of the tab

  2. Open View indexed pages and find the chunks marked Flagged

    Each one shows the reason it was flagged.

  3. If the text is harmless, click Mark safe & enable

    The agent can use it again, and the same text won't be flagged on future re-crawls. Changed your mind? Click Hide again. Every review is recorded in the audit log.

Plan limits

Pages, files and indexed characters are counted across all sources in your workspace.

FreeStarterGrowthPro
Website pages201005003,000
Files31050500
Characters indexed300K2M10M50M
Re-crawlManualWeeklyDailyDaily

When you hit a limit, the dashboard tells you which one, for example “File limit reached: your plan allows 10 files. Delete a file or upgrade.” Remove sources you no longer need, or upgrade your plan.

Troubleshooting

›A source shows the status error

The reason is shown under the source name. Common causes: the site blocks crawlers in robots.txt, the URL needs a login, a PDF is scanned, or a plan limit was reached.

›The crawl found fewer pages than my site has

Check whether the source says capped and raise max on the row, then re-run it. Make sure the pages are linked from your site or listed in your sitemap. Pages behind a login and pages that robots.txt blocks aren't crawled.

›The agent gives an outdated answer

Click Re-index now on the source, or set its schedule to Weekly or Daily. For a fact that must always be exact, add a Q&A pair or an approved answer.

›The agent says it doesn't know, but the answer is on my site

Open View indexed pages and check the page is there with chunks. If it has 0 chunks, the content probably isn't in the page text (for example it's in an image). Add it as plain text or a Q&A pair. Then test in the playground and look at the retrieved sources.