What the Crawler Learns, and What It Misses
Last updated:
What the crawler does
It follows your site's links, reads the text of each page, splits it into passages and indexes them so relevant ones can be found later.
Your navigation, headings and page structure all help it. A well-organised site produces a better assistant, which is one more argument for tidy structure.
What it handles well
- Ordinary content pages with real text
- Service and product descriptions
- FAQ pages, which are ideal
- Blog posts and guides
- Policy and terms pages
What it misses
| Content | Why | Fix |
|---|---|---|
| Text inside images | Not readable as text | Add it as text on the page |
| Behind a login | Not reachable | Upload as a document |
| Loaded only after interaction | Not present in the page | Provide it another way |
| Linked PDFs | Sometimes reachable, sometimes not | Upload directly |
| Prices in a system, not on a page | Not published | Upload a price list |
Prices are the most common gap. If your prices are not written anywhere the crawler can reach, the bot cannot quote them — and it should say so rather than guess.
Improving what it learns
- Write the answers to your top questions as real pages
- Put specifics on the page, not in a downloadable file
- Use the words customers use, not only your internal terms
- Remove or update pages that contradict current reality
Old pages are the hidden problem
A superseded page still on your site is just as visible to the crawler as the current one. Two versions of a policy produce inconsistent answers.
Before crawling, spend twenty minutes finding and removing anything out of date. It is the single highest-value preparation.
Frequently asked questions
How long does a crawl take?
Will it crawl our whole site?
Can we exclude pages?
Does it re-read pages automatically?
Not sure what your site actually says?
Crawling it for a chatbot is a surprisingly good content audit. Try it at spideychat.com.