AI search

AI crawler

Short definition

An AI crawler is the automated program an AI company sends across the web to gather content for training a model or answering questions.

Crawlers such as GPTBot, Google-Extended or ClaudeBot move through sites much like a classic search engine crawler, but what they collect serves a different purpose: training an AI model, or feeding it context when generating an instant answer.

A site can allow or block these crawlers through robots.txt. Blocking them risks the content never appearing in an AI-generated answer at all; allowing them raises a separate concern about content being copied or summarised without permission.

Whether a site allows these crawlers can be checked in two places: reading how robots.txt handles user agent names such as GPTBot, ClaudeBot, PerplexityBot and Google-Extended, or looking through server access logs to see whether those names actually visit, since a blocked crawler leaves no trace there at all, which is the clearer signal when robots.txt rules look ambiguous.

Why it matters

A business that wants its content to appear in AI-generated answers needs to allow the relevant crawlers in. That decision is no longer a technical footnote, it directly shapes a brand's visibility in AI-driven search.

Illustrative example

An SEO agency found that a client's robots.txt was accidentally blocking every AI crawler, and the brand was not appearing in any AI assistant's answers as a result.

Related terms

Related article

If you don’t know where to start, that’s fine; you’re in the right place.

Your project might already be clear in your head, or still just an idea. Either works. On a short call we talk through where you are and where you could go, together.

Let’s set up a call
Let’s talk about your project