AI crawler
Short definition
An AI crawler is the automated program an AI company sends across the web to gather content for training a model or answering questions.
Crawlers such as GPTBot, Google-Extended or ClaudeBot move through sites much like a classic search engine crawler, but what they collect serves a different purpose: training an AI model, or feeding it context when generating an instant answer.
A site can allow or block these crawlers through robots.txt. Blocking them risks the content never appearing in an AI-generated answer at all; allowing them raises a separate concern about content being copied or summarised without permission.
Whether a site allows these crawlers can be checked in two places: reading how robots.txt handles user agent names such as GPTBot, ClaudeBot, PerplexityBot and Google-Extended, or looking through server access logs to see whether those names actually visit, since a blocked crawler leaves no trace there at all, which is the clearer signal when robots.txt rules look ambiguous.
Why it matters
A business that wants its content to appear in AI-generated answers needs to allow the relevant crawlers in. That decision is no longer a technical footnote, it directly shapes a brand's visibility in AI-driven search.
Illustrative example
An SEO agency found that a client's robots.txt was accidentally blocking every AI crawler, and the brand was not appearing in any AI assistant's answers as a result.
