SEO

Robots.txt

Short definition

Robots.txt is a simple text file at a site's root that tells search engines and other crawlers which pages they may or may not crawl.

A site can close off sections that have no business appearing in search results, an admin panel or a shopping basket page, through robots.txt. The file is a directive rather than a rule, well-behaved crawlers respect it, but it is not an enforceable barrier.

Get robots.txt wrong and an entire site can end up blocked from crawling by accident, disappearing from search results altogether. More recently, the file has also been used to state whether AI crawlers may collect content as training data.

Robots.txt is often confused with the noindex tag, though the two do different jobs: a disallow rule stops a page from being crawled, but if another page links to it, a search engine can list that address without a description, removing a page from the index needs a noindex tag or a removal request. How the file is being read can be checked through Google Search Console's URL inspection tool.

Why it matters

A misconfigured robots.txt can make an entire site, or its most important pages, invisible to a search engine. Getting it right keeps the pages that should stay hidden hidden, and guarantees the ones that matter get crawled.

Illustrative example

During a software update, a business accidentally left a robots.txt line that blocked the whole site, and it vanished from search for a week.

Related terms

Related article

If you don’t know where to start, that’s fine; you’re in the right place.

Your project might already be clear in your head, or still just an idea. Either works. On a short call we talk through where you are and where you could go, together.

Let’s set up a call
Let’s talk about your project