Robots.txt & Sitemap Checker
Check your robots.txt and sitemap files on one screen, no sign-up required.
This tool first reads your domain's robots.txt file: whether it exists, its size, the number of rule groups, a warning for a Disallow: / that blocks the whole site, any Sitemap: lines, the Crawl-delay value and syntax errors such as unknown directives. It then reads the sitemap named in robots.txt, or /sitemap.xml if none is named; if that sitemap is an index it steps into the first two sub-files, downloading three sitemap files at most in total.
The report shows the sitemap type, whether it is a plain list or an index, the number of addresses, the share that carries a lastmod date, the newest and oldest lastmod values, addresses pointing at a different domain, a mix of http and https, how close the file sits to the official 50,000 address or 50 MB limit, and whether hreflang tags (xhtml:link) are present. It also sends a HEAD request to the first five addresses and lists their status codes.
This check makes at most ten outbound requests, so on a very large sitemap index it looks at only the first few parts rather than the whole tree; the address count shown reflects what was read, not every page your site has. If robots.txt blocks the whole site or no sitemap can be found, you will see that flagged separately, but only a tool such as Google Search Console can confirm how a search engine actually crawls your site.
FAQ
What if I have no sitemap file?
If the tool finds no Sitemap: line in robots.txt and nothing at /sitemap.xml either, it flags that as missing. In that case, it is worth creating and publishing a sitemap so search engines can discover your addresses.
Does it read every file in a large sitemap index?
No, it downloads three sitemap files at most and only steps into the first two sub-files of an index; the report states this limit clearly rather than claiming to represent the whole site.
Is the domain I enter stored?
No, nothing is kept permanently; if the same domain is checked again shortly afterwards, it is held in server memory for about ten minutes for a faster reply, then cleared.
Does setting Crawl-delay work on Google?
Google ignores this directive; Bing and some other crawlers may respect it, so the tool shows the value but its effect depends on which crawler reads it.
