Scoreling

Google's rules · 8 crawlers · free

Robots.txt checker

Enter a page to test the site's robots.txt against it, or paste a file you are about to upload. See which crawlers may reach the page, the rule that decides for each, and every line that does not do what it seems to.

or paste a robots.txt

How Google reads robots.txt

Three rules decide almost every case, and the checker applies all three.

One group per crawler

A crawler obeys the group that names it and ignores the rest. Only a crawler with no group of its own falls back to User-agent: *.

The longest rule wins

Allow: /shop/sale beats Disallow: /shop because it is longer. On a tie, Allow wins.

Paths match from the start

* matches anything and $ marks the end, so Disallow: /*.pdf$ blocks every PDF.

Questions about robots.txt

Where is robots.txt?

Always at the root of the host: https://example.com/robots.txt. A file anywhere else is ignored, and each subdomain needs its own, so shop.example.com does not read the one on www.example.com.

How do I fix "Blocked by robots.txt" in Search Console?

Find the rule that blocks the page with the checker above: enter the page's URL and it shows which line decides for Googlebot. Remove or narrow that Disallow, or add an Allow for the page, then use Validate fix in Search Console. If the page should stay out of search, leave the rule and remove the page from your sitemap instead.

Does robots.txt keep a page out of Google?

Not reliably. It stops Google from crawling the page, but a blocked page that other sites link to can still be indexed with just its address. To keep a page out of search, let Google crawl it and use a noindex robots meta tag.

Should I block GPTBot?

It depends on what you want. GPTBot collects pages to train OpenAI's models; blocking it does not remove you from ChatGPT search, which uses OAI-SearchBot. Many publishers block the training crawlers (GPTBot, ClaudeBot, Google-Extended, CCBot) and allow the search ones.

How long does Google take to see a change?

Google usually caches robots.txt for up to a day. In Search Console, the robots.txt report shows the version Google last fetched and lets you ask for a recrawl.