One group per crawler
A crawler obeys the group that names it and ignores the rest. Only a crawler with no group of its own falls back to User-agent: *.
Google's rules · 8 crawlers · free
Enter a page to test the site's robots.txt against it, or paste a file you are about to upload. See which crawlers may reach the page, the rule that decides for each, and every line that does not do what it seems to.
or paste a robots.txt
Three rules decide almost every case, and the checker applies all three.
A crawler obeys the group that names it and ignores the rest. Only a crawler with no group of its own falls back to User-agent: *.
Allow: /shop/sale beats Disallow: /shop because it is longer. On a tie, Allow wins.
* matches anything and $ marks the end, so Disallow: /*.pdf$ blocks every PDF.
Always at the root of the host: https://example.com/robots.txt. A file anywhere else is ignored, and each subdomain needs its own, so shop.example.com does not read the one on www.example.com.
Find the rule that blocks the page with the checker above: enter the page's URL and it shows which line decides for Googlebot. Remove or narrow that Disallow, or add an Allow for the page, then use Validate fix in Search Console. If the page should stay out of search, leave the rule and remove the page from your sitemap instead.
Not reliably. It stops Google from crawling the page, but a blocked page that other sites link to can still be indexed with just its address. To keep a page out of search, let Google crawl it and use a noindex robots meta tag.
It depends on what you want. GPTBot collects pages to train OpenAI's models; blocking it does not remove you from ChatGPT search, which uses OAI-SearchBot. Many publishers block the training crawlers (GPTBot, ClaudeBot, Google-Extended, CCBot) and allow the search ones.
Google usually caches robots.txt for up to a day. In Search Console, the robots.txt report shows the version Google last fetched and lets you ask for a recrawl.