How robots.txt rules are matched
- Pick the group. A crawler follows the group for its most specific user-agent (e.g.
Googlebot-NewsbeforeGooglebot), and falls back toUser-agent: *. Groups naming the same agent are combined. - Longest match wins. Inside the group, the rule with the longest matching path decides.
Allow: /blog/publicbeatsDisallow: /blog. - Ties go to Allow. If an Allow and a Disallow are the same length, the URL is allowed.
- Wildcards.
*matches any sequence of characters, and$marks the end of the URL:Disallow: /*.pdf$blocks PDFs only. - Paths are case-sensitive and include the query string.
How to add a sitemap to robots.txt
Add Sitemap: https://example.com/sitemap.xml on its own line, using the full URL. It applies to all crawlers regardless of where it sits in the file. Check the sitemap itself with the sitemap validator.
Common robots.txt mistakes
- Disallow: / left over from staging, which blocks the entire site.
- Blocking CSS or JavaScript that Google needs to render pages.
- Using Disallow to de-index pages. It stops crawling, not indexing; use noindex instead.
- Typos in directives (
Dissallow), which are silently ignored. - Unsupported directives such as
Noindex:orCrawl-delay, which Google ignores.
robots.txt and AI crawlers
AI companies use separate user-agents for search and for training. If you want to appear in AI answers, keep the search crawlers allowed even if you block the training ones. See the LLM SEO tool for a full AI-search readiness check.