Hostinger promo code — 20% off hosting
Hostinger coupon for cheap web hosting & WordPress. Enter promo code YAHMUBASH14E — works for Pakistan (PK) and global checkout.
YAHMUBASH14E
Write robots.txt for AI crawlers in seconds. The recommended policy blocks training bots (GPTBot, ClaudeBot, Google-Extended, CCBot) and allows search bots (OAI-SearchBot, ChatGPT-User, Claude-Web, PerplexityBot, Googlebot).
Generate robots.txt rules for AI crawlers
Ready to publish
Generated robots.txt
Save as robots.txt at your domain root. Re-check with our
AI Crawler Checker.
The recommended default is block training, allow AI search. Other options block GPTBot only, block all AI crawlers, or allow everything.
Disallow GPTBot, ClaudeBot, Google-Extended, CCBot, Applebot-Extended, and Amazonbot. Allow OAI-SearchBot, ChatGPT-User, Claude-Web, PerplexityBot, and Googlebot.
Googlebot, Bingbot, and major AI bots can fetch public pages.
Disallow GPTBot, ClaudeBot, PerplexityBot, Google-Extended, CCBot, and more.
Keep other AI crawlers allowed; block OpenAI GPTBot only.
Block ClaudeBot and Claude-Web; allow other bots.
Block Perplexity's crawler only.
Opt out of Google's generative-AI training crawler token.
Download a commented starter file and edit per-bot rules manually.
Allow search bots if you want to be cited. Block training bots if you do not want content used to train models. They are separate robots.txt tokens.
| Crawler | Vendor | Purpose | Typical policy |
|---|---|---|---|
GPTBot |
OpenAI | Foundation-model training | Training — usually block |
OAI-SearchBot |
OpenAI | ChatGPT search indexing | Search — allow to be cited |
ChatGPT-User |
OpenAI | User-triggered fetches | Search — allow live answers |
ClaudeBot |
Anthropic | Claude training crawler | Training — usually block |
Claude-Web |
Anthropic | Claude user/web fetches | Search — allow citations |
PerplexityBot |
Perplexity | Perplexity AI answers | Search — allow citations |
Google-Extended |
Generative AI training opt-out token | Training — block independently of Googlebot | |
CCBot |
Common Crawl | Open web corpus | Training — usually block |
Applebot-Extended |
Apple | Apple AI training | Training — usually block |
Amazonbot |
Amazon | Amazon AI indexing | Training — usually block |
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: CCBot
Disallow: /
User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: /
User-agent: Claude-Web
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: *
Allow: /
User-agent: GPTBot
Disallow: /
User-agent: *
Allow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Claude-Web
Disallow: /
User-agent: *
Allow: /
User-agent: PerplexityBot
Disallow: /
User-agent: *
Allow: /
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: PerplexityBot
Disallow: /
User-agent: *
Allow: /
Also try the AI Crawler Checker for a full crawl readiness report.
For most public sites in 2026: allow Googlebot and AI search crawlers (OAI-SearchBot, ChatGPT-User, Claude-Web, PerplexityBot) so you can be cited. Block training crawlers (GPTBot, ClaudeBot, Google-Extended, CCBot) if you do not want content used for model training. llms.txt is not a training opt-out.
An AI crawler is an automated bot that fetches public web pages for AI search, answers, or model training. Examples include GPTBot, ClaudeBot, and PerplexityBot.
Block training crawlers if you want to limit model training use of your content. Allow search crawlers like OAI-SearchBot if you want ChatGPT or AI search visibility.
Add a User-agent: GPTBot block with Disallow: / before your wildcard Allow rules. Use this generator and select Block only GPTBot.
Use User-agent: * with Allow: / and avoid Disallow rules that block AI user-agents. Select Allow all AI crawlers in this tool to generate a permissive starter file.
Google-Extended is a robots.txt token to manage whether Google may use your content for generative AI training surfaces. It is separate from Googlebot search crawling.
robots.txt can block well-behaved crawlers that respect it, but it is not DRM. It is a public crawl policy — not authentication or legal protection.
Sponsored links — we may earn a commission at no extra cost to you.