Your robots.txt says allow. Your firewall says 403.
Key takeaways
- Cloudflare has default-blocked AI crawlers on new zones since July 2025 and rewrote managed robots.txt on millions of domains.
- Three failure signatures: a hard block, a JavaScript challenge wearing a 200, and a thin body.
- Blocking answer-surfacing bots removes you from answers now; blocking training crawlers does not.
- A restrictive wildcard group silently catches every bot without a group of its own.
- robots.txt cannot fix an edge rule. The exemption has to be made at the edge.
The signature AI-visibility failure of 2025 and 2026 is a site whose robots.txt welcomes every answer engine while its edge quietly returns 403 or a JavaScript challenge to the same bots. Nothing in your analytics reports it, robots.txt linters pass you, and the only symptom is silence.
Why does this happen to sites that never chose it?
Because it is frequently a default rather than a decision. Cloudflare has default-blocked AI crawlers on new zones since July 2025 and rewrote managed robots.txt across millions of domains. Web application firewalls commonly answer unfamiliar user-agents with a JavaScript challenge — and a challenge is unsolvable by definition for a crawler that does not execute JavaScript. Bot management scores non-browser clients as suspicious. Every one of these is a sensible security posture that also happens to remove you from AI answers, and none of them updates your robots.txt to say so.
How do you test your own edge in five minutes?
Fetch the same URL with a browser user-agent and with an AI crawler user-agent, and compare status and body size:
curl -sSo /dev/null -w "%{http_code} %{size_download}\n" -A "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/126 Safari/537.36" https://example.com/
curl -sSo /dev/null -w "%{http_code} %{size_download}\n" -A "PerplexityBot/1.0 (+https://perplexity.ai/perplexitybot)" https://example.com/
Repeat for each bot in the table below. Three failure signatures matter, and the second and third are the ones people miss:
| Signature | What you see | Meaning |
|---|---|---|
| Hard block | 403, 401, 429 or 503 for the bot, 200 for the browser | The edge is refusing outright |
| Challenge | 200, but the body contains an interstitial fingerprint | A block wearing a success code — unsolvable without JavaScript |
| Thin body | 200, real page, but far fewer bytes than the browser got | A stripped or cached stub; the content is not there |
Which bots actually matter?
The classes have opposite stakes, and treating "AI bots" as one bucket is the most common analytical error in this field. Blocking an answer-surfacing bot removes you from answers immediately. Blocking a training crawler does not affect whether you are cited today at all — it is a separate, legitimate business decision about your content being used for model training.
| User-agent | Class | Blocking it means |
|---|---|---|
| OAI-SearchBot | Answer surfacing | Removed from ChatGPT search results |
| ChatGPT-User | User-initiated fetch | ChatGPT cannot open your page on request |
| PerplexityBot | Answer surfacing | Removed from Perplexity answers |
| Claude-SearchBot | Answer surfacing | Removed from Claude's search results |
| Bingbot | Index prerequisite | Removed from Copilot and much of ChatGPT's grounding |
| Googlebot | Index prerequisite | Removed from AI Overviews and AI Mode |
| GPTBot | Training | Excluded from model training; no effect on today's citations |
| ClaudeBot | Training | Excluded from model training; no effect on today's citations |
| Google-Extended | Training | Excluded from Gemini training; no effect on AI Overview inclusion |
What is wildcard collateral damage?
A restrictive User-agent: * group catches every bot that does not
have a group of its own. Sites paste a published AI blocklist to stop training
crawlers, and it silently takes the answer-surfacing bots with it because those
were never named. Under longest-match agent-token semantics, the only reliable fix
is an explicit group per bot you want to allow. Our audit computes exactly this
set difference: which citation-critical bots have no dedicated group and therefore
inherit your wildcard.
How do you fix it properly?
- Add an explicit allow group for every answer-surfacing bot, above the wildcard.
- Exempt those user-agents from bot management, rate limiting and challenge rules at the edge — this is the step that is actually missing, and robots.txt cannot do it for you.
- Exempt
/robots.txtand/sitemap*.xmlfrom every edge rule. Some configurations block bots on the very file that would grant permission. - Re-test with the curl commands above and confirm status, and body size, match the browser.
- Verify in your server logs, matching source IPs against the operator-published IP feeds. This is the step that converts probable into proven.
Always caption a 403 finding as a user-agent-level block detected from a non-operator IP. Spoofed-user-agent probes can trigger fake-bot rules, and the honest report says so.
— AnswerOpen check catalog, note on CA-04
Our audit runs this matrix for nine user-agents against your homepage and a sample of content URLs, then cross-joins it with your declared robots.txt policy. Related: how the probe pass works · what comes after access · have us run it for you. Operator documentation: OpenAI, Perplexity.
Primary sources
Everything asserted above traces to one of these. Operator documentation changes often; check the current version before relying on any of it.
- OpenAI crawler documentation — the authoritative list of GPTBot, OAI-SearchBot and ChatGPT-User behavior
- Anthropic crawler documentation — ClaudeBot, Claude-SearchBot and Claude-User
- Perplexity bot documentation — PerplexityBot and Perplexity-User
- Robots exclusion protocol (Wikipedia) — the longest-match user-agent semantics that make wildcard groups dangerous
Frequently asked questions
Why would my firewall block AI crawlers if my robots.txt allows them?
Because the two are configured in different places by different people, and the edge often blocks by default. Cloudflare has default-blocked AI crawlers on new zones since July 2025, and web application firewalls commonly answer unfamiliar user-agents with a JavaScript challenge that no non-rendering crawler can solve.
How do I test whether my edge is blocking AI crawlers?
Fetch the same URL twice with curl, once with a browser user-agent and once with a crawler user-agent such as PerplexityBot, and compare the status code and downloaded byte count. A different status, or a much smaller body, is your answer.
Which bots should I definitely allow?
The answer-surfacing ones: OAI-SearchBot, ChatGPT-User, PerplexityBot, Perplexity-User, Claude-SearchBot and Claude-User, plus Googlebot and Bingbot as index prerequisites. Whether you allow the training crawlers GPTBot, ClaudeBot and Google-Extended is a separate business decision with no effect on today's citations.
Is blocking GPTBot the same as blocking ChatGPT?
No, and this is the most consequential misunderstanding in the field. GPTBot is a training crawler. OAI-SearchBot and ChatGPT-User are what fetch pages for answers. Blocking the first affects model training; blocking the second two removes you from ChatGPT's answers immediately.
A 403 showed up in my probe. Is it definitely a block?
It is a user-agent-level block detected from a non-operator IP address. Fake-bot rules can produce the same signature against a spoofed user-agent, so the honest next step is checking your own server logs and matching source IPs against the operator-published IP feeds.