Home › Resources › Crawlable is not indexed

Crawlable is not indexed, and indexed is not cited

Updated · By Mike Hsu, founder and principal engineer, AnswerOpen

Key takeaways

  • Four sequential gates: edge access, index presence, passage retrieval, and model selection.
  • Indexing is three separate prerequisites: over 87% of ChatGPT citations match Bing top organic, Claude aligns about 86.7% with Brave, and Google AI surfaces need Google.
  • Roughly 90% of ChatGPT citations come from pages ranking 21st or worse, because engines retrieve passages for sub-queries.
  • Each gate has a distinct symptom, and they are constantly mistaken for one another.
  • Fix them in order: gate one is usually a config change, gate four is a content programme.

Crawlable is not indexed, indexed is not retrieved, and retrieved is not cited. Four separate gates sit between your server and your brand name appearing in an AI answer, and passing one says nothing about the next. Most AEO advice collapses them into a single idea, which is why so much of it fails to move anything.

What are the four gates?

  1. Access. Your edge returns a real 200 with the full body to the engine's crawler. Not robots.txt saying yes — the actual response.
  2. Indexing. The underlying search index has your page. This is usually somebody else's index: over 87% of ChatGPT citations match Bing's top organic results, Claude aligns about 86.7% with Brave, and Google's AI surfaces require Google indexing. Three separate prerequisites, not one.
  3. Retrieval. Your page comes back for the sub-query the engine actually issued. Engines fan a single prompt into anywhere from two to a dozen sub-queries, so this is rarely the head term you optimized for.
  4. Selection. The model picks a passage from your page over the competing passages it retrieved, and attributes it.
Four sequential gates between a server and an AI citation: access, indexing, retrieval, and selection, each with the failure symptom underneath.
Each gate has a distinct failure symptom, and they are frequently mistaken for each other.

Does ranking still decide it?

Less than most people assume, and in a counterintuitive direction. Only around 38% of AI Overview citations come from Google's top ten, and roughly 90% of ChatGPT citations come from pages ranking twenty-first or worse for the related query. That is not evidence that ranking is irrelevant — it is evidence that the engine is not searching your head term. It decomposed the question and retrieved passages for the parts. A page that answers one specific sub-question completely beats a page that ranks well for the broad one.

Which gate is failing, by symptom.
SymptomProbable gateFirst thing to check
No engine has ever mentioned youAccessFetch your homepage with an AI crawler user-agent and check status and word count
Google's AI mentions you, ChatGPT never doesIndexingBing Webmaster Tools — are you in the Bing index at all?
You are cited for your brand, never for topicsRetrievalWhether any page answers the specific sub-question, under a heading phrased as that question
Competitors are quoted from pages weaker than yoursSelectionWhether your answer is self-contained in the first 30% of the page

What does query fan-out change about content?

It changes what a page should be. If one prompt becomes a dozen retrievals, then the unit of competition is the passage, not the page. That argues for descriptive question-form headings, sections short enough to survive chunking (current retrieval benchmarks converge on roughly 256–512-token chunks), and each section answering its own heading without depending on the paragraph above it. Position matters too: 44.2% of ChatGPT citations come from the first 30% of a page's content, and 78.4% of question-tied citations attach to a heading.

Gate one is usually a one-line configuration change. Gate four is a content programme. People routinely buy the programme while the configuration is still broken.

— AnswerOpen diagnostic note on ordering the four gates

What should you do about it, in order?

  1. Prove access first. Everything downstream is unmeasurable until this passes.
  2. Verify indexing separately in Google, Bing and Brave. They are different prerequisites and you can fail one while passing the others.
  3. Rewrite for sub-questions: question-form headings, self-contained sections, the direct answer before the context.
  4. Only then work on selection — statistics, quotations, cited sources, tables. See what actually correlates.

The reason the order matters is cost. Gate one is usually a one-line configuration change. Gate four is a content programme. People routinely buy the programme while the configuration is still broken.

Related: testing gate one · how we measure all four. Ground truth worth having: Bing Webmaster Tools publishes AI performance data that no third-party tool can replicate.

Primary sources

Everything asserted above traces to one of these. Operator documentation changes often; check the current version before relying on any of it.

Frequently asked questions

Why does my site rank in Google but never appear in ChatGPT?

Most likely an indexing gate, not a content problem. ChatGPT's grounding leans heavily on Bing, where over 87% of its citations match top organic results. If you are absent or weak in Bing's index, Google rankings will not help you.

Does ranking still matter for AI citation?

Loosely. Only around 38% of AI Overview citations come from Google's top ten, and about 90% of ChatGPT citations come from pages ranking 21st or worse for the related query. Engines decompose a prompt into sub-queries and retrieve passages, so answering one sub-question completely beats ranking for the broad term.

What is query fan-out?

An engine turning one user prompt into several search queries, often two to a dozen, then retrieving passages for each. It means the unit of competition is a passage under a heading, not a whole page.

How do I tell which gate is failing?

By symptom. No engine mentions you at all points to access. Google's AI mentions you but ChatGPT never does points to indexing. Cited for your brand but never for topics points to retrieval. Competitors quoted from weaker pages points to selection.

Where can I get real citation data rather than estimates?

Bing Webmaster Tools publishes AI performance data showing actual Copilot citations and grounding queries. It is first-party telemetry, and it costs nothing.