Home › Resources

Resources: what is measurable about AI search citation

Updated

Key takeaways

  • Four articles, each sourced to published studies or primary operator documentation.
  • Numbers that could not be traced to a source were deleted rather than rounded.
  • Where sources genuinely conflict, we say so instead of picking the flattering study.
  • Start with the crawler-access article if you get no AI mentions at all.
  • Every claim carries an evidence tier defined on the what-we-check page.

Four pieces on what is actually measurable about AI search citation, written from the same evidence base the audit engine scores against. Every claim carries its source, and the numbers we could not trace were deleted rather than rounded.

The articles

AI crawlers do not execute JavaScript

Hundreds of millions of observed fetches say GPTBot, ClaudeBot and PerplexityBot read raw HTML only. What that means for a client-rendered site, and why it makes static auditing sufficient rather than second-best.

Crawlable is not indexed, and indexed is not cited

Three separate gates sit between your server and a citation. Most AEO advice collapses them into one, which is why so much of it does not work.

What actually correlates with AI citation

The causal evidence, the correlational evidence, and the widely repeated multipliers that could not be traced to any source at all.

Your robots.txt says allow. Your firewall says 403.

The signature failure of 2025–26, why a robots.txt linter cannot see it, and how to test your own edge in about five minutes with curl.

Where should you start?

Suggested reading order by the symptom you have.
If this is your situationRead this first
You are getting no AI mentions at all, from any engineRobots says allow, firewall says 403
Your site is a React, Vue or Next.js client-rendered appAI crawlers do not execute JavaScript
You rank well in Google but never appear in AI answersCrawlable is not indexed
You are deciding what content work is worth doingWhat actually correlates

Never print an untraceable multiplier in customer-facing copy. Three independent research passes each caught vendor numbers that did not survive source-tracing, so the direction stays and the number goes.

— AnswerOpen research standard, August 2026

How is this material sourced?

From primary operator documentation and large-sample published studies, with each finding tagged by evidence tier. Where sources conflicted — freshness is the clearest case, where one dataset shows 65% of AI bot hits going to content under a year old while another finds the median cited page is around 500 days old for evergreen queries — we say so and audit the honest, machine-visible signal rather than picking the flattering study. The full tier definitions are on what we check.

Primary sources

Everything asserted above traces to one of these. Operator documentation changes often; check the current version before relying on any of it.

Frequently asked questions

Which article should I read first?

If no engine mentions you at all, start with the article on robots.txt versus your firewall. If your site is a client-rendered application, start with the JavaScript article. If you rank in Google but never appear in AI answers, start with crawlable is not indexed.

Are these articles a sales pitch?

They are written so that a competent developer can act on them without buying anything. Every test in them can be run with curl. We sell doing it at scale, across every page, and repeatedly.

How current is this material?

It reflects evidence and operator documentation reviewed in August 2026. This field moves fast: crawler user-agent strings, edge defaults and snippet directive semantics have all changed materially within eighteen months, so verify anything critical against the operator's current documentation.

Do you publish the raw research?

The evidence tiers, category weights and check definitions behind every article are summarized on the what-we-check page, and the full catalog ships inside every audit report.