How an answer engine audit actually runs
Key takeaways
- One crawl, five passes: discovery, edge probing, per-page checks, scoring, fix plan.
- Nine real AI crawler user-agents are compared against a browser baseline on status, challenge markers and body parity.
- 61 deterministic checks per page, each a pure function of raw HTML and headers, so results reproduce exactly.
- Four per-engine readiness lenses fall out of the same crawl at no extra cost.
- The fix plan is ranked by weight times affected pages, and names every URL.
An audit runs in five passes over one crawl: discovery, edge probing, per-page checks, scoring, and a ranked fix plan. It uses no headless browser and makes no model calls at audit time, so it is deterministic — run it twice on an unchanged site and you get the same numbers.
Pass 1 — how does it find your pages?
Sitemap-first. We fetch /robots.txt, follow every declared
sitemap (including nested index sitemaps), and fall back to link discovery from
the homepage where a sitemap is missing or short. A 679-page production site
completed in about ten minutes and roughly five thousand HTTP requests, with
polite concurrency and a delay you can raise for fragile origins.
Pass 2 — what does the multi-user-agent probe do?
Your homepage and a sample of content URLs are fetched nine times over, once per real AI crawler user-agent string, plus a Chrome baseline. Each response is graded on three axes:
- Status — anything other than 200 for a bot that a browser gets 200 for.
- Challenge markers — interstitial fingerprints that a non-JavaScript client can never solve.
- Body parity — word count against the browser baseline. A 200 that returns a stub is still a block.
Then the finding is cross-joined against what /robots.txt
declares. Agreement is a pass. Disagreement — policy says allow, edge says no —
is the highest-value finding the tool produces.
We always caption a bot-level block as detected from a spoofed-user-agent probe rather than from an operator IP. Fake-bot rules can produce the same signature, which is precisely why log verification exists as the tier that upgrades probable to proven.
Pass 3 — what runs on each page?
Every fetched page runs through the check catalog: 61 checks in the current engine build, each a pure function of the raw HTML and response headers. No check depends on a rendered DOM, a screenshot or a model opinion, which is what makes the result reproducible and cheap enough to run on ten thousand pages.
| Field | Example |
|---|---|
| Check id | AS02 |
| Category and weight | answer structure, 14 |
| What it looks for | A direct answer of 40–450 characters in the first paragraph, before any background |
| The fix | Open with a one to three sentence direct answer to the H1's question; models lift the first self-contained answer block |
| Evidence tier | measured position analysis of 1.2M answers and 18K verified citations |
Pass 4 — how is the 0-100 built?
Each check carries a weight inside its category; category scores roll up into the weighted total described on the home page. Alongside the AEO score we compute a separate classic SEO baseline, because search indexability is a hard prerequisite rather than a competing metric: over 87% of ChatGPT citations match Bing's top organic results, Claude aligns about 86.7% with Brave, and Google's AI surfaces require Google indexing. We also derive four per-engine readiness sub-scores from the same crawl at no extra cost.
| Lens | What it keys on |
|---|---|
| ChatGPT / SearchGPT | OAI-SearchBot and ChatGPT-User access, plus the Bing indexability path |
| Google AI Overviews / Gemini | Snippet directives, indexability, Googlebot parity |
| Perplexity | PerplexityBot and Perplexity-User access, freshness signals |
| Bing Copilot | noarchive and nocache directives, Bingbot access, sitemap health |
Pass 5 — what happens to the findings?
Failures are grouped into a fix plan ranked by weight times affected pages, so the first item is always the one that buys the most score for the least work, with every affected URL named. On a fix engagement those become staged edits: each target file is backed up, the diff is generated from your site's own entity data rather than a generic template, and nothing is written to your origin until you approve it. Then we re-run the audit and show you the delta.
Vague suggestions are the single most common complaint against this category of tool. A fix that a developer cannot paste is not a fix; it is homework.
— from our teardown of the five leading paid AEO platforms, August 2026
What do you actually receive?
- A dashboard with five views: overview, per-page table, ranked fix plan, methodology, and the run log.
- The raw JSON report — every check result for every page, yours to keep and diff against later runs.
- The probe matrix: nine user-agents by sampled URLs, with status, challenge and parity for each cell.
- An off-site action plan covering the factors we deliberately do not score.
Related reading: the full check catalog, what a policy-versus-edge mismatch looks like, and what an engagement costs.
Primary sources
Everything asserted above traces to one of these. Operator documentation changes often; check the current version before relying on any of it.
- OpenAI crawler documentation — the authoritative list of GPTBot, OAI-SearchBot and ChatGPT-User behavior
- Anthropic crawler documentation — ClaudeBot, Claude-SearchBot and Claude-User
- Perplexity bot documentation — PerplexityBot and Perplexity-User
- Google robots meta tag reference — how nosnippet, max-snippet and noindex gate AI surfaces
Request an audit of your site
There is no fake "instant scan" button on this site. A real audit crawls your sitemap, fetches every page, and runs nine live user-agent probes against your edge — it takes minutes of machine time, not milliseconds, and it is run by a human who reads the result before you see it. Send the URL and you get the headline score plus your single largest blocker back, at no charge.
Frequently asked questions
Does the audit slow my website down?
It is a polite, rate-limited, read-only crawl with configurable concurrency and delay. A 679-page site took about ten minutes and roughly five thousand requests. If your origin is fragile we slow the crawl down on request.
Why do you not render JavaScript?
Because no major AI crawler does. Rendering the page would measure something the engines never see, at much higher cost. Skipping it is what makes a ten-thousand-page crawl affordable and what makes the measurement match reality.
What does a body parity failure mean?
It means a bot user-agent received a 200 response, but the body contained far fewer words than the same URL served to a browser. That is a block wearing a success code, and it is invisible to any tool checking status codes alone.
Can a probe result be a false positive?
Yes, and we caption it accordingly. Our probes send real crawler user-agent strings from our own infrastructure rather than from the operator's IP ranges, so a fake-bot rule can produce the same signature as a genuine block. Server-log analysis with operator-IP verification is what upgrades a probable finding to a proven one.
What do I actually receive at the end?
A dashboard with overview, per-page, fix plan, methodology and run-log views; the raw JSON with every check result for every page; the nine-by-N probe matrix; and an off-site action plan covering the factors we deliberately do not score.