Home › What we check

What we check, and how strong the evidence is

Updated

Key takeaways

  • 61 deterministic checks across five weighted categories, on every page.
  • Crawler access carries the largest weight at 30, because it is the failure that produces total silence.
  • Schema is deliberately held to 8 points: a matched 1,885-page analysis found no citation lift from adding JSON-LD.
  • Every finding carries an evidence tier, so you can tell a controlled experiment from a vendor claim.
  • llms.txt scores zero and off-site factors are reported rather than scored.

The engine runs 61 deterministic checks per page across five scored categories, and every finding carries an evidence tier that tells you how strong the underlying proof is. We publish the tiers because most numbers in this market do not survive source-tracing.

What do the evidence tiers mean?

A recommendation is only as good as the reason behind it. Each check is labelled with one of four tiers, and we say so in the report rather than flattening everything into "best practice":

The four evidence tiers used throughout the report.
TierMeansExample
causalControlled experiment with a measured effect sizeAdding quotations lifted generative visibility about 41% in the peer-reviewed GEO study (KDD 2024)
measuredLarge-sample observational correlation, source traceable44.2% of ChatGPT citations fall in the first 30% of page content
officialDocumented behavior from the engine operatorMicrosoft documents NOARCHIVE as full exclusion from Copilot answers
mechanismFollows from how retrieval demonstrably works, no effect size claimedSelf-contained sections survive chunking; chunk sizes converge around 256–512 tokens

What you will not find in our reports: untraceable multipliers. Claims such as "tables are cited 2.5 times more often" appear across this market and could not be traced to any source, so the direction survives in our catalog and the number does not.

Crawler access — 30 points

Ten checks covering whether an engine can reach you at all. robots.txt exists, parses, and returns text rather than an HTML error page. No answer-surfacing bot is disallowed — and the distinction matters, because blocking OAI-SearchBot, PerplexityBot or Claude-SearchBot removes you from answers now, while blocking the training crawlers GPTBot, ClaudeBot and Google-Extended does not. We compute wildcard collateral damage: which citation-critical bots have no group of their own and therefore silently inherit a restrictive User-agent: *. We probe the edge with nine live user-agents. And we check that robots.txt and your sitemaps are themselves fetchable by bots, because some configurations block bots on the very file that would grant them permission.

Answer structure — 26 points

Fifteen checks on whether a model can lift an answer out of the page. A single H1. A direct answer in the opening paragraph rather than a warm-up. Question-form headings, since 78.4% of question-tied citations attach to a heading. Sections short enough to survive chunking. A visible key-takeaways block. Concrete statistics, quotations and cited sources — the three levers with actual causal evidence behind them, measured at roughly +34%, +41% and +29% respectively. Tables and lists. An honest, visible updated date. At least 300 words of genuine answer, with no reward for padding.

Extractability — 24 points

Thirteen checks on whether the bytes a crawler receives contain your content. Main content present in raw HTML with no JavaScript wall. Semantic landmarks. No nosnippet, no max-snippet under 160, no noarchive or nocache — a legacy noarchive tag is an invisible Copilot kill switch. A clean 200 with a self-referencing canonical and no redirect hops. Alt text coverage, paragraph length, text-to-markup ratio, response weight and speed, declared language, viewport, and revalidation headers.

Entity clarity — 12 points

Ten checks on whether an engine can work out who you are well enough to name you: an Organization or LocalBusiness block carrying the legal name and contact details, sameAs pointing at at least two authoritative profiles, a substantive linked About page, visible contact details corroborating the schema, author bylines with credentials on articles, and brand-string consistency between your titles, footer and structured data.

Answer schema — 8 points

Eleven checks, deliberately low-weighted. A matched difference-in-differences analysis of 1,885 pages that added JSON-LD found no citation lift, and language models demonstrably extract facts from even invalid JSON-LD as plain text. So we do not reward schema for existing. We check that it parses at all — one syntax error silently voids every structured-data block on the page — that the page type is right, that published and modified dates exist and are honest, that publisher links to the Organization by @id rather than floating free, that any FAQ markup is word-identical to the visible text, and that the schema headline has not drifted away from the H1 and title.

What do you deliberately not score?

Two things, and both are reported instead.

  • llms.txt. Scored at zero. Across 137,210 domains, 97% of llms.txt files received no bot requests at all, and Google has said it will not support the format. We note whether you have one; it earns nothing.
  • Off-site factors. Branded mentions, review volume, Reddit and YouTube presence, Wikidata, knowledge-graph presence and cross-web contact consistency all matter more than most on-page work — and none of them are things your site controls. They arrive as a ranked action plan beside the score, never folded into it. Scoring what a customer cannot change is how an audit gets falsified by the customer's own dashboard.

See also: what actually correlates with AI citation and how a run is structured.

Primary sources

Everything asserted above traces to one of these. Operator documentation changes often; check the current version before relying on any of it.

Request an audit of your site

There is no fake "instant scan" button on this site. A real audit crawls your sitemap, fetches every page, and runs nine live user-agent probes against your edge — it takes minutes of machine time, not milliseconds, and it is run by a human who reads the result before you see it. Send the URL and you get the headline score plus your single largest blocker back, at no charge.

Request an audit See pricing

Frequently asked questions

Why is schema weighted so low?

Because the best available evidence does not support it as a growth lever. A matched difference-in-differences analysis of 1,885 pages that added JSON-LD found no citation lift, and language models demonstrably extract facts from even invalid JSON-LD as plain text. Our schema checks police validity and contradiction rather than rewarding presence.

Why do you not score llms.txt?

Across 137,210 domains, 97% of llms.txt files received zero bot requests, and Google has stated it will not support the format. We note whether you have one, and it earns nothing.

What is an evidence tier?

A label on every finding saying how strong the proof behind it is. Causal means a controlled experiment with a measured effect size. Measured means a large-sample observational correlation with a traceable source. Official means documented behavior from the engine operator. Mechanism means it follows from how retrieval demonstrably works, with no effect size claimed.

Do you score off-site factors like brand mentions?

No, deliberately. They matter more than most on-page work, but your site does not control them, and an audit score that includes factors a customer cannot change will be contradicted by that customer's own data. They arrive as a ranked action plan beside the score.

Which single check finds the most problems?

The multi-user-agent edge probe. A robots.txt linter is a commodity, but almost nobody probes the firewall layer where the 2025 and 2026 failures actually live.