Methodology v2026.09.9

How the audit actually works

Most AI-visibility audits are one prompt to one model, dressed up as analysis. This page explains what ours measures, how it's scored, what we refuse to claim — and a dated changelog of every methodology change, so you can see the audit is maintained, not fossilised.

Maintained by James Eltherington, founder.

Measured AI visibility

We don't infer your AI visibility from on-page signals — we measure it. The audit builds a set of questions your customers actually ask, puts them to ChatGPT, Perplexity, Gemini, Google AI Overviews, Grok and Claude, and records the answers: whether you appear, where you rank, and who gets recommended instead of you. The queries and results are in your report — you can re-run any of them yourself and check.

Engines don't give the same answer twice, so a single query proves nothing — that's why this is a structured probe across engines and queries, scored in aggregate, rather than one screenshot of one lucky answer.

How we read your site

We read the HTML your server returns and do not run JavaScript, because that is what the AI crawlers do: GPTBot, ClaudeBot, PerplexityBot and Google-Extended all fetch the raw page. Content that only appears after JavaScript runs will not appear in your audit, and it does not appear to them either. When a site is built that way we say so at the top of the readiness report rather than let a low score look like a mistake. The one exception is a bot wall: when a site answers with a challenge page we render the homepage in a real browser to get past it, and the report still records that the wall is there.

How scoring works

Your Readiness score is grounded in deterministic checks: measured signals from your live site, evaluated in code against a fixed rubric. The same site produces the same score, every run. On top of that, an AI analysis pass adds depth — but it is anchored to the measured signals, not free to improvise.

Your AI Visibility score comes from the measured probe above. The two blend into one Overall score across six categories:

We publish the categories, not the individual checks and weights — that's the part competitors would like us to hand over.

The QA gate & refund

Before a report reaches you it passes an automated QA gate: completeness of every section, presence and consistency of scores, and evidence backing the findings. A report that fails the gate is automatically retried; if it still can't pass, it is held for human review — and if we can't deliver, you get a full refund. You will never receive a report that didn't pass.

What we refuse to do

The GEO industry has a snake-oil problem. These are written rules enforced in our report pipeline — not marketing copy:

We run it on ourselves — current score: 40/100

Every methodology version is run against citeify.ai before it ships — same pipeline, same scoring, no special treatment. Our current published result (scored under v2026.08.2): Readiness 80/100, measured AI Visibility 0/100, Overall 40/100. Zero visibility means that across 120 recorded AI answers to buyer questions in our own niche, no engine cited citeify.ai. We publish that willingly: it is the honest starting point of a brand that is weeks old, and the number every future re-audit gets measured against.

The instrument itself is on display in the run history: two audits of the unchanged site scored an identical 76 Readiness — same site, same score — and after a set of fixes our own report recommended, the next run measured 82. Under v2026.08.2 the same site scores 80: the two-point move is the methodology tightening (schema and llms.txt recalibrated against published evidence — see the changelog below), not the site. When a score moves, we can tell you which one changed.

Read our full audit report (PDF) →

Methodology changelog

AI search changes monthly; an audit that doesn't is worthless. Every change to what we measure is versioned and dated here. Your PDF report carries the methodology version it was produced under.

v2026.09.9

September 2026 (update 9)

  • Full audits and Tracker runs now put every question to six engines: Grok and Claude join ChatGPT, Perplexity, Gemini and Google AI Overviews. The free check keeps the original four. The visibility score averages whichever engines answered, so scores from before this update remain comparable engine by engine.
  • The AI visibility score is shown as a band: the same run recomputed on each of its five samples, so a two-point wobble is not read as a move. When the question set has been edited, the change since the previous run is computed on the questions both runs asked.

v2026.09.8

September 2026 (update 8)

  • You can now steer the buyer questions we ask the AI engines. Tell us what matters in a few lines and we rewrite the set from it, or edit, pin, remove and add the exact questions yourself, on the dashboard for each domain. Up to ten are included; more can be added.
  • A first paid audit now waits for you to check its questions before it runs, and starts on its own after a day if you do nothing.
  • Changing the questions starts a new baseline: the trend line marks the run where the set changed, and each question in the table shows its own verdicts from earlier runs, so unchanged questions keep their history and edited ones start afresh.

v2026.09.7

September 2026 (update 7)

  • The PDF report is now rendered directly from the measured data, with no model in the loop. It carries the same content as the dashboard in the same order (overview, AI visibility, readiness, the fix), so the two can never disagree, and any past report can be re-issued in seconds under the current design.
  • The executive summary and quick wins written during the audit are now kept with the audit and printed in the PDF. Reports produced before this update are re-issued with a summary built from their scores.
  • The AI visibility and readiness sections each open with a short written summary of what was measured: which questions were lost and to whom, and which categories and findings matter most. Shown on the dashboard and in the PDF.

v2026.09.6

September 2026 (update 6)

  • The report is reordered around the question you actually asked. It now opens with the verdicts: every buyer question, a cited / mentioned / absent result per engine, and the competitors recommended instead. The readiness diagnosis follows as "why", and the findings, 30-day plan and coding-LLM briefs close as "the fix". Each report has its own page with deep-linkable sections.
  • Every question now maps to your own site: when an engine cited you, the report names the page it cited; when you were absent, it names the crawled page best placed to answer that question, or says plainly that nothing on the site does yet.
  • A new panel puts "built to be cited" (your citability readiness) next to "actually cited" (the measured citation rate across engines), with a one-line reading of the gap.
  • Full audits now ask each question five times per engine instead of three, so a verdict reflects a majority of samples rather than a lucky or unlucky answer.
  • Free checks now read sites behind bot-protection challenges with the same headless fallback the full audit uses.

v2026.09.5

September 2026 (update 5)

  • The £19.99 Quick Scan is retired and replaced by a free Visibility Check. The old scan graded your homepage foundations and never asked an AI engine anything; the free check runs the same 10 buyer questions the full audit uses, on the same four engines, once each, and shows you the verdict per question and the competitors recommended instead. Your email is confirmed before the check runs.
  • A free check and a later full audit of the same domain now ask the same questions, so the audit deepens the table you have already seen (three samples per question instead of one) rather than starting a new one.
  • The "ready" email now leads with the headline from your report, such as how many questions you were absent from and who was recommended in your place.

v2026.09.4

September 2026 (update 4)

  • A full audit no longer ships if its measured AI-visibility check could not run. Previously, if the answer engines were unreachable, the report would quietly go out without the visibility page and the headline score would fall back to readiness alone. Now the audit is held for human review and re-run, so every published full audit carries a real measurement.
  • Corrected what we claim to measure: live AI visibility is measured on ChatGPT, Perplexity, Gemini and Google AI Overviews. Bing Copilot has no queryable answer service, so we score readiness for it but do not present it as measured. Removed references to features that are not yet available to buy.

v2026.09.3

September 2026 (update 3)

  • The buyer-intent questions behind your visibility score are now tailored to what your business actually does and where it operates. A local or regional business is measured on the questions its real customers ask — including “near me” and place-specific searches — and against competitors at its own scale, not global market leaders it never competes with. This makes the visibility score, and the competitors it names, far more relevant and actionable.

v2026.09.2

September 2026 (update 2)

  • Our crawler is now a well-behaved, clearly identified visitor: it spaces out its requests so it never trips a site’s rate-based security, and it announces who it is so hosts can recognise it.
  • When a site’s security serves our crawler a “please wait, verifying…” challenge, we now read the real page by rendering it in a full browser — so good content behind a strict security wall is measured fairly instead of scored as empty. Crucially, the report still flags that the wall exists: AI engines like ChatGPT and Perplexity cannot pass these challenges, so a site behind one is effectively invisible to them, and we call that out as a critical issue to fix.

v2026.09.1

September 2026 (update 1)

  • Strengthened the crawl-integrity safeguards after a real-world miss: bot-challenge detection now recognises many more challenge-page variants, checks the exact pages our crawler was served (catching intermittent challenges), and retries once before concluding a site cannot be read. Found by re-auditing a site we knew was healthy.

v2026.09

September 2026

  • Added crawl-integrity safeguards: if a site serves our scanner a bot challenge instead of its real content, or a score moves implausibly far between audits, the audit is now held for human review rather than published. No one should ever receive a confident report scored against a security wall.
  • Anything our crawler is blocked from reading is now always reported as "could not be verified", never as missing from your site.
  • If an AI engine is unavailable while your audit runs, the report now says so explicitly; your visibility score is averaged only over the engines that actually answered.

v2026.08.4

August 2026 (update 4)

  • Reports now show the actual buyer-intent questions we put to each AI engine, with the outcome per engine for every question (cited, mentioned, or absent) and which competitors were recommended instead. The measurement itself is unchanged: this data was always collected, and it is now published in full so you can verify any result yourself.

v2026.08.3

August 2026 (update 3)

  • Audits now always include the site’s About/team page when judging authorship and expertise — a named expert who exists but isn’t surfaced is now reported as “surface it”, not “missing”. Found by re-auditing our own site.
  • Hardened finding categorisation so every issue files under the category that owns it (a security finding can no longer appear under Content Quality).
  • FAQ guidance sharpened: visible question-and-answer content is what earns citations; FAQ markup alone is no longer recommended as a fix following Google’s May 2026 rich-results retirement.

v2026.08.2

August 2026 (update 2)

  • Modernised AI-crawler coverage to the current 2026 user-agent landscape — including the newer Anthropic, Mistral and DuckDuckGo crawlers — from a single canonical list.
  • Added a CDN-level access check: some CDNs now block AI crawlers by default even when robots.txt looks open. The audit now tests how AI crawlers are actually served.
  • Down-weighted llms.txt to an informational check, in line with the evidence that major engines don’t consume it (it retains minor value for documentation sites).
  • Recalibrated structured-data guidance to current evidence: schema is scored for entity clarity and machine readability, not as a claimed citation booster; FAQ rich-result guidance updated for Google’s May 2026 deprecation.
  • Corrected platform facts across the audit (Google-Extended vs AI Overviews controls, ChatGPT’s own search index, Microsoft Copilot naming) and added machine-readable freshness detection.

v2026.08

August 2026

  • Unified the scoring model so every layer of the audit — analysis, report and PDF — is produced from a single documented rubric.
  • Visibility probe now samples each question multiple times per engine (AI answers vary run to run — one sample is a coin flip) and re-uses a domain’s question set between audits, so score changes measure visibility movement, not question drift.
  • Tightened our evidence policy: every claim in a report must trace to a measured signal or published research. Estimated traffic and revenue projections are now prohibited in our reports.
  • Recalibrated research citations against the current academic literature on generative engine optimisation (2024–2026).
  • Every report PDF now carries its methodology version, tied to this changelog.

v2026.07

July 2026

  • Added the measured AI-visibility probe: your customers’ real questions are put to ChatGPT, Perplexity, Gemini and Google AI Overviews, and the answers recorded.
  • Moved readiness scoring to deterministic, code-computed checks grounded in measured site signals — the same site now always produces the same score.
  • Added multi-page citability analysis and automated QA gating on every report.

See it applied to your site.

About 25 minutes from domain to report. QA-gated, refund-backed.