Skip to content
Status on 31 July 2026: in build, not open yet. Measurement has not opened, nothing on this page can be bought today, and CitedProof has no customers — so there are no logos or testimonials here and there will not be any invented ones. Everything below describes what a Proof Pack contains when it opens. What you can do today: leave a domain and we measure it free on day one, and read the methodology, which is already published in full — including the sizes of change this test cannot detect.
AI answer measurement · citedproof.com

Your buyers are asking AI about your category today. Here is what it answers.

CitedProof runs your real buyer questions through ChatGPT, Google AI Overview and Google AI Mode, shows you the sentences that decide whether you get named, hands you the exact fix written out in full, and re-measures on day 30 to tell you whether it worked.

Free at launch, one domain. The paid Proof Pack is $29, charged once. No card to see the access report.

Specimen · output format 1 prompt · 3 engines

“best mid-range espresso machine for a small café”

ChatGPT Not mentioned
“For a small café I’d look at the Rocket Appartamento, the Lelit Bianca, or the Profitec Pro 500.”
3 brands named · yours absent · 6 sources cited · none yours
Google AI Overview Mentioned · position 4
“Other options include your brand, though pricing is not listed publicly.”
Cited source: a forum thread, not your site
Google AI Mode No reading

The surface returned no answer for this prompt. This run is excluded from the denominator — it is not counted as an absence, and it does not become a zero.

Excluded · reason recorded · re-runs on the next cycle
Illustrative rendering. Not a customer measurement.
The one promise

We tell you what the engines said, what to change, and whether the change worked — or we tell you we could not tell.

That last clause is the whole company. Every other tool in this category prints a number no matter what came back. When too few engines answer, or too few runs came back usable, CitedProof prints NOT ENOUGH DATA and the stored reason — which of the two conditions failed, and by how much. A zero appears only when the denominator is full and your brand genuinely did not appear.

REAL CHANGE NO CHANGE NOT ENOUGH DATA

Three verdicts. There is no fourth, and none of them is a grade out of 100.

How it works

Three steps. You do the second one.

CitedProof measures, specifies and verifies. The work of shipping the fix stays with you or your developer — which is why every finding arrives already written, not described.

  1. Measure

    We read your site, write the buyer questions, and run them

    Paste a domain. We crawl it, build the prompt set from what you actually sell, and run every prompt against ChatGPT, Google AI Overview, Google AI Mode and the Google result page. Before any of that we check 22 AI crawler tokens — OAI-SearchBot, ChatGPT-User, GPTBot, Claude-SearchBot, Claude-User, ClaudeBot, PerplexityBot and the rest — at robots.txt, at your CDN, and at the HTTP layer. Twenty of them we fetch as; the two Google and Apple publish no user-agent for are read from robots.txt and labelled that way.

  2. Fix

    Every finding ships finished, not described

    Not “add pricing to this page” — the rewritten block, ready to paste. Not “unblock the crawler” — the exact robots.txt lines and the curl command that proves it returns 200. Each one names who does it, how long it takes, and which prompts we will re-measure to check it.

  3. Prove

    Day 30, we re-run the same prompts and apply a statistical test

    Same prompts, same engines, same method. Baseline against re-measurement with a two-proportion test and a Wilson confidence interval — never a raw before-and-after subtraction. You get one of three verdicts and the numbers that produced it: both rates, both intervals, both sample sizes and the p-value.

What you keep

Four artefacts, exportable, yours

A dashboard you stop paying for is a dashboard you lose. Everything CitedProof produces is a document or a file you can download and keep — the report, and a JSON export carrying your prompts, every measurement snapshot with its denominators, and every finding.

AI Access Report

A binary verdict per crawler: reachable or blocked, and by what — robots.txt, WAF, CDN rule or HTTP status. OpenAI states plainly that sites opted out of OAI-SearchBot are not shown in ChatGPT search answers [3]. This check is close to free to run and is the only finding on this list that can be true within a week of fixing it.

Includes your exposure to Cloudflare’s 15 September 2026 change, which begins blocking crawlers classified as Training and Agent by default on newly onboarded domains [4].

Action Packages

One per finding, and each contains all five parts or we do not publish the finding at all: the defect with its evidence; why it matters with an evidence grade and source; the finished artefact to paste; the execution plan with owner and duration; and the proof plan naming the prompts we will re-measure.

Evidence grades run A to D. A grade-D finding is labelled grade D on the page you hand your client.

Anti-Waste Rulebook

The list of things you are being sold that the evidence says do nothing — checked against your own site, so it names what you are spending on. Most reports of this kind are longer than the fix list.

Nobody is assigned to these. That is the point of the tier.

The verification verdict

On day 30: REAL CHANGE, NO CHANGE, or NOT ENOUGH DATA, with both rates, both confidence intervals and the number of prompts behind them. NO CHANGE is a normal, publishable outcome and we do not dress it up — it means we could not detect a change at this sample size, not that your fix failed.

Exactly what that sample size can and cannot detect is published as a table on the methodology page. Your prompts, findings and measurement snapshots export as JSON from the dashboard at any time, including after you cancel.

The anti-waste stance

We start by telling you what to stop paying for

Three of the most widely sold “AI visibility” tactics have been measured. Two do nothing and one measured negative. We ship this list before we ship a single recommendation, because a tool that only ever adds to your to-do list is selling you work, not judgement.

97%

of published llms.txt files received zero requests during May 2026. Of the 137,210 domains studied, 28% publish one. The largest category of fetcher was SEO audit tools, at 21.7% — not AI bots.

Ahrefs, n = 137,210 domains, May 2026 [1]
−4.6%

AI Overview citations after pages added JSON-LD schema for the first time — statistically significant, and negative. AI Mode +2.4% and ChatGPT +2.2% were both within noise. Difference-in-differences, 1,885 treated pages against ~4,000 matched controls.

Ahrefs, n = 1,885 treated / ~4,000 control, Aug 2025 – Mar 2026 [2]
0.194

Spearman correlation between number of site pages and AI visibility — the weakest of every signal measured across 75,000 brands. Publishing more pages is the most expensive lever with the least evidence behind it.

Ahrefs Brand Radar, n = 75,000 brands; correlation, not causation [5]
Google says it in its own words. “You don’t need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search… Google Search ignores them.” — Google Search Central, 15 May 2026 [6]
Why a verdict and not a delta

The surface moves on its own. A before-and-after number hides that; a test does not.

Three independent parses of the same 10,000 keywords on Google AI Mode, on the same day, from the same location, agreed on only 9.2% of exact URLs across all three runs [7]. On a surface that unstable, “your visibility went from 18% to 24%” is a coin toss written in a serious font.

So we do not sell you repetitions either. In a 12,933-response variance study, going from 5 repeats to 10 reduced variance by 0.00030 — while adding languages or engines was roughly fifteen times more effective per unit of budget [8]. Daily refresh is a pricing feature, not a statistical one.

What we do instead: fix the prompts, fix the engines, fix the method, and compare the two proportions with a pooled two-proportion z-test at p < 0.05, printing a Wilson interval beside each rate. One verdict, pooled across every engine — a per-engine verdict is never the headline, because four verdicts at that threshold are wrong about one time in five. If the test does not clear the bar, the answer is NO CHANGE, which means we could not detect a difference at this sample size. If either side is too thin to compare, the answer is NOT ENOUGH DATA, and we name the condition that failed — the engines that did not answer, or the usable runs we are short of.

What a verdict looks like

Baseline · 12 prompts × 3 engines 36 runs
Usable answers (the denominator) 11
Engines that answered 1 of 3
Mentions 2

A run where the engine returned nothing is excluded from the denominator. It is never counted as an absence, because “the engine was down” and “you were not recommended” are different facts.

NOT ENOUGH DATA

The stored reason is printed verbatim: “only 1 of 3 engines answered; 2 are required before a cross-engine rate is meaningful”. Eleven usable answers would have cleared the volume floor on their own — one engine agreeing with itself is not the cross-engine measurement we sell. Both conditions are set out on the methodology page.

Pricing

Two prices. Both include the day-30 verdict.

No credits, no per-brand meter, no seat charge, no annual lock. The free AI Access Report needs an email and a domain — not a card.

Proof Pack

One brand. You own it, you fix it.

$29 once

  • Baseline across ChatGPT, Google AI Overview, Google AI Mode and Google SERP
  • Crawler-access verdict across 22 AI crawler tokens
  • Every finding as a full Action Package with paste-ready copy
  • Anti-Waste Rulebook checked against your own site
  • Day-30 re-measurement and verdict
Join the launch list

The free report comes first. Nobody should pay $29 before seeing what we found for nothing.

Agency

You manage brands that are not yours.

$129 / month

  • 7 client workspaces, flat — the price does not rise as you win clients
  • White-label: your logo, your colour, your name on every PDF
  • Recurring re-measurement and a verdict per shipped action
  • One login per account today, and seats are never a billing line. Cancel from the dashboard, no email required
  • The published methodology page, so your client can check our work
See the agency case

Roughly $18.43 per client brand per month at 7 brands.

Start here

Get your AI Access Report

CitedProof is in build. Leave a domain and we run your access report the day measurement opens, and email it once. We do not run a drip sequence, and there is no second email unless you reply.

Questions

Answered plainly

What does CitedProof actually measure?

Your real buyer questions, run against ChatGPT, Google AI Overview, Google AI Mode and the Google result page. For every run we record whether your brand was mentioned, whether it was recommended, its position, the sentiment, and which sources the answer cited. Separately we check 22 AI crawler tokens — including OAI-SearchBot, ChatGPT-User, GPTBot, Claude-SearchBot, Claude-User, ClaudeBot and PerplexityBot — against your site. Twenty are fetched as themselves; Google-Extended and Applebot-Extended have no published user-agent, so they are read from robots.txt and the report says so.

We do not query Google’s Gemini grounding API. Google’s terms forbid analysing grounded results, so that engine is not in our set and will not be added. The methodology page quotes the clause.

Why would you refuse to show me a score?

Because a score built on too few valid runs is a fabrication, and it is the specific failure this product exists to avoid. If engines did not answer, those runs are excluded rather than counted as an absence, and the report prints NOT ENOUGH DATA along with the reason — which engines did not answer, or how many usable runs it is short of.

A zero appears only when the denominator is full and your brand genuinely did not appear. That zero is real, and we show it.

Do I need an llms.txt file?

No. Across 137,210 domains, 28% publish a valid llms.txt and 97% of those files received zero requests during May 2026 [1]. Google stated on 15 May 2026 that it ignores such files [6]. It is in our Anti-Waste Rulebook and we will not generate one for you.

Is being blocked to AI crawlers really fatal?

It is a rate lever, not an on/off switch, and we say so. Rate-normalised, domains blocking GPTBot showed a citation propensity of 0.003 against 0.417 for non-blockers across 1,058 domains — the authors state plainly that the data cannot prove causation [9].

The honest counterweight: a separate study of 4M citations found sites blocking GPTBot still appeared in 88.2% of the prompts examined [10]. So we sell unblocking as removing a handicap, never as a guarantee of citation. Both studies are on the methodology page.

Will schema markup help me get cited?

The best causal evidence available says no: AI Mode +2.4% and ChatGPT +2.2% (both noise), AI Overviews −4.6% and significant, across 1,885 first-time JSON-LD pages against ~4,000 matched controls [2]. The widely quoted “cited pages are 3× more likely to have schema” is confounding, not cause.

Keep schema for rich results if you want them. It is not a lever on AI citation, and we will not bill you to add it.

How do you charge, and can I get a refund?

The Proof Pack is $29 charged once, through TAP Payments. The agency plan is $129 per month and we do not hold your card: a payment link is sent one day before each renewal and you pay it yourself, so nothing is ever taken automatically and doing nothing is how you stop. Cancel from the dashboard in one click. Card details never touch our servers.

If a Proof Pack fails to produce a completed measurement, we spot it and refund you without being asked — a person here approves the payment, usually within one business day — and we do not ask why. The conditions are written out on the refunds page.

Who is behind this?

CitedProof is built and operated by KIRA Holdings. It is a small operation by design: no sales team, no call booking, published email support at hello@citedproof.com.

Sources for every number on this page

  1. Ahrefs, “We studied 137,210 domains to see if llms.txt is being used” — 28% publish a valid file; 97% received zero requests in May 2026. ahrefs.com/blog/llmstxt-study
  2. Ahrefs, difference-in-differences on schema and AI citations — 1,885 treated pages, ~4,000 matched controls, Aug 2025 – Mar 2026. ahrefs.com/blog/schema-ai-citations
  3. OpenAI, crawler documentation — “Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers.” developers.openai.com/api/docs/bots
  4. Cloudflare’s change to default crawler handling from 15 September 2026, reported secondarily; we label this Grade C until Cloudflare’s own policy page states it. ppc.land report
  5. Ahrefs Brand Radar correlation study — 75,000 brands, Spearman; number of site pages 0.194–0.205, the weakest signal measured. Correlation, not causation. ahrefs.com/blog/ai-brand-visibility-correlations
  6. Google Search Central, AI features and your website, 15 May 2026. developers.google.com/search/docs/fundamentals/ai-optimization-guide
  7. SE Ranking, Google AI Mode study — 10,000 keywords, 9,734 responses, three independent parses on 2026-06-20; 9.2% exact-URL agreement across all three. seranking.com/blog/ai-mode-research
  8. arXiv 2607.13304 — 12,933 responses, 20 brands, 3 engines; variance reduction from 5→10 repeats = 0.00030. arxiv.org/html/2607.13304
  9. cloro.dev — 1,058 domains, citations per Google-organic appearance; GPTBot blockers 0.003 vs 0.417. Authors’ caveat: “The data can’t prove causation.” cloro.dev/research/ai-crawler-blocks
  10. BuzzStream — 4M citations, 3,600 prompts; sites blocking GPTBot still appeared 88.2% of the time. buzzstream.com/blog/news-block-ai-bots-citations

No customer names, logos or testimonials appear on this site. We have not shipped a paying customer yet, and inventing one would be the same lie we are selling against. When we have results to publish, they will be published with sample sizes and dates like everything above.