See what makes AI recommend your business.
Lesli.com -- AI Visibility & SEO

Method and baseline / results pending

Measuring AI Citation One Sub-Query at a Time

Every AI visibility tool I can find measures whether a brand gets mentioned. None measure the thing that decides it: which of the engine’s own fan-out sub-queries your passage wins. Looking one level down turned up something I was not expecting, which is that on one sub-query the engine is confidently wrong and the authority that could correct it is not cited at all.

Lesli Rose · 14 July 2026 · Baseline n=13 sub-queries, one site. Verdict due 28 August 2026.

This is not a results post. It is a method and a baseline, published before the outcome is known, because a design flaw found now costs nothing and a design flaw found on 28 August costs the whole experiment. The things I most want attacked are at the bottom.

01 / The gap

Prediction and Measurement Exist. Nobody Has Joined Them.

Google confirms that AI Mode uses query fan-out, decomposing one question into concurrent sub-queries across subtopics. Independent measurement puts Google AI Mode at roughly nine sub-queries per prompt, and finds that about 95% of those sub-queries carry no recorded search volume. Ahrefs finds only 12% to 38% of AI citations come from the head query’s organic top ten.

So the unit of competition is the sub-query, and keyword tools structurally cannot see it. Two halves of a solution already exist:

Qforia (iPullRank)

Predicts what the fan-out sub-queries probably are. It does not check who is cited for them.

Profound, Peec, Otterly, Brand Radar, the Semrush AI toolkit

A category that raised north of $300M. Every one of them measures citation at the prompt and brand level, then aggregates to share of voice.

Nobody closes the loop: predict the fan-out, observe real citation per sub-query, find the sub-queries you do not answer, write passages for them, and re-verify the lift against a control. That loop is what I am testing.

02 / Method

Five Steps, and One of Them Is the Whole Point.

The steps are numbered because they are genuinely sequential and the control has to be fixed before any content moves.

01

Predict the fan-out

Decompose the head question into sub-queries across thirteen types: reformulation, related, implicit, comparative, entity-attribute, generalization, procedural, temporal, credibility, local, negative or risk, follow-up, disambiguation. The count is judged by complexity, not fixed.

Two sanity checks. If a keyword tool would find most of them, they are too head-shaped and I regenerate. If none of them is uncomfortable for the business, the fan-out is sanitised and useless, because real engines do not protect brands.

02

Observe who is actually cited

Run every predicted sub-query against live Google AI Mode and record the real cited sources. This is the step nobody else runs, and it is what turns a guess into a measurement.

Cost is about USD 0.004 per sub-query, so a full map runs to a few cents.

03

Build a coverage matrix

One row per sub-query: who is cited today, are we, do we even have a passage, and a verdict.

The most valuable row is never the missing page. It is covered but not cited, where a passage exists and still loses.

04

Fix the passage, not the page

Citation is decided at passage level, by comparing an embedding of a passage against the answer claim (Google patent US11769017B1). So the deliverable is an answer block, not an article.

Front-loaded answer, self-contained enough to survive being lifted out of the page, and carrying a cited statistic or an authoritative quotation, which are the only two levers with peer-reviewed experimental support behind them (the GEO paper, KDD 2024).

05

Hold a control, then re-verify weekly

A set of real gaps is deliberately left untreated. Semrush ran a fan-out experiment without a control, ChatGPT cut citations platform-wide mid-test, and their treated pages read as a failure.

Without a control you cannot separate your lift from platform drift, and any claim you make afterwards is unfalsifiable. Weekly re-verification, not daily, because daily is noise. A trend line makes platform shifts visible as steps that move the control too.

03 / Baseline

One Registry, Thirteen Sub-Queries, 46%.

Test site is a breed registry I own, so I can ship content against it without asking anyone’s permission. Head question: “how do I register an American Bulldog litter”. Target: abra1st.com.

13Sub-queries
6Cited
46%Citation rate
$0.04Run cost

A brand-level tracker reports that as “46% AI visibility, healthy.” The sub-query view says something entirely different, and it only becomes visible once you sort the rows.

Sub-query (predicted)TypeCited today (observed)Us?
how do I register an American Bulldog litterheadnkc, abra1st, youtubecited
American Bulldog registry options comparedcomparativeakc, nkc, ukc, abra1stcited
is ABRA a legitimate American Bulldog registrycredibilityabra1st, pedigreedatabasecited
how much does it cost to register an American Bulldog litterentity-attrnkc, akc, abra1stcited
UKC vs NKC for American Bulldog registrationcomparativeukc, abra1st, nkccited
what is a pedigree certificate for an American Bulldogrelatednkc, abra1st, akccited
what registry should I use for American Bulldog puppiescomparativeukc, akc, adoptapetnot cited
what documents do I need to register a litter of puppiesproceduralakc, royalkennelclub, ckcusanot cited
how long does American Bulldog litter registration takeentity-attrnkc, abkcdogs, akcnot cited
do I need DNA testing to register an American Bulldog litterimplicitakc, embarkvet, redditnot cited
can I register an American Bulldog with the AKCimplicitakc, quora, redditnot cited
American Bulldog registration scams to avoidnegativeakc, youtube, telegraphnot cited
how do I register American Bulldog puppies without papersnegativeyoutube, justanswer, akcnot cited

The registry is cited for its identity and invisible for its process. Every sub-query it wins is about who it is, how it compares, what it costs, what a pedigree certificate means. Every sub-query it loses is operational: what documents, how long, is DNA required, what goes wrong, which registry should I actually pick.

Those operational questions are the ones a breeder asks before choosing a registry, and at that moment the registry is not in the room. The National Kennel Club and the AKC take those instead, on nine sub-queries each, because they publish deep procedural documentation and this registry does not.

The sharpest row is “what registry should I use”. That is the decision itself. The brand is named in the answer text and still not cited. It is losing the recommendation while being mentioned, which is a distinction a brand-mention tracker cannot draw at all.

04 / The part I did not expect

The Engine Is Confidently Wrong, and the Authority Is Absent.

One sub-query in the fan-out was “can I register an American Bulldog with the AKC”. Google AI Mode opens its answer, in bold, with this:

“Yes, you can register an American Bulldog with the American Kennel Club (AKC)”

The AKC’s own Foundation Stock Service page, which the answer cites, says this:

“FSS® breeds are not eligible for AKC registration.”

The American Bulldog is an FSS breed. FSS is a recording service, not registration. The answer goes on to soften itself, but the lead sentence, the one that is bolded and front-loaded and therefore the one most likely to be lifted, contradicts the source it is standing on. A breeder acting on it could pay for something that is not what they think it is.

I did not catch this. I looked at the FSS breed list, saw the American Bulldog on it, decided the AI was basically right, and moved on. The registrar who has run the breed’s archive since 2005 read one sentence and said “I don’t think you can register your dogs with the foundation stock registry.” She was right and I was wrong.

Two things follow. First, this is the Tow Center finding made concrete: eight engines, 200 tests, over 60% of citations wrong. Second, and more useful: the breed’s actual registry is not cited anywhere in that answer. The error and the authority gap are the same problem. The entity that could have corrected the record was not in the room.

That reframes what a citation gap costs. It is not only lost traffic. It is the engine confidently answering in your domain, wrongly, without you.

05 / Limits

What I Cannot Claim.

The fan-out is predicted, not observed.

No public API exposes Google's real internal sub-queries. Prediction and observation stay in separate columns and I never merge them.

n=1 site, 13 sub-queries, one head question.

This is a pilot, not a study.

A parser bug nearly cost me the competitor data.

AI Mode embeds most citations as inline markdown links in the answer text, not in the structured references array. My first pass read only the array and undercounted competitors badly. Fixed before these numbers were taken. I mention it because anyone building this will hit it.

I killed the right headline for the wrong reason, and had to be corrected by a human.

I assumed AKC does not register the breed, found it listed in the Foundation Stock Service, decided I was wrong, and dropped it. The registrar knew FSS is a recording service and not registration. Domain expertise caught what the method missed. Whatever this tool becomes, it does not replace the person who knows the field.

The evidence base under all of this is thin.

One peer-reviewed controlled study exists (GEO, KDD 2024). Two independent tests, from Ahrefs and SearchVIU, find schema markup does not lift AI citation, and Google's own May 2026 guide says structured data is not required, content should not be pre-chunked, and llms.txt is not needed. Most of what the field sells is folklore with a percentage attached to it.

Peer review

The Things I Want You to Attack

01

Is a predicted fan-out good enough to measure against?

If my thirteen sub-queries are not the ones Google actually issued, I am measuring citation on a fictional surface. Is there a validation step you would run that I have not thought of?

02

Is the control set doing the job?

Three untreated gaps against four treated ones, re-verified weekly for 45 days. Is that enough to separate a real lift from AI Mode's own volatility, or is the noise floor higher than I think?

03

Is the identity-versus-process split real, or is it an artifact of how I wrote the sub-queries?

This is the finding I am least sure of and most excited by, which is usually a bad sign.

04

Does a citation gap on a sub-query where the engine is factually wrong deserve its own severity tier?

Being absent is one problem. Being absent while the engine invents an answer inside your domain is a different one, and I do not have a good way to score it yet.

If it fails, I will publish that too. A method that can only succeed is not a measurement.

Citation data observed via the Google AI Mode endpoint on 14 July 2026. Sources named inline. Verdict, including the failure case, due 28 August 2026.

Which Sub-Queries Are
Answering Without You?

The same method runs against any business. The visibility report shows what AI systems say about you, who they cite instead, and where the gap sits.

Run Your Visibility Report

Related: The Entity Anchor Method