Competitor AI citations are the brands and sources an AI assistant names when it answers the prompts your buyers actually type. You measure them by freezing a prompt set, running it on each engine on a schedule, logging every brand named in every run, and dividing your run count by the total across all brands. That ratio is your share of citation. The share is measurable. The cause behind it is not.
Most teams arrive here from a screenshot. Someone asks an assistant for a shortlist, a rival is named, you are not, and the screenshot lands in a channel by lunchtime. One answer is an anecdote. What follows is the method we use at DevCommX to turn it into a number a board can read: the prompt set, the logging schema, the arithmetic, and the honest limits. Measuring your own visibility is a different job, covered in our guide to measuring LLMO and AI visibility.
The short answer: citation share is measurable, citation cause is not
An assistant's answer is a sample, not a record. The same prompt can return different brands on a second run, in a different mode or locale, or for a signed in user. The only defensible unit is a distribution measured across many runs, and a single screenshot proves only that the answer was possible.
First party data behaves the same way. Google's generative AI performance report reports impressions from AI Overviews and AI Mode, and it carries impressions only, with no clicks, CTR, position or query data. The follow up rule sits in Search Console's definition of impressions, position and clicks: a follow up question inside AI Mode counts as a new query with its own impression, position and click data. Google introduced those reports in June 2026. The surface is real and countable, and it still does not say why one brand was chosen over another.
So separate the two questions permanently. Question one is descriptive: across a fixed prompt set, how often is each brand named, and what share of all naming events does each hold. That is answerable to a stated precision. Question two is causal: why was the rival named. That is not answerable from outside the model. Measure question one, then form hypotheses about question two and test them as changes.
Why your own visibility number tells you almost nothing on its own
Suppose you are named in 15 percent of runs. That is meaningless until you know what filled the other 85 percent. If no brand at all is named in most of them, the category is not being answered with shortlists and your problem is prompt selection. If four rivals split them, you have a competitive gap. If one rival holds most of them, you have a single displacement problem with a named owner.
Relative measurement is also more robust here. Engines change models, retrieval and interface without notice, so an absolute presence rate moves for reasons unrelated to you. A share computed on the same run set moves both numerators together, so a change in your share is more likely to mean something. That is the argument for tracking competitor AI visibility alongside your own rather than instead of it.
Do not back into share from traffic. Pew Research Center found users clicked a search link in 8 percent of visits with an AI summary against 15 percent without, and clicked an in summary source in only 1 percent of visits. Assistant referral data is therefore a thin, biased sample of citation events. Keep it as its own metric using the method in tracking ChatGPT referral traffic in GA4, and never use it as a proxy for share.
Building the prompt set that represents a real buying decision
The prompt set is the instrument, and a sloppy set gives a precise number about the wrong thing. Build it in five families. Category definition prompts ask what a category is and who operates in it. Shortlist prompts ask for recommendations, which is where who does ChatGPT recommend actually gets answered. Comparison prompts name two vendors and ask for a verdict. Objection and pricing prompts ask cost or weakness. Implementation prompts ask how to do the job, which is where documentation gets cited.
Three rules keep the set honest. Never put your own brand in a shortlist prompt, because naming yourself guarantees the mention you are measuring. Write prompts as full sentences, since Pew's data shows AI summaries appeared for 8 percent of one or two word Google searches but 53 percent of searches of ten words or more. And freeze the wording, versioning any change instead of editing in place.
Size the set to your decision, not to a round number. We start at thirty to sixty prompts for one ICP and one category, which is a working rule rather than a published standard, with separate persona variants where the buying committee genuinely differs. Pricing prompts deserve their own attention because they retrieve differently, as we cover in optimising pricing pages for AI search.
What to log for competitor AI citations: per prompt, per engine, per competitor
Five flat tables are enough, and a spreadsheet is a legitimate implementation. The point is that every share you report can be recomputed from these rows by someone who was not in the room. A number that cannot be traced to a run_id is an assertion, not a measurement.
Three conventions matter more than the schema. We run each prompt at least three times per engine in a clean session, which is our house rule rather than a published standard, because one execution is one draw from a distribution. Record engine_mode explicitly, since search mode, reasoning mode and plain chat are different measurements that must never be pooled. And count a brand once per run however often it repeats, because repetition within one response is formatting, not a second observation.
The roster table carries the weight nobody expects. Keep legal name, product names and common misspellings as aliases for every competitor, plus an other bucket for brands outside it. That bucket is usually the most interesting part of the exercise, because it is where a new entrant appears months before anyone adds them to a battlecard.
Calculating competitor citation share
Define the terms once, in writing, and never vary them. A run is one execution of one prompt on one engine. Let R(e) be the runs on engine e in the period, and M(b,e) the runs in which brand b was named at least once. Presence rate is P(b,e) = M(b,e) divided by R(e). That is where most teams stop, and it is the number that cannot be compared across periods when engines change.
Citation share needs a denominator built from brands, not runs. Let T(e) be the sum of M(b,e) across every roster brand plus the other bucket. Citation share is then S(b,e) = M(b,e) divided by T(e). Because T(e) is the sum of the numerators, shares add to 100 percent across brands by construction, even though several brands can appear in one run. That is what makes the number readable in a deck.
An illustrative worked example, with invented figures used only to show the arithmetic. Take 40 prompts run 3 times on one engine, so R = 120. Suppose you are named in 18 runs, Rival A in 54, Rival B in 33, and brands outside the roster account for 61 observations. Then T = 18 + 54 + 33 + 61 = 166. Your citation share is 18 divided by 166, or 10.8 percent. Rival A sits at 32.5 percent, Rival B at 19.9 percent, and the other bucket at 36.7 percent. Your presence rate is 18 divided by 120, or 15 percent, against Rival A at 45 percent.
Two extensions earn their keep. For a cross engine number, either sum numerators and denominators across engines, or weight each engine's share by a factor you can defend in writing, such as your own referral mix. Never weight by guessed market share. Second, build a head to head table on the same runs: Rival A and not you, you and not Rival A, both, neither. The cell where a rival appears without you is the addressable gap.
Establish a noise floor before interpreting anything. Run the identical prompt set twice in one week with nothing changed, and take the spread between the two share numbers as your measurement noise. Report that spread beside every share and refuse to read any movement smaller than it. This is what stops AI citation tracking becoming a monthly ritual of explaining variance as strategy.
Reading the result: the four patterns worth acting on
Pattern one, absent everywhere while rivals cluster on one source type. Your sources table shows it: most citations point at a handful of review pages or roundups you do not appear on. The work is source coverage, not on site content, which is why brand mentions behave differently from backlinks in AI search. Pattern two, present only through your own domain. Every citation is owned_by_brand. Models are quoting your marketing back at buyers with nothing independent corroborating it, which breaks the moment a comparison prompt runs.
Pattern three, present but never recommended. Your brand_role column reads compared or aside. You are in the consideration set and losing the verdict, which is a proof and positioning problem rather than a visibility one, and the fixes resemble those in how to get cited by ChatGPT. Pattern four, a clean split by engine. Present on one engine and absent on another usually points at access or indexation, not content.
Pattern four has a concrete checklist. Google states a page must be indexed and eligible to be shown with a snippet before it can appear as a supporting link in AI features. OpenAI documents separate crawlers: OAI-SearchBot and GPTBot are independently controllable in robots.txt, while ChatGPT-User covers user triggered fetches that OpenAI says robots.txt rules may not apply to. Anthropic publishes its own crawler guidance. A robots.txt rule written two years ago is a common and fixable cause of a one engine blackout. For the engine specific angle, see how to get cited by Perplexity in B2B.
Where the tools help with competitor AI citations, and where they stop
Tooling is good at the mechanical parts: scheduling runs, executing at volume, matching brand strings against a roster, storing evidence, rendering a chart. Past a few hundred runs a month, doing that by hand wastes a person. The category and its tradeoffs sit in our LLMO checklist and tooling notes, and this post does not rank products.
Where tools stop matters more. A vendor's default prompt set measures a category, not your deal, so it reports a share that is precise and irrelevant unless you supply your own prompts. Role labelling, the difference between recommended and merely compared, is a judgement call automation only approximates. The other bucket needs a human. Modes drift, and a tool that silently changes which mode it queries hands you an artificial step change. One rule covers it: export the raw rows monthly, because a share you cannot recompute is not a measurement you own.
Reporting it upward without overclaiming
Report five things and nothing more. The share per engine, with run count and date range. The noise floor from the repeat run. The head to head cell against your most cited rival. The pattern you diagnosed. And the dated change you made, so the next period has something to attribute against. A share without a run count has no error bars, and executives will read it as far more solid than it is.
Say the non claim out loud. You are not claiming pipeline attribution from citations, and Pew's 1 percent in summary click rate is the reason. Put Google's generative AI performance report beside your panel and label the difference: Search Console counts your impressions, your panel measures relative share across brands. If budget is the real conversation, our SEO and AEO budget model frames it, and the category boundaries sit in LLMO vs SEO vs GEO vs AEO.
Measure Competitor AI Citations With DevCommX
If a rival keeps appearing where you do not, the first deliverable is a measurement you can defend, not a content sprint. DevCommX builds the prompt set with your sales team, stands up the logging schema above, runs the baseline and the repeat that sets your noise floor, and hands over the raw rows. That is the starting point of our AEO services engagement. For what the same operating model produces elsewhere: 40+ qualified demos in ~6 weeks, from our AI SDR work on a fully scoped programme with a defined ICP. Book a call and bring the screenshot that started this.
References
- Google Search Console Help, Generative AI Performance Report, source for Search Console impressions from AI Overviews and AI Mode, a report that carries impressions only.
- Google Search Console Help, What are impressions, position, and clicks?, source for an AI Mode follow up question counting as a new query with its own impression, position and click data.
- Google Search Central Blog, Introducing Search Generative AI Performance Reports, source for the June 2026 launch of those reports.
- Pew Research Center, Google Users Are Less Likely to Click Links When an AI Summary Appears, source for the 8 versus 15 percent click figures, the 1 percent click rate on in summary sources, and the summary trigger rates by query length.
- Google Search Central, AI Features and Your Website, source for the requirement that a page be indexed and snippet eligible to appear as a supporting link in AI features.
- OpenAI, Overview of OpenAI Crawlers, source for the distinct roles of OAI-SearchBot, GPTBot and ChatGPT-User, for OAI-SearchBot and GPTBot being independently controllable in robots.txt, and for OpenAI's note that robots.txt rules may not apply to ChatGPT-User fetches.
- Anthropic Help Center, Anthropic Crawlers and How Site Owners Can Block Them, source for Anthropic's separate crawlers and robots.txt opt out.
FAQ
How do I track AI citations?
Fix a prompt set that reflects a real buying decision, run every prompt on every engine you care about on a schedule, and log one row per run plus one row per brand named in that run. Record the engine, the mode, the prompt version and the date on every row. Without those four fields you get numbers you cannot compare next month.
Can I see which sources ChatGPT uses?
You can see the sources an assistant surfaces in a given answer, because linked citations sit alongside the response, and you can log those URLs. You cannot see the retrieval or ranking logic behind them, and the same prompt can return a different source set on the next run. Treat logged URLs as observations from a sample, not a complete source list.
How do I know if a competitor is cited more than me?
Count runs, not mentions. For each brand, count how many runs named it at least once, then divide that by the total of those counts across every brand plus an other bucket. Compare the resulting shares over the same run set and date range. A rival named in, say, 54 of 120 runs against your 18 is a gap you can size and act on.
How do I measure competitor AI citations?
Measure competitor AI citations as a share rather than a count. Run a frozen prompt set across engines, log every brand named per run, then divide each brand's run count by the sum of all brand run counts to get citation share. Establish a noise floor first by running the identical set twice with no changes, and report only movements larger than that spread.
How many prompts and runs do I need for a reliable citation share?
There is no published threshold worth quoting, so derive your own. Start with enough prompts to cover your buying decision, which for us usually means thirty to sixty, run each at least three times per engine, then measure your noise floor by repeating the whole set unchanged. If the spread between the two identical passes is wide, add runs before you add interpretation.
Does a competitor AI citation send me traffic?
Rarely, and you should not budget as if it does. Pew Research Center found users clicked a source cited inside an AI summary in just 1 percent of visits, so a citation is better read as a brand exposure event than a traffic channel. Measure it as presence and share, and keep referral traffic as a separate, caveated number.









































































.webp)


























