
We asked ChatGPT and Grok the same question, "what is the best agentic AI platform", and collected 15 citations with zero overlapping domains. That result is why LLM citations by engine is the measurement that matters. Each engine builds answers from its own retrieval stack, so a blended visibility score hides where you actually appear. Track every engine separately, compute citation share per engine, and the work becomes clear: you can see which answer surfaces need repair and which content formats each engine rewards.
One Prompt, Two Engines, Zero Overlap: The Raw Data
The teardown is simple. We ran one commercial prompt through ChatGPT and Grok, collected every cited domain from both answers, and found no shared ground. Fifteen citations. Two completely separate reading lists. I expected thin overlap. Zero still landed differently.
0overlapping domains across 15 total citationsThe prompt was deliberately competitive: "what is the best agentic AI platform" is the kind of question a buyer types when budget is already approved. Both answers arrived confidently. Both cited sources. Yet the two engines behaved like editors at rival magazines who have never read each other's publication.
| ChatGPT cited | Format | Grok cited | Format |
|---|---|---|---|
| eesel.ai | Vendor comparison blog | gumloop.com | Vendor platform page |
| ampcome.com | Agency listicle | kore.ai | Vendor platform page |
| softwaresuggest.com | Software directory | gartner.com | Analyst page |
| Four more domains | Directories and roundups | Five more domains | Vendor and analyst mix |
| Overlap: 0 domains | Total: 15 citations |
What ChatGPT cited and why it feels like a listicle diet
ChatGPT's shortlist reads like someone who eats comparison content for breakfast. eesel.ai publishes vendor comparison posts, ampcome.com is an agency running listicles, and softwaresuggest.com is a classic software directory. These pages arrive pre-chewed. Think numbered feature grids and verdict paragraphs. ChatGPT reached for pages that already look like answers, and that says something useful about what its retrieval layer rewards.
What Grok cited and why it reads more editorial
Grok's list has a different flavor entirely. gumloop.com and kore.ai are vendors speaking in their own voice, and gartner.com brings the analyst register. There is less "top 10" scaffolding and more primary-source material. Same question, but Grok assembled its answer from vendors and analysts rather than aggregators. Two engines, two editorial sensibilities. A brand can look strong to one editor while the other has never heard of it. A blended score will not show that gap. It gives you a lukewarm middle that describes neither engine.
Why Do ChatGPT and Grok Cite Completely Different Sources?
The tempting explanation is that one engine is wrong. The honest explanation is more interesting: they are sampling different libraries with different tastes, and divergence is the expected outcome.
Different indexes, different candidate pools
Each engine sits on its own retrieval stack. Its crawler and index shape the candidate pool before any ranking decision happens. A page that ChatGPT's index refreshed last week may sit stale or absent in Grok's. You cannot be cited from an index you were never pulled into. That simple fact explains more zero-overlap results than any ranking mystery.
How AI engines pick sources once candidates are retrieved
Once candidates exist, each engine applies its own reranking logic, and editorial personality enters. ChatGPT's browsing behavior visibly favors pages already structured as answers, which is why directories and comparison posts dominated its citations in our teardown. Grok pulls from a different slice of the web and folds in its own real-time signals. We have written a full breakdown of how AI assistants choose their sources, so I will not rebuild retrieval from scratch here. The short version: retrieval is a pipeline of preferences, and every engine tuned its preferences independently.
Why freshness and format preferences diverge
Freshness windows differ too. An engine that re-crawls aggressively may cite this month's roundup, while an engine with a slower window cites the evergreen analyst page that has held its ranking for a year. Put format preference on top of freshness and you get two engines that could agree in theory, yet rarely do on subjective commercial queries. Factual questions with one canonical source produce overlap. "Best of" questions produce fifteen citations and an empty intersection.
What Does a Blended AI Visibility Score Actually Hide?
This is where the teardown stops being trivia and starts costing people money.
The averaging trap, with numbers
Run the failure mode with real numbers. Say your brand appears in 4 of 10 ChatGPT answers for your prompt set and 0 of 10 Grok answers. The blended dashboard reports 20 percent visibility. That number describes neither reality. On ChatGPT, you have a foothold worth defending. On Grok, you do not exist, and that is an emergency if your buyers live there. The 20 percent tells you to feel mildly okay, which is the wrong feeling for both situations.
Our zero-overlap result is the extreme case that proves the rule. When two engines share no sources at all, averaging their citation data produces pure noise dressed up as a KPI. You are computing the mean of two unrelated distributions and putting it on a slide.
Mentions versus citations before you even split by engine
There is a second layer of confusion underneath the first. A mention is your brand named in the answer text. A citation is your domain linked as a source. They are different signals with different causes, and plenty of tools quietly blend them. Blend mentions with citations, then blend engines with each other, and you have compounded two distortions into one confident-looking number. Keep the signals separate and keep the engines separate. Then the dashboard starts telling the truth.
How to Track LLM Citations by Engine
The method is unglamorous and it works. You are building a longitudinal log, and the discipline is in what you refuse to merge.
Build the prompt set before you build the dashboard
Start with 10 to 30 prompts your actual buyers would type, from "best X for Y" to "X vs Z". Freeze the wording. A prompt set that drifts is a measurement that lies, because you can never tell whether the answer changed or your question did.
Log per engine, per run, per domain
The minimum viable spreadsheet has five columns, and every citation gets its own row.
- Engine: which model answered, logged exactly, never merged.
- Prompt: the frozen wording from your set.
- Date: every run, even when nothing changed.
- Cited domain: the root domain of each source linked.
- Position: first citation carries more weight than fifth.
Position deserves a defense because people skip it. Being the first source an engine leans on is a different commercial reality from being a trailing footnote. If you only log presence, you miss the moment you climb from fifth citation to first. Once the log exists, the math is straightforward. We walk through how to calculate citation share step by step in a worked example, so the numbers stay reproducible instead of vibes.
How often to re-run and why drift is the point
Answers drift week to week. Engines re-crawl, rerank, and swap sources without announcement, so a single snapshot is a photograph of weather when what you want is the climate. Weekly runs are the sweet spot for most teams: frequent enough to catch a source swap, spaced enough that the trend line means something. The drift is where the insight lives.
Citation Share by Engine: The Metric That Replaces the Average
Here is the number that should sit where the blended score used to.
The formula, engine by engine
Citation share by engine is your domains cited divided by total citations, computed separately for each engine, over each run of your prompt set. If ChatGPT's answers produce 40 citations across your prompts and 6 point to you, your ChatGPT citation share is 15 percent. Grok gets its own numerator and denominator from its own answers. The full framework, including how this rolls into a broader scoreboard, lives in our citation share hub. Per-engine citation share feeds share of voice across channels. It sharpens the bigger picture rather than replacing it.
Reading a zero honestly
Now the uncomfortable part, and I will apply the medicine to ourselves. On this specific query, our own share of voice is zero mentions and zero citations across both engines. A blended metric would let us bury that inside some rounder number. The per-engine view puts it on the table plainly: two engines, two zeros, two separate problems to solve. That specificity is what makes the metric a prioritization tool. You invest where the gap is largest and where the engine's demonstrated source preferences match content you can realistically produce. A zero on an engine that quotes directories is a listing problem. A zero on an engine that quotes analysts is a much longer game.
15%example ChatGPT citation share: 6 of 40 citations, computed for that engine aloneTurning Per-Engine Data into a Working Program
Measurement without action is a hobby, so here is how the teardown becomes a program.
Audit each engine as its own channel
Run an AI visibility audit as the first structured pass, and let the zero-overlap finding set the ground rule: audit each engine separately, because each one is effectively its own channel with its own editor and submission guidelines. Map where you appear, then do the same for competitors and the source types that dominate each engine's answers for your prompts.
Match content format to engine preference
Then match the play to the appetite. In our teardown, ChatGPT's citations skewed hard toward comparison and directory formats. That suggests earning placement in the roundups and directories it already trusts, then publishing comparison content structured the way it likes to quote. Grok's vendor-and-analyst mix points somewhere else entirely: stronger first-party pages and analyst-facing material. One content push rarely moves both engines equally, and that is fine.
Set timeline expectations honestly too. Engines re-crawl and re-rank on their own clocks, so per-engine wins arrive unevenly. Keep the loop tight: ship a content push, re-run the frozen prompt set, and watch which engine moves first. That first mover tells you whose editorial clock runs fastest for your niche, which is itself useful intelligence for sequencing the next push.
Running the Teardown Yourself with AstroFabric
Everything above is doable with a spreadsheet and a stubborn streak. It is also the kind of work that quietly stops happening in week four, which is why we built it into AstroFabric. The AI visibility agent runs your prompt set across engines on a schedule and logs citations per engine, per run, per domain. The teardown becomes a recurring report instead of a weekend project you keep meaning to repeat.
The citation share math runs in a code sandbox as exact computation, which means the percentages in your report are calculated rather than estimated by a model. A 15 percent share is genuinely 6 divided by 40. Findings arrive wherever your team actually lives: the console, Slack, Telegram, or email. When the agents propose changes based on what the data shows, every write is approval-gated, so nothing moves without a human saying yes.
The honest takeaway from fifteen citations and an empty intersection is simple: measure each engine on its own terms, because the data says they behave like different channels with different editors. If you want that measurement running by next week instead of next quarter, start with AstroFabric and let the first per-engine report make the case for itself.
Frequently asked questions
Why do ChatGPT and Grok cite different sources for the same prompt?
Each engine runs its own retrieval stack: separate crawlers, separate indexes, separate freshness windows and separate reranking logic. In our teardown of 'what is the best agentic AI platform', ChatGPT cited eesel.ai, ampcome.com and softwaresuggest.com while Grok cited gumloop.com, kore.ai and gartner.com. They sample different libraries with different editorial tastes, so divergence is the expected behavior rather than a glitch.
Is zero citation overlap between AI engines normal?
It is more common than most dashboards suggest, especially on competitive commercial queries. Engines differ in what they index, how recently they crawled, and which content formats they prefer to quote. Overlap tends to rise on factual queries with a few canonical sources and fall on subjective 'best of' queries, where each engine assembles its own shortlist from its own candidate pool.
How do I track LLM citations by engine without special tooling?
Fix a prompt set, run each prompt against each engine on a regular schedule, and log engine, prompt, date, cited domain and citation position in one row per citation. Never merge engines into a single column. From that log you can compute citation share per engine and watch drift over time. A spreadsheet works at small scale, and automation earns its keep once the prompt set grows.
What is citation share by engine and how is it calculated?
Citation share by engine is the percentage of citations in an engine's answers that point to your domains, computed separately for each engine. If ChatGPT's answers to your prompt set produce 40 citations and 6 are yours, your ChatGPT citation share is 15 percent. Grok gets its own calculation from its own answers. Keeping the numbers separate shows exactly where you are strong and where you are absent.
Should I optimize for every AI engine at once?
Start with the engine where your buyers actually ask questions and where your per-engine data shows the largest fixable gap. The teardown shows engines reward different source types, so one content push rarely moves every engine equally. Treat each engine as its own channel with its own editor, sequence your effort accordingly, and re-measure after each push to see which engine responds first.
Sources
Every playbook on this blog ships as a runnable mission.
Open a workspace and the playbook library is waiting - describe the outcome and the agents carry it end to end, on your plan's monthly credits.