How to Calculate Citation Share: A Worked Example

The citation share formula, a 40-prompt sample set, a sampling cadence and a spreadsheet template, with every calculation worked through step by step.

ArticleBY THE ASTROFABRIC TEAM · AUG 14, 2026 · 12 MIN READ

Dark abstract grid of glowing data cells converging into a single rising line of light, suggesting counted AI answer mentions resolving into one visibility metric

Citation share is your brand's citations divided by all brand citations across a fixed prompt set, model set and time window. To calculate it, freeze 40 buyer-style prompts, declare your engines and competitive set, run each prompt three times per engine, and log every answer as a row. Divide your mention count by total mentions. In the sample dataset below, 149 brand mentions out of 1,240 across 480 logged answers gives 12.0 percent. Report presence rate and per-engine share of model alongside that figure. Those three numbers answer three different questions.

Citation share in one formula

The formula is short on purpose: citation share equals your brand's citations, divided by all brand citations across the same prompt set, the same model set and the same time window. The rest of this post is about locking those denominators so they mean the same thing every week. The division itself is the easy part.

People treat three numbers as interchangeable. Hold them apart:

  • Presence rate - the share of answers you appear in at all, regardless of who else is mentioned.
  • Citation share - your slice of all brand mentions logged across the competitive set.
  • Share of model - the same citation-share math, but cut per engine instead of pooled across all of them.

A brand can post high presence and low citation share in the same window if it shows up often while three competitors show up just as often. Read that gap as a diagnostic.

The numerator: what counts as a citation

Decide once whether a citation means any mention of your brand name, a linked source pointing to your domain, or both counted separately. The cleanest approach logs both: a brands_mentioned count and a brands_linked count, so you can calculate citation share on either basis and compare them.

The denominator: total mentions or a fixed competitive set

There are two defensible denominators. The first is every brand mention that appears in the logged answers, whoever it belongs to. The second is a fixed competitive set you name in advance - your brand plus a specific list of competitors - and you count mentions only within that list. Either works. Switching between them mid-quarter does not.

Citation share vs presence rate vs share of model

The unit of observation underneath all three metrics is identical: one answer, to one prompt, from one model, on one date, logged as a single row. Presence rate is a yes/no rollup of that row. Citation share is a mention-count rollup across all rows. Share of model is the same mention-count rollup filtered to rows from one engine. Read the hub definition of citation share for the conceptual grounding, then treat everything below as the arithmetic.

What do you need before you calculate citation share?

Five inputs need to exist before the first prompt runs. None of them should change mid-measurement without a note in your log.

The five inputs checklist
  • A frozen prompt set, unchanged for at least one quarter
  • A declared model set (for example ChatGPT, Perplexity, Gemini, Google AI Overviews)
  • A declared competitive set, named before you start counting
  • A sampling cadence, written down and followed even on slow weeks
  • A logging schema, built before the first run so every row is comparable

Rules that must apply symmetrically to competitors

Whatever you decide about unlinked mentions, alias handling, or subsidiary brand names, apply it identically to every competitor in the set. A rule that is generous to your own brand and strict toward competitors will inflate your number in a way that falls apart the moment someone checks the raw answers.

The logging schema, column by column

Every run becomes one row with these columns: prompt ID, prompt text, engine, run date, run index, answer text, brands mentioned, brands linked, whether your domain was cited, and the position of the first mention in the answer. Nothing else needs to exist yet. The calculation tabs come later.

Freeze before you measure
Most citation-share tracking fails because the prompt set or the competitive set changed between weeks, even when the arithmetic is sound. A moving denominator produces a moving number for reasons that have nothing to do with visibility.

Building the prompt set: 40 questions that mirror real buying behavior

A workable starting set has 40 prompts split into four buckets of ten. Splitting the buckets matters because they behave differently. Averaging them into one number hides the story you came here to see.

The four prompt buckets

  1. Category definition - "what is [category]" style questions a buyer asks before they know vendor names.
  2. Comparison - "X vs Y" and "best [category] for [use case]" questions asked mid-evaluation.
  3. Problem or job-to-be-done - questions phrased around a task rather than a product category.
  4. Vendor-name and alternatives - questions that already include a brand name, yours or a competitor's.

Sample prompts you can copy

  • Definition: "What does an AI visibility platform actually do?"
  • Comparison: "What's the best tool for tracking brand mentions in ChatGPT answers?"
  • Problem: "How do I know if my content is getting cited by AI assistants?"
  • Vendor-name: "What are alternatives to [competitor] for AI search monitoring?"

Write them the way a buyer types them into a chat window. Keyword-tool phrasing produces a different set of answers, and that difference alone changes what you get back.

Prompt IDs, weights and a change log

Give each prompt a stable ID and an importance weight from 1 to 3, based on how close it sits to a purchase decision. Vendor-name and comparison prompts typically weigh more than definition prompts. Keep a dated appendix of every prompt added or retired, so a future shift in the trend line has an audit trail instead of a mystery.

How often should you sample AI answers?

A single query to a single engine is an anecdote. Generated answers vary run to run even with identical prompts, which is why three to five runs per prompt per engine is the working minimum before you trust a count.

Runs per prompt and why variance matters

Run the same prompt three times on the same day and you will often see different brands surface in different orders, sometimes with one run omitting a brand another run included. That variance is normal model behavior. It is the reason single-run numbers mislead.

A weekly cadence with a launch-week pulse

For most teams, run the full 40-prompt set weekly. During a launch or a competitive push, add a lighter daily pulse covering 8 to 10 decision-stage prompts so you can see early movement without re-running the entire set every day.

Rolling windows over single snapshots

Standardize run conditions: a fresh session each time, no personalization carried over, the same country and language settings, and a recorded timestamp on every row. Then compute the headline metric on a rolling four-week window rather than a single week, so one volatile week does not read as a trend reversal. Store every raw answer text. Re-scoring history under a new rule is the only way to fix a methodology change without discarding your baseline. This is the same discipline behind cross-channel share of voice reporting, just applied to generated answers instead of publications.

The worked example: from 480 answers to one number

Here is an illustrative dataset you can mirror with your own numbers: 40 prompts x 4 engines x 3 runs produces 480 logged answers in one week. Every figure below is a sample number for the arithmetic. Replace them with your own logs once you are running this for real.

  1. Count observations. 480 rows total, 120 rows per engine.
  2. Count total brand citations. Across the competitive set, those 480 answers contain 1,240 brand mentions.
  3. Count your mentions. Your brand appears 149 times across those rows.
  4. Divide. 149 / 1,240 = 0.120, so raw citation share is 12.0 percent.
  5. Compute presence rate separately. Your brand appears in 96 of the 480 answers, so presence rate is 20.0 percent. That is a different question with a different answer.
  6. Cut by engine. On one engine, 58 of your 149 mentions occur within 310 total brand mentions, giving 18.7 percent share of model. On another, it's 21 mentions within 300 total, giving 7.0 percent.
  7. Cut by prompt bucket. Definition and comparison prompts almost always diverge, and that gap becomes your next content brief.
12.0%Raw citation share in the sample: 149 of 1,240 brand mentions
SAMPLE-DATASET
EngineAnswers loggedYour mentionsTotal brand mentionsPresence rateRaw citation shareWeighted citation share
ChatGPT1205831026.7%18.7%15.1%
Perplexity1202130012.5%7.0%5.8%
Gemini1203732019.2%11.6%9.0%
AI Overviews1203331016.7%10.6%8.2%
Total4801491,24020.0%12.0%9.4%

Every column in that table traces back to a raw log row. Nothing here is estimated.

How do you weight prompts and models fairly?

Unweighted citation share treats a low-stakes definition prompt exactly like a high-stakes comparison prompt. That flatters most brands that index well on easy, top-of-funnel questions.

The weighted formula

Weighted citation share is the sum of (mention count x prompt weight x engine weight) for your brand, divided by the same sum across the whole competitive set. It is the identical division as before, with each mention scaled before it gets counted.

Prompt weights from 1 to 3

Applying prompt weights of 1 to 3 to the sample dataset moves the raw 12.0 percent down to 9.4 percent weighted, because in this illustrative case the brand over-indexes on low-stakes definition prompts and under-performs on the comparison prompts that carry more weight. That is a meaningfully different story than the raw number tells alone.

Engine weights and how to justify them

Two defensible options exist for engine weighting: equal weight across all engines for simplicity, or weight by your own referral data if you have visibility into which engines actually send traffic. Either is fine as long as you document the choice and change it at most once per quarter, restating history whenever you do. Report raw and weighted figures side by side, always, so nobody has to trust an adjustment they cannot see.

The spreadsheet template you can build in 20 minutes

Five tabs are enough to run this entire process by hand.

  • Tab 1, Prompts: prompt_id, prompt_text, bucket, weight, added_date, retired_date.
  • Tab 2, Runs (the raw log): run_id, prompt_id, engine, run_date, run_index, answer_text, brands_mentioned, brands_linked, our_domain_cited, first_mention_position.
  • Tab 3, Competitors: brand_name, aliases, domain, in_competitive_set flag, so alias handling stays consistent across every count.
  • Tab 4, Calc: COUNTIFS and SUMPRODUCT formulas that turn the raw log into raw share, weighted share, presence rate and per-engine share of model.
  • Tab 5, Trend: one row per week per engine, feeding a single chart with a rolling four-week line.

The formulas that do the work

A SUMPRODUCT that multiplies mention count by prompt weight, summed across the competitive set, gives you the weighted denominator. The same formula filtered to your own brand's rows gives you the weighted numerator. COUNTIFS handles the simpler raw counts and the presence-rate rollup directly from Tab 2.

Two mistakes that corrupt the trend line

Mistakes to design out early
  • Alias mismatches, where a competitor's shortened name or subsidiary brand gets miscounted or missed entirely
  • Overwriting last week's rows instead of appending new ones, which silently destroys your historical trend

Reading the number: what movement actually means

Before reading anything into a week-over-week change, check run-to-run variance within that same week. If three runs of the same prompt already disagree by several points, a small move in the headline number is noise.

Signal versus variance

Once you have ruled out variance, three patterns tend to explain most real movement.

Three diagnostic patterns and their fixes

  • Presence rising, citation share flat - the category is getting more crowded and you're gaining ground more slowly than competitors, even though you show up more often.
  • Strong on definition, weak on comparison - a content gap on comparison and alternatives pages, which usually needs new pages rather than edits to existing ones.
  • High mention count, low linked-source count - models know your name but aren't retrieving your pages, which is a retrievability problem rather than a brand-awareness one.

When three or more of these patterns show up at once, that is the point to run a full AI visibility audit rather than patching one bucket at a time.

The review rhythm

Review the raw numbers weekly, write the narrative monthly, and revisit the prompt set itself quarterly. That rhythm keeps the metric from becoming either ignored or over-reacted to.

Running the calculation on a schedule instead of by hand

Manual measurement works fine for a quarter. Then the cadence slips. Someone is on vacation during the weekly run, a competitor renames a product, and the spreadsheet quietly stops being comparable to itself. The fix is scheduling the same arithmetic. Redesigning it every time the cadence slips is how the baseline gets lost.

From spreadsheet to scheduled job

AstroFabric's AI visibility agent runs your frozen prompt set across surfaces on the cadence you set, and its code sandbox performs the exact computation behind raw share, weighted share, presence rate and per-engine share of model, so those figures are arithmetic rather than estimates. If you want a deeper walkthrough of the mention-logging side of this before automating it, track brand mentions in ChatGPT and Perplexity covers the manual version step by step.

Where the numbers get delivered

Results land in the console, over REST, through MCP, or as scheduled digests in email, Slack or Telegram, depending on how your team wants to see them. Metered tool capabilities and credit-based pricing keep sampling volume a deliberate choice rather than an open-ended cost, and approval-gated writes mean that when the content agent turns a comparison-prompt gap into a brief, nothing publishes without a human sign-off first.

The six-step recurring checklist
  • Freeze the prompt set for the quarter
  • Declare the model set and the competitive set in writing
  • Log every raw answer, not just the counts derived from it
  • Compute raw and weighted citation share on the same run
  • Report presence rate and share of model alongside citation share
  • Review the prompt set itself once per quarter

For background on how the browsing and retrieval behavior behind these answers actually works, OpenAI's documentation on how ChatGPT search works and Perplexity's notes on how citations get generated are both useful reading before you finalize your logging rules.

If you'd rather have this running on a schedule than rebuilt in a spreadsheet every quarter, sign up for AstroFabric and point the AI visibility agent at your prompt set.

Frequently asked questions

What is the citation share formula?

Citation share equals your brand's citations divided by all brand citations from the same prompt set, model set and time window, expressed as a percentage. The weighted version multiplies each mention by a prompt importance weight and an engine weight before dividing. Publish both figures together, along with the competitive set you counted, so anyone reading the number can reproduce it.

How many prompts do I need for a reliable calculation?

Forty prompts split evenly across definition, comparison, problem and vendor-name buckets is a solid starting set for one category. Fewer than twenty makes bucket-level cuts too thin to read. Freeze the set for a quarter, log additions and retirements with dates, and expand only when you enter a genuinely new topic area.

How often should I sample to track citation share?

Run the full prompt set weekly with three to five runs per prompt per engine, and report on a rolling four-week window. During a launch, add a daily pulse on eight to ten decision-stage prompts. Answers vary between identical queries, so repeated runs are what separate a real shift from ordinary generation variance.

What is the difference between citation share and share of model?

Citation share is your slice of all brand mentions across every engine you track. Share of model is the identical calculation restricted to one engine, so you get a separate figure for ChatGPT, Perplexity, Gemini and AI Overviews. Engines retrieve differently, and the per-engine cut usually tells you where to focus first.

Should unlinked brand mentions count as citations?

Either rule works as long as you apply it to every brand identically and state it in your reporting. Many teams log both: brands_mentioned and brands_linked as separate columns. That pairing is useful on its own, because plenty of mentions with few links means assistants know your name but are not retrieving your pages.

Do I need software, or is a spreadsheet enough?

A five-tab spreadsheet handles the math for a quarter, and building it teaches you exactly what the number contains. Cadence is where manual tracking usually slips. Scheduling the same runs and the same arithmetic keeps the trend line intact, whether that scheduling lives in a script or in a platform that runs prompt sets for you.

Sources

⟨ RUN IT INSTEAD OF READING IT ⟩

Every playbook on this blog ships as a runnable mission.

Open a workspace and the playbook library is waiting - describe the outcome and the agents carry it end to end, on your plan's monthly credits.

⟨ KEEP READING ⟩
PlaybookAI search & GEO

The AI visibility audit you can run this week

A complete audit in five steps: build the question set, measure presence across models, diagnose absences by pipeline stage, rank the moves, set the cadence - with a presence-rate calculator.

Aug 13, 2026 · 9 min read