
AI Visibility Sample Size and Prompt Sampling Methodology
A defensible AI visibility sample size: about 40 unbranded prompts per engine, three replicates each, across three or four engines, run weekly and reported monthly. That is roughly 480 runs per cycle and plus or minus 10 points of margin at 20% citation share. Detecting a 5-point shift needs roughly 1,000 effective observations, so pool weekly runs into monthly rollups. Report the sample size and interval next to every number.
What are you actually estimating when you measure AI visibility?
Sample-size math only starts to help once you are precise about what the number even represents. Most teams say "our AI visibility is 22%" without ever stating what population that 22% describes. A perfectly fine measurement then turns into a confusing quarterly debate.
Prompt run as the unit of observation
The atomic unit here is a single prompt, sent to a single engine, at a single point in time, producing a single answer that either cites you or does not. A weekly rollup is an aggregate of those binary outcomes. So is a monthly trend. So is a competitor comparison. Treating the prompt run as the unit, rather than "the topic" or "the query," is what lets you apply ordinary proportion statistics to the result.
Coverage versus share within an answer
Two related but distinct numbers get merged constantly:
- Coverage - the share of prompts where your brand appears anywhere in the answer, at all.
- Share within answer - your share of voice among the brands that show up when several are cited together.
A brand can post high coverage and still own a weak share within the answer, appearing often but as a minor mention behind two competitors. Reporting only a blended figure hides which of those two problems you actually have. Read more on share of voice tracking to see how this split plays out across channels beyond AI answers.
Writing the estimand down first
Write the estimand in one sentence before you pick a sample size: "our citation share on mid-funnel comparison prompts in ChatGPT during October." That sentence locks the engine, the time window and the prompt segment. A stakeholder can no longer compare your November comparison-prompt number to your October full-set number as if they were the same measurement.
AI visibility sample size: how many prompts per engine do you need?
Fix the estimand and the question becomes a standard sample-size problem for a proportion.
The margin-of-error formula in plain form
At 95% confidence, the margin of error for a proportion is approximately:
1.96 x sqrt( p(1-p) / n )
where p is your observed citation share and n is the number of prompt runs per engine. At p = 0.20 and n = 30, that works out to roughly plus or minus 14 points. A reported 20% could honestly sit anywhere from 6% to 34%.
±14 pts95% margin of error at n=30 prompt runs, p=20% citation shareWhy square-root scaling changes your budget math
Precision improves with the square root of n, not with n itself. Halving your margin of error costs roughly four times the runs. Doubling the budget will not get you there. That fact should reshape how a team allocates its AI visibility spend. A one-off 1,000-prompt sweep looks impressive. A steady weekly cadence at a sane size compounds into more reliable trend data for less total spend, because pooling across periods buys the precision that one giant sweep buys once.
Floors for directional, reportable, and segment-level reads
As a practical set of floors per engine:
- 20 prompts - directional signal only, useful for spotting gross anomalies.
- 40 prompts - supports monthly reporting at a reasonable margin.
- 100+ prompts - supports segment-level cuts, such as by intent stage or geography.
| Effective obs. per engine/period | 95% margin at p=20% | Smallest reliable period-over-period change | Weekly runs (4 engines, 3 replicates) | Fit for |
|---|---|---|---|---|
| 20 | ±18 pts | not detectable | ~27/engine | Direction only |
| 40 | ±12 pts | ~15 pts | ~53/engine | Early monthly trend |
| 100 | ±8 pts | ~11 pts | ~133/engine | Reportable monthly trend |
| 250 | ±5 pts | ~7 pts | ~333/engine | Segment-level cuts |
| 600 | ±3 pts | ~4 pts | ~800/engine | Competitor-level share of voice |
That table is the whole methodology in one place. Pick the row that matches the decision you need to make, then build the run schedule that produces it.
Why replicates matter more than most prompt sampling methodology allows for
Sample size alone is not the full story. A single run of a prompt does not behave like a single independent draw.
Non-determinism as a measurement property
Assistants are non-deterministic by design. The same prompt, on the same engine, on the same day, can retrieve different sources and cite different brands across separate calls. Run each prompt exactly once and you have no way to separate genuine prompt-to-prompt variation in the market from ordinary run-to-run instability in the model.
Choosing a replicate count
Three replicates per prompt is a workable default. That count is enough to see whether a prompt's citation behavior is stable or volatile. It will not quadruple your run budget the way five or six replicates would.
Effective sample size and the design effect
Replicates of the same prompt are correlated with each other, so 40 prompts times 3 replicates is not 120 independent observations. With an intra-cluster correlation of roughly 0.5, the effective sample size lands closer to 60.
"n = 120 runs, effective n = 60, margin roughly plus or minus 10 points" is a more honest and more defensible line in a board deck than a bare "120 runs." Anyone who has run citation share numbers past a skeptical stakeholder knows the second version invites the exact question the first one already answers.
How many engines should you split your prompt budget across?
Prompt budget and engine count trade off directly against each other. This is where a lot of otherwise careful sampling designs quietly fall apart.
The precision cost of breadth
Every additional engine divides your fixed run budget further. Six engines at a fixed total budget gives each one roughly a third of the precision that two engines would get from the same spend. Breadth feels thorough. You buy it with the exact precision you need to say anything confidently about any single engine.
Why blended cross-engine averages mislead
Never average citation share across engines into one headline figure. Retrieval behavior differs between engines. Source preferences differ. Freshness windows differ. A blended number hides the only detail you can act on. A "22% average" can easily be 40% in one assistant and near zero in another, and those two situations call for entirely different remediation work, as covered in more depth in an AI visibility audit.
Tiered coverage for long-tail engines
A sensible default is three or four engines, reported separately, chosen by where your category's buyers actually ask questions. If you need broader coverage than that:
- Run the full prompt set on your two priority engines.
- Run a reduced core subset on the remaining engines.
- Keep the prompt wording identical across every engine so comparisons stay like-for-like rather than an artifact of differently phrased questions.
How often should you run the set, and what can each cadence detect?
Estimating a single level is one problem. Detecting whether that level has changed between two periods is a harder one, because both periods carry their own error.
Minimum detectable change by period size
As a rough rule for comparing two proportions at p = 0.20 with 80% statistical power, detecting a 10-point shift needs roughly 250 effective observations per period. Detecting a 5-point shift needs roughly 1,000. A citation share that moves from 18% to 24% in a single week is almost always inside the noise band. The same move sustained across a pooled month is worth attributing to something real.
Weekly for direction, monthly for reporting
Run weekly for anomaly spotting and directional awareness. Pool into monthly figures for anything you report externally or to leadership. Search Engine Land's coverage of AI search behavior is a useful reminder of how quickly retrieval patterns shift week to week. A single weekly number should never carry the weight of a trend line on its own.
Frozen core plus rotating exploratory prompts
Keep roughly 80% of the prompt set frozen for continuity and rotate the remaining 20% quarterly to capture emerging phrasing and new competitor framing. Time-stamp every run. Never backfill. Retrieval indexes change, so a prompt re-run in November is not a measurement of October, no matter how tempting it is to patch a gap in the log that way.
Prompt sampling methodology: stratification, weighting and what to exclude
The run counts above only mean something if the underlying prompt list actually represents how buyers ask questions.
Four strata worth separating
Stratify prompts into at least four intent stages:
- Category definition ("what is a [category] tool")
- Solution comparison ("X vs Y for [use case]")
- Vendor shortlist ("best [category] for [segment]")
- Objection or pricing ("is [category] worth it," "how much does X cost")
Weighting by commercial value
Allocate prompts proportionally to commercial value, not evenly across strata, then weight the results back when you report an overall figure. A strong showing on low-value definitional prompts will otherwise flatter a weak shortlist number that actually determines revenue outcomes.
- Prompts are written the way buyers actually phrase them, including conversational and sloppy forms, not as keyword strings.
- Branded prompts are excluded from the headline citation share and tracked as a separate diagnostic set.
- Every logged run captures the prompt, engine, timestamp and full answer text.
- Every cited URL and every brand mentioned in the answer is recorded, not only your own hits.
- The prompt list version is dated, so a rotation is visible on the trend chart rather than mistaken for a market move.
Branded prompts and other contaminants
"What does [your brand] do" will almost always cite you. Leave that prompt in the headline set and you inflate the number. You also hide the real story on the comparison and pricing prompts that actually decide shortlists. Keep a small branded set for a different purpose entirely: checking how accurately assistants describe you. That check is a good companion exercise to a wider effort to track brand mentions in ChatGPT and Perplexity.
A worked example: 40 prompts, 3 replicates, 4 engines
Putting the whole design together in one concrete plan makes the abstractions above much easier to apply.
The design table
- Prompt set: 40 unbranded prompts, stratified roughly 12/12/10/6 across the four intent strata.
- Replicates: 3 per prompt.
- Engines: 4, reported separately.
- Cadence: weekly runs, pooled monthly.
- Volume: 40 x 3 x 4 = 480 runs per week, 120 runs per engine per week, roughly 2,000 runs per engine per month.
Weekly versus monthly precision
The weekly per-engine read has an effective n around 60, so an observed citation share of 20% carries a margin of roughly plus or minus 10 points. Treat that number as direction only. The monthly pooled per-engine read has an effective n around 240. That narrows the margin to roughly plus or minus 5 points, which is precise enough to support a reported trend line.
Reading a movement correctly
A move from 18% to 24% inside a single week sits comfortably inside the weekly noise band. It should not trigger a strategy debate on its own. The same move, sustained across a pooled monthly figure, clears the detectable-change threshold. Attribute it to a specific cause. A content push can do it. So can a competitor launch, or a change in how a given engine retrieves sources.
What segment reporting costs
Reporting the comparison stratum on its own is not a free slice of the total 480 weekly runs. That stratum needs its own effective n to support a standalone claim. That is exactly why the sample-size matrix above lists "segment-level cuts" as needing roughly 250 effective observations, not the 60 that the full weekly design happens to produce per engine.
Running the sampling design without running it by hand
None of the arithmetic here is difficult. The burden is operational. You need hundreds of runs per week. You need answers parsed and URLs deduplicated. You need a schema that stays stable across four differently-behaved engines.
- In AstroFabric, the AI visibility agent handles scheduled prompt runs through metered tool capabilities, so you define the design once - prompt set, replicate count, engine list, cadence - and a credit ceiling caps the spend exactly where you set it.
- Confidence intervals, design effects and weighted rollups are computed in a code sandbox for exact numbers, rather than approximated inside a language model's own reasoning.
- Results land wherever your team already works: console, REST, MCP, widget, email, Slack or Telegram, so a weekly direction check does not require anyone to open a separate tool.
- Anything the content agent proposes in response to a sampling signal stays approval-gated, so a genuine citation-share drop turns into a reviewed content brief rather than an unreviewed page going live on its own.
Vendors covering this space, including Profound and Semrush, converge on the same underlying point: the value is in a consistent, well-designed run schedule, not in any single sweep, however large.
If you want the design above running on a schedule, with the statistics computed exactly and the results routed to wherever your team already reads them, sign up for AstroFabric and set the prompt list, replicate count and credit ceiling once.
Frequently asked questions
What is the minimum number of prompts for AI visibility tracking?
Twenty prompts per engine is the practical floor, and it buys direction only. At a 20% citation share, twenty runs give a margin of roughly plus or minus 18 points, so almost any weekly movement is noise. Forty prompts per engine with three replicates each supports monthly reporting. One hundred or more per engine is where segment-level cuts by intent stage start to hold up.
Why do I need to run the same prompt more than once?
AI assistants are non-deterministic. The same prompt sent twice to the same engine can retrieve different sources and cite different brands. Single runs blend prompt-level variation with run-to-run variation, so you cannot tell whether a change came from the market or the sampler. Three replicates per prompt captures that instability at a reasonable cost and lets you report an effective sample size honestly.
Should I average citation share across engines?
Report per engine. Retrieval mechanics, source preferences and freshness windows differ enough that a blended average hides the only detail you can act on. A blended 22% might be 40% in one assistant and 4% in another, and those two situations call for entirely different work. Use the same prompt list across engines so the comparison stays like-for-like.
How large a change in citation share is actually meaningful?
It depends on your effective sample size per period. At around 250 effective observations per period and a 20% baseline, you can reliably detect shifts of about 10 points. Detecting a 5-point shift takes roughly 1,000. That is why weekly numbers work as anomaly detection and monthly pooled figures work as reporting. Publish the interval next to the point estimate.
How often should the prompt set change?
Keep roughly 80% of the set frozen so your trend line stays comparable across months, and rotate the remaining 20% quarterly to capture emerging questions and new competitor framing. Log every version of the set with dates. When you swap prompts, note the change on the chart so a step in the data is never mistaken for a market move.
Do branded prompts belong in the sample?
Track them separately. Questions containing your own name will nearly always surface you, so including them inflates headline citation share and masks weakness on the category and comparison prompts that decide shortlists. Run a small branded set to monitor how accurately assistants describe you, and keep your reported visibility figure to unbranded buyer questions.
Sources
Every playbook on this blog ships as a runnable mission.
Open a workspace and the playbook library is waiting - describe the outcome and the agents carry it end to end, on your plan's monthly credits.