How to Track Brand Mentions in ChatGPT and Perplexity

Learn how to track brand mentions in ChatGPT and Perplexity: build a prompt set, sample answers on a schedule, and score AI share of voice you can report.

ArticleBY THE ASTROFABRIC TEAM · AUG 13, 2026 · 16 MIN READ

Abstract dark visualization of layered glass panels with connected nodes, a few glowing teal to suggest tracked brand mentions across AI answers

To track brand mentions in ChatGPT and Perplexity, build a fixed set of 40 to 150 buyer prompts, run each prompt several times per cycle across the assistants you care about, and store every raw answer with its model version, date, and cited sources. Then tag each answer for brand presence, answer position, competitor mentions, sentiment, and citations, and compute AI share of voice as your mentions divided by all brand mentions in the panel. Re-run the same panel on a weekly or monthly schedule so movement reflects real visibility changes rather than the natural variance in generated answers.

What does it mean to track brand mentions in ChatGPT?

A mention is your brand named somewhere in an assistant's answer text. A citation is your domain listed or linked as a source behind that answer. These two things move independently: a model can name your brand from something it learned during training without citing any page, and a page can quietly inform an answer while your brand name never appears. Any tracking system that conflates the two will mislead you about what's actually working.

AI answers are also generated probabilistically. The same prompt, asked twice on the same day, can return different competitor sets, different phrasing, and a different position for your brand in the response. That means measurement here looks more like sampling a population than checking a search rank. You're not looking for "the answer" - you're estimating how often, across many draws, your brand shows up and how it's described.

This is where AI share of voice becomes useful: your mentions divided by all brand mentions across a fixed prompt set, tracked over time. It's the same logic as share of voice in traditional media, just applied to generated text instead of published articles. The goal is a repeatable panel you can re-run every week or every month and defend in front of a board without hedging on methodology.

Scope matters too. ChatGPT, Perplexity, Google AI Mode, Claude, and Gemini each retrieve and generate differently - some browse live, some lean more on training data, some show sources inline and some don't show them at all. A single-assistant view will overstate or understate your position depending on which one you picked.

Mentions vs citations vs sentiment

Treat these as three separate columns in your data, not one blended score. Presence tells you whether you showed up. Citation tells you whether a page of yours earned the retrieval. Sentiment tells you whether the mention helps or hurts you - being named alongside a caveat about pricing or support is a different outcome than being named as the recommended option.

Why AI answers vary between runs

Model versions update, retrieval indexes refresh, and some assistants factor in session context or account history. None of that is a flaw in your measurement - it's the environment you're measuring. The fix is repetition: sample enough runs per cycle that your presence rate stabilizes even while individual answers keep shifting.

The three metrics worth reporting

Report AI share of voice, citation share, and a sentiment-adjusted presence rate for your top commercial prompt tier. Everything else - raw mention counts, one-off screenshots, single-run "wins" - is color commentary, not a metric to act on.

Build the prompt set that becomes your measurement panel

Start from the questions buyers actually type, not the keywords you'd optimize a landing page for. Category definitions ("what is [category] software"), best-of and alternatives lists, direct comparisons, pricing questions, integration questions, and objection prompts ("is [competitor] better for enterprise") all belong in the panel because they mirror how people actually use these assistants during evaluation.

A working panel usually lands between 40 and 150 prompts, grouped into intent tiers so you can segment results later instead of staring at one flat number. A B2B SaaS company might run 20 unbranded category prompts, 20 comparison prompts, 15 pricing/integration prompts, and 15 branded accuracy prompts - enough to be stable, not so many that scoring becomes a burden.

Split the panel into unbranded prompts, where you're competing for inclusion in a consideration set, and branded prompts, where you're competing for accuracy about your own product. Both matter. Losing an unbranded prompt costs you a lead you never see. Losing a branded prompt - an assistant stating a wrong price or a discontinued feature - can cost you a deal that's already in motion.

Freeze the phrasing once you've built the panel and version it like you would a codebase. Editing prompts mid-quarter breaks your trendline just as surely as changing your analytics tagging mid-campaign. Add locale or persona variants only where your buying committee genuinely differs - a finance-persona pricing prompt versus an IT-persona integration prompt, for instance - rather than multiplying variants for their own sake.

Prompt taxonomy: unbranded, branded, comparison, objection

Unbranded prompts test discovery. Branded prompts test accuracy. Comparison prompts ("X vs Y") test whether you're framed as the strong option or an afterthought. Objection prompts ("is X too expensive for small teams") test whether the assistant is repeating a stale criticism you've already addressed.

How many prompts you need for stable numbers

Below about 30 prompts, week-to-week swings are mostly noise. Above 150, scoring overhead grows without much added signal for most companies. Most teams find their sweet spot in the 60-100 range, weighted toward the tiers closest to revenue.

Versioning and change control for the panel

Log every prompt with an ID, a creation date, and a change history. When you must retire or reword a prompt, close out its old trendline and start a new one rather than splicing the two together.

Collect answers on a schedule you can trust

Pick a cadence that matches how fast the underlying situation changes. Weekly sampling suits active campaigns, product launches, or periods right after a content push, since it lets you see how quickly a change in your published content shows up in answers. Monthly sampling is plenty for a steady-state category baseline.

Run each prompt multiple times per cycle - three to five repetitions is a reasonable floor - and store every raw response rather than just the summary tag. You'll want the raw text later, both to re-score under new definitions and to pull exact quotes for a report.

Capture metadata alongside every run: model version, date, region, whether web browsing or retrieval was active, and whatever you can observe about session state. Two answers pulled a week apart under different model versions are not the same data point, and mixing them will hide the real cause of any change you see.

Keep the answer store warm - queryable, not archived to cold storage - because your tagging definitions will evolve. Being able to re-score last quarter's answers against this quarter's sentiment rubric is far more useful than starting from zero every time you refine the methodology.

Log cited URLs as their own field, separate from the mention tag. The source list is where content strategy decisions actually get made, and it deserves its own row in your data rather than a footnote inside the mention record.

Manual sampling vs API sampling vs monitoring tools

Manual sampling - a person typing prompts into ChatGPT and Perplexity and logging results in a spreadsheet - works fine for validating a new panel at small scale. API-based sampling scales better once you're running 60+ prompts across five assistants on a weekly cadence, since it removes the labor cost of repetitive querying. Purpose-built monitoring tools sit on top of either approach and handle scoring, storage, and trend reporting so the workflow doesn't live entirely in a spreadsheet.

Metadata fields to capture on every run

At minimum: prompt ID, assistant, model version, timestamp, region, browsing/retrieval status, and a flag for whether the account was signed in. These fields are what let you explain a spike instead of just reporting one.

Handling personalization and memory effects

Some assistants adjust answers based on account history or prior conversation turns. Sample from a clean, signed-out or fresh-session state wherever the assistant allows it, so your panel measures the general answer surface rather than a personalized one that only you would ever see.

Score the results: turning raw answers into share of voice

Once answers are collected, tag each one for brand presence, position in the answer (first mentioned, buried in a list, mentioned only in a caveat), competitor set named alongside you, sentiment, and citation source if one exists. This tagging layer is what turns a folder of text into a metric.

Compute a presence rate per prompt tier first - what percentage of comparison prompts mention you, what percentage of pricing prompts mention you - then aggregate to a single AI share of voice figure across the whole panel. The tier-level numbers are usually more actionable than the blended one, since they point directly at which content team owns the fix.

Weight prompts by commercial value rather than treating every prompt equally. A pricing comparison prompt where you're absent is a bigger problem than a general category-definition prompt where you're absent, and your weighted score should reflect that.

Track citation share as its own line rather than folding it into the mention score. Citation share is the lever you control most directly, since it traces back to specific pages you can rewrite, publish, or promote.

Set a variance band - based on your repeated-run data - before you start reacting to week-over-week movement. If your presence rate naturally swings by five points between identical runs, a five-point week-over-week change isn't news yet.

The share of voice formula, step by step

Count mentions of your brand across the panel, count total brand mentions across all competitors in those same answers, and divide. Report the result per tier and as a panel-wide weighted total. This is straightforward arithmetic, but it needs to run on exact counts rather than eyeballed impressions if you want the trendline to hold up under scrutiny.

Weighting prompts by revenue intent

A simple three-tier weighting - low, medium, high commercial intent - is enough for most teams. Assign pricing, comparison, and objection prompts to the high tier, integration and feature prompts to medium, and general category prompts to low.

What counts as a meaningful week-over-week change

Anything inside your measured variance band is noise. Anything outside it, especially if it's concentrated in one tier or tied to a specific competitor's cited page, is worth a follow-up.

Perplexity brand tracking: the citation-first playbook

Perplexity surfaces its sources inline with every answer, which makes it the cleanest environment for auditing exactly which pages earn retrieval. Where ChatGPT often names brands without showing its work, Perplexity gives you a source list you can map, page by page.

Classify every cited URL by content type: your own docs, a comparison page, a third-party review site, a forum thread, or a news article. This classification tells you immediately whether you're winning citations through owned content or losing them to sites you don't control.

Identify third-party sources that consistently outrank your own pages for a given prompt and treat them as placement targets - review sites, comparison roundups, and community threads where getting your product accurately represented is a realistic and valuable goal.

Watch the citation set after you publish or update a page, and log how many sampling cycles pass before that page starts appearing in the source list. That lag is one of the more concrete numbers you can bring to a content planning meeting, since it tells you how long a fix takes to show up.

Compare Perplexity's source mix against ChatGPT's answers for the same prompts. Where the two diverge - Perplexity citing a recent blog post that ChatGPT never mentions, for instance - you're seeing the gap between live retrieval and training-time knowledge, which is a useful diagnostic for generative engine optimization work more broadly.

Citation share is a leading indicator

Mention rate tells you where you stand today. Citation share tells you what's about to change, since a newly cited page tends to show up in more mentions a few cycles later.

Auditing the source list, not just the answer

Read the sources before you read the answer text. The source list explains why the answer says what it says, and it's the part you can actually act on.

Third-party placements that drive inclusion

If a review site or comparison page is consistently cited ahead of your own domain, treat outreach or contribution to that page as a legitimate placement task, not a side project.

Publish-to-citation lag and how to measure it

Tag each page update with a publish date, then note the first sampling cycle where that page appears in a citation list. Average this lag across several updates to get a planning number your content team can rely on.

How do you turn AI visibility monitoring into content decisions?

Every prompt where you're absent should become a content brief or a placement task - not a vague note in a monthly report. Absence on a comparison prompt points to a comparison page. Absence on an integration prompt points to a docs gap. The panel only pays for itself once it routes directly into the queue your writers and strategists work from.

Fix factual drift first. An assistant stating an old price, a discontinued plan, or a feature you no longer support costs you deals in real time, and it should jump the queue ahead of any inclusion work.

Prioritize prompts where a competitor is cited from a page you could clearly outperform. If a thin, outdated comparison page is earning citations against your product, a thorough, current version of the same comparison is often the fastest win available.

Rebuild the pages that already earn citations before you write new ones. A page an assistant already trusts enough to cite will absorb added depth faster than a brand-new page has to earn trust from scratch - the compounding effect is real and it's usually the higher-leverage move.

Report outcomes as presence gained per prompt tier, tied explicitly to the pages or placements that changed. That framing connects the work to pipeline conversations in a way a single blended score never will, and it's a natural fit for teams already running agentic AI workflows across content and demand generation.

From absence to brief: a triage workflow

Sort absences by commercial tier, then by whether the fix is a new page, an update to an existing page, or a third-party placement. Assign each to the team that owns that lever and set a re-check date tied to your sampling cadence.

Correcting inaccurate brand claims in AI answers

Publish a clear, current, well-structured page that states the correct fact plainly, and monitor the specific prompt that surfaced the error until the correction shows up in the answer.

Reporting AI visibility to revenue leadership

Lead with the tier-level presence numbers tied to commercial-intent prompts, not the blended panel score, and pair each number with the specific action taken since the last report.

Running the workflow with AstroFabric

AstroFabric ships eight specialist agents - audit, performance, market intelligence, AI visibility, pipeline, content, demand generation, and design - and the AI visibility agent is built to own exactly this loop: maintaining prompt panels, sampling answers across assistants, and scoring share of voice on a schedule.

SurfaceCitation transparencyRetrieval behaviorSampling difficultyRecommended cadencePrimary metric to track
ChatGPTLow to moderateMix of training recall and browsingModerateWeeklyMention rate and position
PerplexityHighLive retrieval, sources shownLow to moderateWeeklyCitation share
Google AI ModeModerateLive retrieval, search-linkedModerateWeeklyCitation share
ClaudeLowMostly training recallModerateMonthlyMention rate and sentiment
GeminiModerateMix of training and live resultsModerateMonthlyMention rate and position

Metered tool capabilities handle the collection work - running the panel across assistants on your chosen cadence - while the code sandbox computes share of voice, citation share, and variance bands through exact calculation rather than an estimated or eyeballed score. That distinction matters once you're reporting these numbers upward: exact arithmetic on stored raw data holds up to scrutiny in a way a rough summary doesn't.

Approval-gated writes mean that when the content agent proposes a new comparison page or an update to fix factual drift, it waits for your sign-off before anything ships. The system surfaces the recommendation and the reasoning; you keep control of what goes live.

Findings reach your team through whichever surface it already works in - the console, the REST API, MCP, the embeddable widget, email, Slack, or Telegram - so a tier-level presence drop can land as a Slack alert the same day it's detected rather than waiting for a monthly deck.

Credit-based pricing lets you scale panel size and sampling frequency to match the value of the answers you're chasing - a smaller monthly baseline panel for a stable category, or a larger weekly panel during a launch, without committing to a fixed seat count regardless of what you're actually measuring.

Which agent does what in the loop

The AI visibility agent runs the panel and scores presence and citations. The content agent drafts the fixes. The pipeline agent can connect presence changes to the deals moving through your funnel, once the numbers are in place.

Exact computation instead of estimated scores

Share of voice, weighted scores, and variance bands are computed in the code sandbox from your stored raw answers, so the same inputs always produce the same output - no rounding drift between reports.

Choosing a cadence and panel size

Start with a monthly baseline panel of 60-80 prompts to establish your position, then add a weekly, smaller high-intent panel once you're actively publishing fixes and want to see the lag.

Common measurement mistakes and how to avoid them

A single-run screenshot of a good answer feels like proof, but it's one draw from a probabilistic system. Always sample repeatedly - at least a handful of runs per prompt per cycle - before drawing any conclusion from the results.

Mixing model versions inside one trendline is one of the more common errors, since a jump in your presence rate might simply reflect a model update rather than anything you did. Tag every run with its model version and split trendlines across major version changes.

Counting any mention as a win ignores position and sentiment, which shape whether a buyer actually acts on the answer. A brand named last, with a caveat about support quality, is not the same outcome as being named first as the recommended pick, even though both count as a "mention."

Ignoring branded prompts leaves factual errors about your own product to accumulate quietly. Unbranded visibility gets the attention because it's tied to new pipeline, but branded accuracy protects deals already underway, and it deserves equal weight in the panel.

Treating AI visibility as a separate initiative from your regular content system slows down every fix that matters. The panel is only useful if absences and citation gaps flow directly into the same briefing and publishing process your content team already runs, rather than sitting in a report nobody actions.

Sample size and confidence

If repeated identical runs swing by more than a few points, treat anything inside that range as noise rather than signal.

Keeping trendlines comparable across model updates

Log model version on every run and annotate your trendline chart at each known update, so a viewer can separate "the model changed" from "our content changed."

Connecting visibility data to the publishing pipeline

Route every flagged absence or factual error directly into your content team's existing brief queue, tagged with the prompt and tier it came from, so nothing sits stranded in a monitoring dashboard.

Try it in AstroFabric

If you'd rather not build and maintain this panel by hand, AstroFabric's AI visibility agent runs the prompt panel, samples answers across ChatGPT, Perplexity, and the other assistants that matter to you, and scores share of voice and citation share with exact computation - with every proposed fix routed through approval-gated writes before it ships. Sign up to set up your first panel and start seeing where your brand stands in the answers your buyers are already reading.

Frequently asked questions

How often should I sample AI answers?

Weekly sampling suits active campaigns where you are publishing and want to see publish-to-citation lag. Monthly sampling is enough for a stable category baseline. The key is consistency: run the same panel, at the same cadence, with several repetitions per prompt so you can separate real movement from the natural variance in generated answers.

How do you calculate AI share of voice?

Count how many answers in your prompt panel mention your brand, then divide by the total brand mentions across all competitors in those same answers. Express it as a percentage per prompt tier and as a weighted total, giving higher weight to commercial prompts like comparisons and pricing. Track citation share separately, since it measures a lever you control more directly.

Can I track ChatGPT mentions manually?

Yes, and manual sampling is a reasonable way to validate your prompt set before you automate. Run 20 to 30 prompts, log the answers in a spreadsheet with model version and date, and tag brand presence by hand. Once you want repeatable trendlines and larger panels, automated collection and exact scoring save far more time than they cost.

Why do AI answers mention different brands each time?

Assistants generate answers probabilistically, retrieval results shift as the web updates, and model versions change under the same product name. Personalization and conversation memory add more movement. That is why single-run screenshots mislead and repeated sampling works: your presence rate across many runs is stable enough to trend, while any individual answer is not.

What is the difference between a mention and a citation?

A mention is your brand named in the answer text. A citation is your domain listed or linked as a source behind that answer. You can be mentioned without being cited, when the model recalls your brand from training data, and cited without being mentioned, when a page informs the answer quietly. Report both figures.

Which prompts belong in an AI visibility panel?

Include category definitions, best-of and alternatives lists, head-to-head comparisons, pricing and integration questions, objection-handling prompts, and branded queries about your own product. Unbranded prompts show whether you get included in consideration sets. Branded prompts show whether assistants describe you accurately, which matters just as much for deals already in motion.

⟨ RUN IT INSTEAD OF READING IT ⟩

Every playbook on this blog ships as a runnable mission.

Open a workspace and the playbook library is waiting - describe the outcome and the agents carry it end to end, on your plan's monthly credits.

⟨ KEEP READING ⟩
GuideAI search & GEO

What is AEO? Answer engine optimization, defined properly

AEO is the practice of making your content the answer that AI assistants and answer engines give - the definition, how it differs from SEO and GEO, how answer engines pick sources, and where to start.

Aug 14, 2026 · 8 min read