Questions to Ask a GEO Agency Before You Sign

The vendor-neutral checklist for hiring a GEO agency: measurement proof, prompt methodology, pricing terms, and the red flags that end the call early.

ArticleBY THE ASTROFABRIC TEAM · AUG 21, 2026 · 11 MIN READ

Abstract dark illustration of a luminous checklist grid connecting to a network of glowing nodes, representing due-diligence questions for vetting a GEO agency

Before you hire a geo agency, make measurement the deciding filter. Ask any GEO agency to show live citation share across named AI engines, the prompt set behind the number, and how it recomputes each month. Firms that answer with reproducible, platform-backed data earn a second call; firms that answer with adjectives are selling reports. What follows is the full checklist - methodology, deliverables, pricing, and exit terms - in the order a careful buyer should ask them.

What Does a GEO Agency Actually Sell?

Every pitch in this market sounds the same right now. AI-first, answer-engine-ready, built for the ChatGPT era - the vocabulary has been commoditized so thoroughly that your checklist has to test substance, because the language never will.

Here is the honest definition of the deliverable. A GEO agency sells changes to how AI assistants describe and cite your brand, which means the real work is content built for answer engines, technical structure that makes pages quotable, and entity clarity sharp enough that the models know exactly who you are. TechTarget's coverage of generative engine optimization frames it the same way: influencing what generative systems say is a different discipline from influencing where a blue link ranks. Some firms are genuinely instrumented for that work. Others are rebadging a classic SEO retainer, and our AEO agency guide is the deeper buy-side comparison for telling those two apart at the proposal stage.

GEO, AEO, and AI SEO agency: same buyer, three labels

You will hear GEO agency, AEO agency, and AI SEO agency used almost interchangeably, and for buying purposes they are the same conversation. Whatever the label, you are evaluating one thing: can this firm move the way assistants answer your buyers' questions, and can it prove the movement happened?

The one-sentence test for any pitch deck

One vivid example beats any taxonomy. I sat through a pitch where the deck promised "AI-optimized content across every major assistant," and when I asked to see a single tracked prompt, the room went quiet. That agency was describing hope. This entire post exists to catch that moment in the first call instead of month four.

How Should a GEO Agency Prove Its Measurement?

This is the deciding filter, so ask it first and let weak answers end the process early. Measurement quality predicts everything else. A firm that measures rigorously tends to scope honestly and price transparently, and it hands data over cleanly when the engagement ends, because rigor is a habit rather than a feature.

The proof to demand is specific: a live view of citation share or mention share across named engines, the exact prompt set behind it, and a stated cadence for recomputing the number. Ideally they open a current client's dashboard on the call with the identifying details anonymized. Screenshots pasted into a quarterly slide do not count. A slide is a claim; a dashboard is evidence.

Ask to see the prompt set, then ask who wrote it

The prompt set is the measurement instrument, so inspect it the way you would inspect any instrument. Who wrote these prompts? Do they reflect how your buyers actually phrase questions, or are they keywords with a question mark stapled on? A strong agency talks about prompt design the way a researcher talks about survey design, with visible care about bias and coverage.

Reproducible numbers versus hand-assembled decks

Platform-backed measurement matters because a number computed the same way every month, in tooling you can inspect, beats any report assembled by hand. That is the problem AI visibility tools exist to solve, and it helps to know what the high end looks like as a reference point. On AstroFabric, for instance, an AI visibility agent runs metered prompt checks and computes the resulting share numbers exactly in a code sandbox, so the figure is reproducible rather than estimated. The agency does not need to run that specific stack. They need to clear that bar with whatever they run.

The one question that does most of the work

"Show me the number, the prompts behind it, and how it recomputes next month." An agency that answers fluently has probably done everything else right too.

Which engines they track and why the list matters

Engines diverge, sometimes dramatically, on the very same question. A firm tracking only ChatGPT is reporting one weather station and calling it the climate. The credible answer covers ChatGPT, Perplexity, Gemini, and Google AI Mode at minimum, with a working explanation of why the variance exists and which engines matter most for your buyers.

The Methodology Questions: Prompts, Engines, and Sampling

Once you know they measure, the next call tests whether the measurement is sound. Sample size is the honest tell. Assistants answer the same prompt differently from one day to the next, so a visibility number built on a handful of one-off checks is noise wearing a suit.

4minimum engines a credible tracking program should cover

Ask to see their math. We walked through the statistics in how many prompts you need, and a serious firm will have done its own version of that reasoning - prompts per topic, runs per prompt, and a confidence story behind the resulting share. If the answer you get is "we check the big queries weekly," you have learned what you needed to learn.

How many prompts, sampled how often

There is no single magic number, but a good answer has a recognizable shape: enough prompts to cover your real buying questions, enough repeated runs to smooth out model randomness, and a refresh process that lets the set evolve as buyer language shifts. A static prompt set written six months ago measures a market that no longer exists.

Mentions, citations, and which they report

Mentions and citations are different assets. A mention is the assistant naming you in its answer. A citation is the assistant grounding that answer in your page as a source. Ask which one they optimize first, and why. Expect a reasoned answer - many strong firms chase citations first because grounded sources tend to stick, then treat mentions as the trailing indicator. A shrug at this question is a shrug at the whole discipline. Microsoft's own documentation on grounding and retrieval is a useful primer on why the distinction is architectural rather than cosmetic.

Deliverables and Timelines: What Good Generative Engine Optimization Services Look Like

Now for expectations, because this is where the overpromising lives. GEO compounds over months. Assistants re-crawl and re-ground on their own schedule, and no agency controls that schedule. A firm promising visibility jumps inside two weeks is selling against physics, and you should hear that promise as a warning bell rather than ambition.

Month one versus month six

Month one should be diagnosis and structural work: an entity and citability audit, the prompt baseline, the first round of fixes that make your pages quotable. Months two through six are where content mapped to real buyer questions ships steadily and the visibility report starts showing movement against the agreed prompt set. Ask for that arc explicitly. Specific answers signal an operating firm; vague ones signal a retainer wearing a new label.

The proposal line items that signal real work

The practitioner's version of this is concrete. The best proposal I have seen did not open with a twelve-line scope list. It opened with "here is the exact FAQ block we would restructure on your pricing page, and here is why an assistant currently skips it." One deliverable, named and defended, tells you more than a page of scope language ever will.

What a serious GEO proposal contains
  • An entity and citability audit with named pages and fixes
  • A prompt baseline across at least four engines
  • Content mapped to documented buyer questions
  • Structural changes that make key pages quotable
  • A recurring visibility report tied to the agreed prompt set
  • A stated monthly cadence and review rhythm

Pricing, Contracts, and Exit Terms

Buyers dazzled by a good demo forget the commercial questions, so write them down before the call. What exactly does the retainer cover? What costs extra? How does scope flex when priorities shift mid-quarter? None of these are gotchas - they are the questions a confident firm answers without flinching, and hesitation here is information too.

Who owns the prompt set and the history

Everything the engagement produces should transfer to you on exit: prompt sets, dashboards, content, and the full tracking history. Get that in the contract, in plain language. The measurement history is the asset that makes the whole engagement auditable, and losing it at termination resets your baseline to zero. While you are pricing the relationship, it helps to know what the underlying software costs on its own - our breakdown of GEO tools pricing shows how those platform fees typically stack, which makes it much easier to see what margin you are paying for judgment versus tooling.

Tying fees to a metric you can verify

Ask for the termination story before you sign: notice period, data handover format, and whether reporting continuity survives the relationship. Then anchor the fee itself to measurement. If the agency cannot tie what you pay to a tracked visibility metric, you are purchasing activity, and activity has a way of expanding to fill any retainer.

GEO Agency vs Platform: Which Should You Buy First?

The honest decision logic is simpler than most comparison posts admit. Teams with strategy in-house and hands available usually do better running GEO tools directly, because they keep the learning and skip the margin. Teams that need judgment and execution bought together are the true agency customer, and there is nothing wrong with being one.

When the agency is the right call

Buy the agency when you lack the time or the internal conviction to run the program yourself, when you want a practitioner who has already seen a dozen of these engagements, or when your team is stretched thin enough that even a great platform would sit unused.

When you should own the scoreboard yourself

Increasingly, the right answer is hybrid: you own the platform-backed measurement, and the agency executes against it. The scoreboard stays neutral, every renewal becomes a data conversation, and switching agencies never costs you your baseline. That last part is the core move of this checklist. The measurement layer should outlive the vendor relationship, because that is precisely what makes vendors accountable.

The third option worth knowing

Agentic platforms now sit between tools and agencies: specialist agents handle audit, content, and AI visibility work directly, with approval-gated writes so nothing ships without your sign-off. AstroFabric runs eight of them. Worth evaluating before you commit a retainer.

The Red Flags That Should End the Call

Keep this list light but firm, because each flag has a positive mirror that a strong agency hits naturally:

  • Guaranteed placement in AI answers. Nobody controls model output. The good version sounds like: "here is our baseline, here is the movement we have driven before, here is how we will measure yours."
  • Refusal to share methodology. The good version volunteers prompt counts and engine lists before you think to ask.
  • Case studies with no numbers. The good version shows citation share over time, even anonymized.
  • Reports that never name engines. The good version tells you exactly which assistants were sampled, and when.
  • The subtle one: a firm fluent about rankings that goes quiet on citation share, grounding, or how assistants choose sources is an SEO shop wearing a new jacket. The good version can explain grounding to your CFO in two sentences.

The Complete Question List, Ready for Your Next Call

Group your questions by call stage and the conversation practically runs itself. Discovery: what do you sell, which engines do you track, can you show me a live dashboard. Methodology deep-dive: who wrote the prompt set, how many prompts and runs sit behind each number, how often the set refreshes, mentions or citations first. Commercial: what the retainer covers, who owns the data on exit, how the termination story reads.

Ten questions, in the order to ask them

STRONG VS WEAK ANSWERS
QuestionStrong answerWeak answer
Can you show measurement live?Opens a dashboard on the callPromises a deck later
Which engines do you track?Names four or more, explains variance"The major ones"
Who wrote the prompt set?Built from our buyer research, refreshed quarterlyKeyword list with question marks
How many prompts and runs?Shows the sampling math"We check the big queries"
Mentions or citations first?Reasoned sequencing with a whyTreats them as the same thing
What ships in month one?Named audits and structural fixes"Strategy alignment"
When does visibility move?One to two quarters, compoundingTwo-week guarantees
What does the retainer cover?Line items tied to the prompt setFlexible hours
Who owns data on exit?Everything transfers, in the contract"We can discuss that"
Are fees tied to a metric?Yes, the tracked visibility numberTied to activity delivered

What to do with the answers

Score the answers honestly, and let two or more weak ones end the process. The buyer who insists on platform-backed, reproducible measurement turns agency selection from a personality contest into an evidence review, and that single filter does most of the work.

See the Measurement Bar for Yourself

The fastest way to calibrate is to look at real instrumentation before your next agency call. AstroFabric's AI visibility agent runs metered prompt checks across engines and computes share numbers exactly in a code sandbox, so you will know precisely what a reproducible baseline looks like and what to demand from anyone who wants your retainer. Start free at /signup and walk into that call holding the scoreboard.

Frequently asked questions

What does a GEO agency actually do?

A GEO agency works to improve how AI assistants like ChatGPT, Perplexity, and Gemini describe and cite your brand. The real work covers content structured for answer engines, technical fixes that make pages quotable, entity clarity, and ongoing measurement of citation share across engines. Strong firms show tracked visibility data; weaker ones repackage classic SEO retainers under new language.

What is the single most important question to ask before hiring a GEO agency?

Ask them to show you their measurement live: citation share or mention share across named engines, the prompt set behind it, and the recomputation cadence. Measurement quality predicts everything else about the engagement. A firm with reproducible, platform-backed numbers can be held accountable each month, and a firm without them is asking you to buy on trust.

How long should a GEO agency take to show results?

Expect a phased arc rather than a sprint. Audits and structural fixes land in the first month, content built for answer engines ships steadily after that, and visibility movement typically compounds over one to two quarters as assistants re-crawl and re-ground sources. Any agency promising citation jumps within two weeks is overpromising against how these systems actually update.

Should I hire a GEO agency or buy a platform instead?

Buy the agency when you need judgment and execution together; run a platform directly when you have strategy in-house and hands available. The strongest arrangement is often hybrid: the client owns platform-backed measurement while the agency executes against it. That keeps the scoreboard neutral, makes renewals a data conversation, and lets you switch vendors without losing your baseline.

What red flags disqualify a GEO agency immediately?

Guaranteed placement in AI answers, refusal to share prompt methodology, case studies without numbers, and reports that never name which engines were measured. A subtler tell is fluency about rankings paired with silence on citation share and how assistants choose sources. Each flag has a positive mirror: strong agencies volunteer methodology, engine coverage, and measurable baselines without being pushed.

Who should own the measurement data when the engagement ends?

You should. Prompt sets, dashboards, visibility history, and all produced content should transfer to the client on exit, and the contract should state this explicitly before signing. Owning the measurement layer is what makes agencies accountable, since it lets you change vendors while keeping the baseline that proves whether the work is compounding or stalling.

Sources

⟨ RUN IT INSTEAD OF READING IT ⟩

Every playbook on this blog ships as a runnable mission.

Open a workspace and the playbook library is waiting - describe the outcome and the agents carry it end to end, on your plan's monthly credits.

⟨ KEEP READING ⟩
ArticleAI search & GEO

How Many Prompts to Measure AI Visibility? The Math

How many prompts, how often, and across how many engines before your AI citation share is trustworthy. Worked sample-size math, a cadence table and defaults.

Aug 14, 2026 · 11 min read