Skip to content
Implementa.

SEO for ChatGPT (GEO) · Guide 34 of 34

Prompts to audit your AI visibility: how to build the battery that makes the audit worth anything

Most people build a GEO audit with four questions written from memory, put their brand name in all of them, and then act surprised that the AI "names them a lot." Of course it does: you asked it to. The audit isn’t worth what the tool or the report is worth; it’s worth what the prompts you ask with are worth. A bad set of questions gives you a pretty, false number; a good one tells you the uncomfortable truth of whether you exist for the model when the buyer asks without knowing you. Here’s how to build that battery, question by question.

Most people build a GEO audit with four questions written from memory, put their brand name in all of them, and then act surprised that the AI "names them a lot." Of course it does: you asked it to. The audit isn’t worth what the tool or the report is worth; it’s worth what the prompts you ask with are worth. A bad set of questions gives you a pretty, false number; a good one tells you the uncomfortable truth of whether you exist for the model when the buyer asks without knowing you. Here’s how to build that battery, question by question.

Why a GEO audit is only as good as its prompts

In a GEO audit the prompt is the measuring instrument, not a detail. If the thermometer is miscalibrated, it doesn’t matter how nice the chart looks: the number lies. And the most common bias is putting your brand in the question —"what do you think of [my company]?"—: that measures reputation, not visibility. The AI will talk about you because you asked, not because it recommends you on its own when the buyer doesn’t mention you.

Real visibility plays out in the questions where your name never appears. That’s where you get in on merit or you don’t. That’s why designing the prompt set is 80% of a well-run audit: it defines what you measure, against whom, and in what language. The rest —running it, logging, scoring— is mechanics. This is the work almost nobody does, and the one that separates a useful audit from theatre with screenshots.

The three prompt types you can’t skip

A serious battery covers three question families. Each measures something different, and if one is missing you measure a slice and think you measure the whole.

TypeWhat it asks (without your brand)What it measures
Category"Best tools for X", "how to choose a Y provider"Whether you get into the answer on merit when no one names you.
Comparative"A vs B", "alternatives to Z", "is X worth it or better Y?"Your relative position against the competitors you’re disputed with.
Problem"How to stop wasting time on...", "why does my X keep failing"Whether you capture the buyer who has the pain but doesn’t know your category yet.

Category ones are your existence baseline. Comparative ones measure share of voice: showing up isn’t enough, it matters whether it’s you or your rival. And problem ones are the highest-value and the ones almost nobody writes, because they force you out of internal jargon and into the language of the customer who doesn’t know you yet. That’s where the earliest, least-disputed buyer is.

How to write prompts the AI answers like a real buyer

A good audit prompt mimics how a person with your problem actually asks, not how you’d describe your product. Five rules so the answer stays honest:

  1. No brand (except the sentiment block). If your name is in the question, you’re not measuring visibility, you’re measuring echo. Reserve two or three branded prompts only to read the tone the AI uses about you, and don’t mix them with the presence metric.
  2. In the customer’s language, not yours. The buyer doesn’t write "document-workflow automation solution"; they write "how to stop entering invoices by hand." Steal the vocabulary from your support tickets, your sales calls and your reviews: that’s how people actually search.
  3. Neutral, no answer baked in. "What’s the best cheap tool in English for small businesses?" has already filtered the answer with your criteria. Ask open and let the model choose; if you narrow, narrow the way the buyer would, not the way that suits you.
  4. Intent-specific, not generic. "Marketing" measures nothing; "how to tell if ChatGPT recommends my store" does. The closer the prompt is to a concrete decision, the more useful the reading.
  5. Repeatable word for word. Each prompt has to be relaunchable identical a week later. If you improvise it each time, you don’t have a time series, you have loose anecdotes.

How many and how to split them: the 15-25 question battery

The healthy size is between 15 and 25 fixed questions. Fewer than 15 and the result is anecdotal —one day you show up, the next you don’t, and you can’t tell signal from noise—. More than 30 and you won’t rerun it by hand each week, which is exactly what makes it useful as a baseline. A split that works:

FamilyApprox. weightIn a battery of 20
Category (no brand)~50%10 questions
Comparative~30%6 questions
Problem~15%3 questions
Brand (sentiment only)~5%1 question

The mistakes that invalidate the battery (and how to avoid them)

Four mistakes turn an audit into a flattering, false photo. Commit one and the rest of the report is moot:

  1. Putting the brand in everything. The classic. You measure that the AI can repeat your name, not that it recommends you. Take it out of the presence questions.
  2. Questions that answer themselves. The prompt with the answer baked in ("why is [my brand] the best for X?") doesn’t audit, it applauds. Neutrality or nothing.
  3. Tiny sample. Four questions aren’t an audit, they’re an impression. Without volume you can’t tell the trend from the noise of a model that varies between answers.
  4. Changing the set each time. The silent mistake: it looks like you improved, but you only changed the ruler. Fix the prompts for at least a quarter; add without removing.

With the battery well built, the rest is execution: run the audit for today’s snapshot, track mentions week by week for the movie, and score each answer with the framework in how to measure your AI visibility. If you don’t want to build and maintain the battery by hand, delegating AI visibility monitoring keeps the instrument live and watched; and when the reading calls for action, GEO optimization works on the content, not the report.

Frequently asked questions

Between 15 and 25 fixed questions. Fewer than 15 and the result is anecdotal: one day you show up, the next you don’t, and you can’t tell signal from noise. More than 30 and the battery becomes impossible to rerun by hand each week, which is exactly what makes the audit useful as a baseline. A healthy split is roughly half category questions (no brand), a third comparative against competitors and the rest problem-based —how someone who doesn’t know you yet phrases the search—. The exact count matters less than two things: that it covers all three families and that it’s always the same, because a battery you change every week doesn’t measure, it guesses.

In most of them, no. The most expensive mistake in a GEO audit is asking "what do you think of [your brand]?": the AI will talk about you because you asked, not because it recommends you on its own. What you measure that way is reputation, not visibility. Real visibility is measured with category questions without your name —"best tool for X", "how to solve Y"— counting whether you show up on merit. Keep a couple of branded prompts only to measure sentiment (the tone the AI uses about you when asked directly), but don’t let them contaminate the presence metric, which is the one that actually tells you whether you exist for the model.

Three families. Category: brandless questions about your type of product or service ("best platforms for X", "how to choose Y") —they reveal whether you get into the answer on merit—. Comparative: questions that pit options against each other ("A vs B", "alternatives to Z") —they measure your relative position against competitors—. And problem: questions in the language of someone who has your problem but doesn’t yet know you or your category exist ("how to stop wasting time on...", "why does my X keep failing") —the highest-value ones, because they capture the buyer before they know who to look for—. Without all three, the audit measures a slice and thinks it measures the whole.

As rarely as possible. The battery is your measuring instrument, and changing it breaks the week-to-week comparison: if you ask differently this week than last, you can’t tell whether the mention rose because you improved or because you changed the question. Fix the set and rerun it identical for at least a quarter. You only revise it when something real changes in your business —you enter a new category, a relevant competitor appears— and even then you add new questions without touching the old ones, so you don’t lose the history. An instrument you recalibrate every week doesn’t measure: each reading comes out on a different scale.

Free AI Impact Plan

The guide is generic. Your plan isn't.

Tell us about your company and we'll ship back a diagnosis with priorities, numbers and what to implement first. No sales call, no charge.

Prompts to audit your AI visibility: how to build the battery that makes the audit worth anything · Implementa