You know AI mentions you. Fine. But "it mentions you" is an orphan fact until you put a rival next to it. The question that decides isn’t "do I show up?", it’s "do I show up more or less than my competition, before or after them, with better or worse context?". A lone number doesn’t answer that. A head-to-head benchmark does: the same set of questions, run for you and three or four competitors, read side by side. Here’s how to set it up, how to read it without fooling yourself, and how to turn the table into a plan.
Your AI visibility means nothing without a rival beside it
Say you audit your presence and show up in 40% of your category’s answers. Is that good? You have no idea, and anyone who tells you it is doesn’t either. If your competitors show up in 70%, you’re last. If they show up in 15%, you dominate. The same 40% is a win or a loss depending on who you compare it to, which is why a visibility audit on its own stops halfway: it tells you where you are, not where you stand against the ones fighting you for the client.
Generative AI has made this sharper than classic SEO. On Google you saw ten links and could sit sixth without drama: the user scrolled and found you. A ChatGPT answer has room for three or four brands, not ten, and if your competitor gets in and you don’t, there’s no second page to recover on. AI visibility is more zero-sum, and in a zero-sum game the only number that matters is the relative one. Measuring your presence without measuring theirs is like timing your lap without knowing the other drivers’: you know your time, not whether you win.
A benchmark isn’t share of voice: the process vs the metric
It’s worth not mixing two things that sound alike. The benchmark is the process: running the same set of questions for you and your rivals and noting who shows up, where and how. Share of voice is one of the metrics that come out of that process —the percentage of answers where you appear versus the group total—. The benchmark is the method; share of voice, one of the numbers it produces. One is the kitchen, the other is a dish.
The distinction isn’t pedantic, it changes what you do. If you chase share of voice as a metric, you end up staring at a percentage and celebrating or suffering as it rises or falls, without knowing what moves it. If you run the benchmark as a process, you see the raw answers: which question leaves you out, with what phrase they describe the competitor you’re missing, which source AI leans on to cite them. That qualitative reading —the "why"— is what a percentage hides. Share of voice tells you you’re losing; the benchmark tells you where they’re beating you. This guide is about the second, because it’s the one that turns into tasks.
| Benchmark (process) | Share of voice (metric) | |
|---|---|---|
| What it is | Running the same prompt set for you and N rivals | The % of your mentions over the group total |
| What it gives you | Who shows up, in what order, with what context and source | A number comparable over time |
| What it’s for | Deciding what to fix and where | Watching the trend and reporting to leadership |
| Relationship | It’s the method that produces the metric | It’s one of the method’s outputs |
How to set up the benchmark: same prompt set, you and 3-4 competitors
An honest benchmark rests on a simple rule: everything the same except the player. Same questions, same engines, same window, same day. If you change the set between one competitor and another, you’re not comparing, you’re inventing. Four steps:
- Pick the real rivals (3-4). Don’t include the whole sector: include who a buyer of yours would weigh in the same decision. Start with the ones that already show up when you ask about your category and add one or two you know fight you for clients even if they don’t appear yet. If an unexpected name comes up when you run the questions, bring it in: AI is telling you who your competition is in its eyes, which doesn’t always match yours.
- Build the brand-free battery (10-15 questions). As in an audit, the questions carry no name —yours or anyone’s—: they’re the ones someone looking for a solution asks, not a brand. "What’s the best software for X?", "which company do you recommend for Y?", "alternatives to [category leader]". That way you measure who gets in on merit, not who shows up once you name them yourself. Save them: it’s your fixed set, the one you’ll rerun identical each cycle.
- Fire the same set in each engine, in incognito. ChatGPT, Perplexity and Google AI read different internets, so run all three. Incognito window or logged out so your history doesn’t contaminate the answer. One pass per engine and question; if an engine gives very volatile answers, two passes and you keep what repeats.
- Record in a table: brand, position and context. For each question and engine, note which brands show up, in what order they appear within the answer, and —this is what almost no one does— with what phrase AI accompanies them. "X, the most complete option" isn’t the same as "X, more expensive". Position gives you the ranking; the phrase gives you the why.
How to read the results: who shows up, in what position, and why
The filled table isn’t the goal; it’s the raw material. Read it in three layers, from the dumbest to the most useful:
- Presence: who appears? The basic layer. Count in how many of the N questions each brand shows up. Here’s where, if you want, you compute share of voice: your appearances over the total. It’s the headline, but it’s the least actionable figure.
- Position: in what order? Coming first in the answer isn’t the same as coming last in a list of five. AI orders by perceived confidence, so each brand’s average position across the set tells you who’s the category’s "default answer" and who’s filler. Climbing in position is worth more than adding one mention further down.
- Context: with what face? The layer almost no one looks at and the one that decides most. Collect the phrases AI uses to present each one. If it describes your competitor with a concrete fact ("integrates with 200 tools", "the fastest to deploy") and you with a generic, there’s the gap: it’s not that they don’t name you, it’s that they have nothing to say about you. That’s fixable, and it’s fixed off your website, in the sources AI cites.
Cross the three layers and the diagnosis appears. A competitor that shows up a lot, high and with concrete facts is beating you structurally: it has reputation in the sources the model weights. One that shows up a lot but low and list-style is vulnerable: presence without authority. And if you show up little but when you do you land well positioned, you have a base to scale. The benchmark doesn’t give you a grade, it gives you a map of where to attack. And that reading is exactly the starting point of how to measure AI visibility with a clear head: first the map, then the metrics.
From the table to the plan: what to do with what you find
A benchmark that ends in a pretty table is an autopsy. The value is in the three or four decisions that come out of it. Translate each finding into an action:
- "The competitor shows up with a fact I haven’t published" → produce that fact in citable form. If they’re cited for "integrates with X" and you integrate too but haven’t told it in any source AI reads, the job is to win that mention, not lament it.
- "I show up, but always last" → it’s an authority problem, not a presence one. Reinforce your weight in the third-party sources —reviews, comparisons, forums— where the model calibrates trust. Climbing in position is reputation work, not copy on your landing.
- "I don’t show up in the questions that sell most" → prioritize those. Not every question in the set is worth the same: the ones closest to purchase are the ones to win first, even if you lose an informational one.
- "A competitor I didn’t have on the radar dominates" → study why. Which sources cite it, with what phrase, in which engine. The benchmark just handed you competitive intelligence you didn’t have.
All of this can be done by hand, and at first it should be: understanding the benchmark with your own hands is worth more than any dashboard. But the spreadsheet gets small for the usual signals: the volume of questions, engines and languages eats your morning; you need automatic history to see whether the gap with each rival opens or closes; or you want continuous monitoring instead of a quarterly snapshot. There, delegating AI visibility monitoring leaves you the benchmark alive and compared without transcribing by hand. And when the table calls for action —win the mention you’re missing, climb in position, get into the question that sells—, GEO optimization does the work on the sources, not another slide. If you’re starting here, the full cluster map lives in the SEO for ChatGPT guide.