Skip to content
Implementa.

SEO for ChatGPT (GEO) · Guide 38 of 38

AI visibility benchmark: how to compare yourself with the competition, prompt by prompt

You know AI mentions you. Fine. But "it mentions you" means nothing until you put a rival next to it: does it name you more or less than your competition? Does your name come first, or theirs? Why does it cite them with data and you in passing? Without that comparison, your visibility is a lone number, and a lone number doesn’t tell you if you’re winning or losing. A head-to-head benchmark fixes it: the same set of questions, run for you and three or four competitors, read side by side. It isn’t a metric —that comes later—, it’s the process that produces one. Here’s how to set it up, how to read it without fooling yourself, and how to turn the table into a plan.

You know AI mentions you. Fine. But "it mentions you" is an orphan fact until you put a rival next to it. The question that decides isn’t "do I show up?", it’s "do I show up more or less than my competition, before or after them, with better or worse context?". A lone number doesn’t answer that. A head-to-head benchmark does: the same set of questions, run for you and three or four competitors, read side by side. Here’s how to set it up, how to read it without fooling yourself, and how to turn the table into a plan.

Your AI visibility means nothing without a rival beside it

Say you audit your presence and show up in 40% of your category’s answers. Is that good? You have no idea, and anyone who tells you it is doesn’t either. If your competitors show up in 70%, you’re last. If they show up in 15%, you dominate. The same 40% is a win or a loss depending on who you compare it to, which is why a visibility audit on its own stops halfway: it tells you where you are, not where you stand against the ones fighting you for the client.

Generative AI has made this sharper than classic SEO. On Google you saw ten links and could sit sixth without drama: the user scrolled and found you. A ChatGPT answer has room for three or four brands, not ten, and if your competitor gets in and you don’t, there’s no second page to recover on. AI visibility is more zero-sum, and in a zero-sum game the only number that matters is the relative one. Measuring your presence without measuring theirs is like timing your lap without knowing the other drivers’: you know your time, not whether you win.

A benchmark isn’t share of voice: the process vs the metric

It’s worth not mixing two things that sound alike. The benchmark is the process: running the same set of questions for you and your rivals and noting who shows up, where and how. Share of voice is one of the metrics that come out of that process —the percentage of answers where you appear versus the group total—. The benchmark is the method; share of voice, one of the numbers it produces. One is the kitchen, the other is a dish.

The distinction isn’t pedantic, it changes what you do. If you chase share of voice as a metric, you end up staring at a percentage and celebrating or suffering as it rises or falls, without knowing what moves it. If you run the benchmark as a process, you see the raw answers: which question leaves you out, with what phrase they describe the competitor you’re missing, which source AI leans on to cite them. That qualitative reading —the "why"— is what a percentage hides. Share of voice tells you you’re losing; the benchmark tells you where they’re beating you. This guide is about the second, because it’s the one that turns into tasks.

Benchmark (process)Share of voice (metric)
What it isRunning the same prompt set for you and N rivalsThe % of your mentions over the group total
What it gives youWho shows up, in what order, with what context and sourceA number comparable over time
What it’s forDeciding what to fix and whereWatching the trend and reporting to leadership
RelationshipIt’s the method that produces the metricIt’s one of the method’s outputs

How to set up the benchmark: same prompt set, you and 3-4 competitors

An honest benchmark rests on a simple rule: everything the same except the player. Same questions, same engines, same window, same day. If you change the set between one competitor and another, you’re not comparing, you’re inventing. Four steps:

  1. Pick the real rivals (3-4). Don’t include the whole sector: include who a buyer of yours would weigh in the same decision. Start with the ones that already show up when you ask about your category and add one or two you know fight you for clients even if they don’t appear yet. If an unexpected name comes up when you run the questions, bring it in: AI is telling you who your competition is in its eyes, which doesn’t always match yours.
  2. Build the brand-free battery (10-15 questions). As in an audit, the questions carry no name —yours or anyone’s—: they’re the ones someone looking for a solution asks, not a brand. "What’s the best software for X?", "which company do you recommend for Y?", "alternatives to [category leader]". That way you measure who gets in on merit, not who shows up once you name them yourself. Save them: it’s your fixed set, the one you’ll rerun identical each cycle.
  3. Fire the same set in each engine, in incognito. ChatGPT, Perplexity and Google AI read different internets, so run all three. Incognito window or logged out so your history doesn’t contaminate the answer. One pass per engine and question; if an engine gives very volatile answers, two passes and you keep what repeats.
  4. Record in a table: brand, position and context. For each question and engine, note which brands show up, in what order they appear within the answer, and —this is what almost no one does— with what phrase AI accompanies them. "X, the most complete option" isn’t the same as "X, more expensive". Position gives you the ranking; the phrase gives you the why.

How to read the results: who shows up, in what position, and why

The filled table isn’t the goal; it’s the raw material. Read it in three layers, from the dumbest to the most useful:

  • Presence: who appears? The basic layer. Count in how many of the N questions each brand shows up. Here’s where, if you want, you compute share of voice: your appearances over the total. It’s the headline, but it’s the least actionable figure.
  • Position: in what order? Coming first in the answer isn’t the same as coming last in a list of five. AI orders by perceived confidence, so each brand’s average position across the set tells you who’s the category’s "default answer" and who’s filler. Climbing in position is worth more than adding one mention further down.
  • Context: with what face? The layer almost no one looks at and the one that decides most. Collect the phrases AI uses to present each one. If it describes your competitor with a concrete fact ("integrates with 200 tools", "the fastest to deploy") and you with a generic, there’s the gap: it’s not that they don’t name you, it’s that they have nothing to say about you. That’s fixable, and it’s fixed off your website, in the sources AI cites.

Cross the three layers and the diagnosis appears. A competitor that shows up a lot, high and with concrete facts is beating you structurally: it has reputation in the sources the model weights. One that shows up a lot but low and list-style is vulnerable: presence without authority. And if you show up little but when you do you land well positioned, you have a base to scale. The benchmark doesn’t give you a grade, it gives you a map of where to attack. And that reading is exactly the starting point of how to measure AI visibility with a clear head: first the map, then the metrics.

From the table to the plan: what to do with what you find

A benchmark that ends in a pretty table is an autopsy. The value is in the three or four decisions that come out of it. Translate each finding into an action:

  • "The competitor shows up with a fact I haven’t published" → produce that fact in citable form. If they’re cited for "integrates with X" and you integrate too but haven’t told it in any source AI reads, the job is to win that mention, not lament it.
  • "I show up, but always last" → it’s an authority problem, not a presence one. Reinforce your weight in the third-party sources —reviews, comparisons, forums— where the model calibrates trust. Climbing in position is reputation work, not copy on your landing.
  • "I don’t show up in the questions that sell most" → prioritize those. Not every question in the set is worth the same: the ones closest to purchase are the ones to win first, even if you lose an informational one.
  • "A competitor I didn’t have on the radar dominates" → study why. Which sources cite it, with what phrase, in which engine. The benchmark just handed you competitive intelligence you didn’t have.

All of this can be done by hand, and at first it should be: understanding the benchmark with your own hands is worth more than any dashboard. But the spreadsheet gets small for the usual signals: the volume of questions, engines and languages eats your morning; you need automatic history to see whether the gap with each rival opens or closes; or you want continuous monitoring instead of a quarterly snapshot. There, delegating AI visibility monitoring leaves you the benchmark alive and compared without transcribing by hand. And when the table calls for action —win the mention you’re missing, climb in position, get into the question that sells—, GEO optimization does the work on the sources, not another slide. If you’re starting here, the full cluster map lives in the SEO for ChatGPT guide.

Frequently asked questions

They’re different links in the same chain. The benchmark is the process: you run the same set of questions for you and your competitors in ChatGPT, Perplexity and Google AI, and note who shows up, in what position and with what context. Share of voice is one of the metrics you pull out at the end —the percentage of answers where you appear versus the group total—. You can run the benchmark and keep only "who appears and why" without computing any percentage; and you can’t compute an honest share of voice without having run the benchmark first. One is the method, the other is the number. This guide is about the method.

Three or four, and chosen, not every name that comes to mind. Include the ones a buyer of yours would genuinely weigh in the same decision: those that show up when AI answers your category question, plus one or two you know fight you for clients even if they don’t appear yet. Fewer than three and you see no pattern; more than five and the table becomes unmanageable and your focus blurs. The goal isn’t a census of the sector, it’s understanding who you lose the mention to and why. If an unexpected name shows up when you run the questions, add it: AI is telling you who your competition is in its eyes.

Every four to eight weeks, aligned with the pace at which models change. A one-day benchmark is a snapshot: it tells you how you stand today against your rivals, not whether the gap is widening or closing. The series is what matters. Rerun the same set of questions, with the same reading rule, and save the date: that’s how you tell a real improvement from a model wobble. Rerunning weekly is spending time on noise; leaving it as a yearly snapshot means learning too late that a competitor has overtaken you. The middle ground —monthly or bimonthly— is where the benchmark becomes a thermometer, not an anecdote.

Yes, and it’s best to start that way. With a battery of 10-15 category questions, an incognito window and a spreadsheet you have the basics: fire each question in ChatGPT, Perplexity and Google AI, note which brands show up and in what order, and compare. The paid tool doesn’t give you a different truth; it saves you the manual transcription, stores the history and lets you scale to more questions, more languages and more competitors without burning the morning. Start by hand to understand what you’re measuring; delegate when the volume eats your time or when you need continuous monitoring instead of a quarterly snapshot.

Free AI Impact Plan

The guide is generic. Your plan isn't.

Tell us about your company and we'll ship back a diagnosis with priorities, numbers and what to implement first. No sales call, no charge.

AI visibility benchmark: how to compare yourself with the competition, prompt by prompt · Implementa