On this page
You asked ChatGPT for “the best tools for X” and your product was not in the list, but two competitors were. That stings, but on its own it is not a benchmark. It is one answer, from one model, on one day. To decide what to fix, you need to know how big the gap actually is, where it is concentrated, and whether it is moving. This article walks through a benchmark method that a solo founder can run in an afternoon and repeat every month.
What “benchmarking” actually means here
In SEO, benchmarking against a competitor is easy to picture: same keyword, who ranks higher. In AI answers there is no ranked list of ten links. There is a generated response that may or may not name you, may name you first or fifth, and may say something warm or lukewarm about you.
So a useful benchmark compares four things, always on the same set of prompts:
- Mention rate: the share of runs where a brand is named at all. This is the raw input to an AI visibility score.
- Position: when the answer is a list, where each brand lands. Being named is not the same as being named first.
- Sentiment: is the brand recommended, described neutrally, or mentioned with a caveat (“good, but expensive”).
- Citations: which URLs the engine leans on when it talks about each brand. This tells you where the visibility is coming from.
Everything else in this article is about getting those four numbers to be comparable across brands, rather than a pile of anecdotes.
Step 1: Pick competitors that buyers actually compare you to
The most common mistake is benchmarking against the company you wish you were competing with. Pick the set the way a buyer would:
- Two or three direct alternatives that show up in your sales calls, churn reasons, and “vs” searches. These are the ones where a visibility gap costs you deals today.
- One category leader as the ceiling. You will not beat them across the board this quarter, but their numbers tell you which prompts are effectively closed.
- Optionally one adjacent product that solves the same job differently. If LLMs keep recommending it for your prompts, that is a positioning problem, not a content problem.
Keep it to three to five names. Every extra competitor multiplies the number of runs you need before the numbers stop wobbling.
Step 2: Build one shared prompt set
The benchmark is only fair if every brand is scored on the same questions. Do not use “your” prompts for you and “their” prompts for them. Build one list and run it for all brands at once. A workable structure for a small SaaS:
| Prompt type | Example | Why it matters |
|---|---|---|
| Category recommendation | ”best invoicing tools for freelancers” | Broadest exposure, most competitive |
| Use-case specific | ”invoicing tool that handles EU VAT for a one-person business” | Where smaller products win |
| Comparison | ”Brand A vs Brand B for freelancers” | Shows how the model frames you against a named rival |
| Problem-first | ”how do I stop chasing late invoices” | Tests whether you are associated with the pain, not just the category |
| Brand-direct | ”is Brand A any good” | Sentiment and accuracy check |
Aim for 30 to 60 prompts to start. If you are unsure about the count, the prompt volume glossary entry covers the trade-off between coverage and noise in more detail. Write prompts the way a person types them, not the way a marketer writes a headline.
Step 3: Sample properly, or the numbers will lie
LLM answers are not deterministic. Ask the same model the same question three times and you can get three different lists. If you run each prompt once, a brand’s “visibility” can swing by 20 or 30 points between runs with nothing changing in the real world.
Three rules that fix most of the noise:
- Repeat each prompt several times per engine. Three to five samples per prompt is the range that turns a coin flip into a usable rate. This is what multi-sampling means in practice.
- Run all brands in the same batch. Same prompts, same day, same model versions. A benchmark taken for you in July and for a competitor in August is not a benchmark.
- Cover more than one engine. ChatGPT, Claude, Perplexity, and Gemini pull from different sources and have different habits. Perplexity in particular is citation-heavy and rewards fresh, well-structured pages, so a brand can look strong there and weak in ChatGPT. Report per engine first, then blend.
Step 4: Normalize into share of voice
Once you have mention rates for every brand, you have two ways to read them.
Absolute visibility answers “how often am I named”: mentions divided by total runs. Useful for tracking your own progress.
Share of voice answers “when the model names anyone in this category, how much of that attention is mine”: your mentions divided by the sum of mentions across all tracked brands. This is the number that makes a benchmark comparable, because it removes the effect of prompts where nobody gets named at all. The full method is laid out in how to measure AI share of voice, and AI share of voice versus market share covers why the two numbers often disagree.
A minimal report per engine looks like this:
| Brand | Mention rate | Share of voice | Avg. position | Net sentiment |
|---|---|---|---|---|
| You | 22% | 14% | 3.4 | neutral |
| Competitor A | 41% | 27% | 1.9 | positive |
| Competitor B | 35% | 23% | 2.6 | positive |
| Category leader | 55% | 36% | 1.3 | positive |
Two things jump out from a table like this even before you dig in: the leader is winning on position as much as on presence, and Competitor A is getting a warmer description than you. Those are different problems with different fixes.
Step 5: Break the gap down by prompt, not by brand
A headline gap (“we are at 14%, they are at 27%”) is a motivational number, not an actionable one. The action lives one level down. Sort your prompt set by the difference in mention rate between you and your closest competitor, and you will usually see three buckets:
- Prompts nobody owns. Mention rates are low for everyone. These are open ground and usually the cheapest wins: a clear “what is X” page, a specific comparison, a well-structured use-case guide.
- Prompts a competitor owns and you do not. Look at the citations. Nine times out of ten the model is leaning on a specific review site, a comparison page, or a G2 or Capterra listing where the competitor is present and you are thin or absent. Citation data is where the benchmark turns into a to-do list, which is why citation tracking matters more than mention rate alone.
- Prompts the leader owns outright. High mention rate, first position, warm sentiment, everywhere. Note them, do not chase them this quarter.
For each prompt in bucket two, write down the top three cited domains for the competitor. If the same third-party site keeps appearing across many prompts, that one site is worth more than a month of blog posts.
Step 6: Turn it into a recurring number, not a one-off audit
The first benchmark tells you where you stand. The value comes from the second, third, and tenth. Visibility in AI answers moves faster than Google rankings: a model update, a new review roundup, or a competitor’s PR push can shift the picture in a couple of weeks.
Practical cadence for a small team:
- Weekly: automated runs, same prompt set, same engines. Glance at the trend, ignore single-week wobbles.
- Monthly: the full comparison table plus the per-prompt breakdown. Update the to-do list from the citation gaps.
- Quarterly: revisit the competitor set and the prompt set. Drop prompts that no longer describe how buyers ask, add new ones from support tickets and sales calls.
Running this by hand across four engines with multi-sampling is hundreds of queries a week, which is why most founders either stop after the first audit or automate it. AskAiRank was built for exactly this loop: you set the prompts and competitors once, it polls ChatGPT, Claude, Perplexity, and Gemini on a schedule, and the competitor comparison and share-of-voice table update themselves. Whichever tool you use, the requirement is the same: identical prompts, repeated sampling, per-engine reporting, and a history you can look back on.
Common ways benchmarks go wrong
- Different prompts per brand. Scores stop being comparable the moment the question changes.
- One run per prompt. You are measuring randomness, not visibility.
- Blending engines too early. A 30% blended score can hide 55% in Perplexity and 5% in ChatGPT, which are two very different situations.
- Counting brand-direct prompts as wins. “Is Brand A good” naming Brand A is not visibility, it is the model reading the question back to you. Keep those prompts for sentiment, not for share of voice.
- Ignoring position and sentiment. Being the fifth, hedged name in a list is not the same as being the confident first recommendation.
Your next step
Pick three competitors, write 30 prompts across the five types above, and run them at least three times each on two engines this week. Put the results in the five-column table, sort your prompts by gap, and pick the two or three prompts in the “nobody owns it” and “cited site we are missing from” buckets. That is a benchmark you can act on and, more importantly, one you can rerun next month to see if the gap moved.
Frequently asked questions
Answers about running a fair, repeatable AI visibility benchmark against rivals.
Three to five is the practical range for a small team. Fewer than three and you cannot tell whether a gap is about you or about the whole category. More than five and the prompt count needed for stable numbers grows faster than the insight you get back. Start with the two or three names buyers actually mention next to yours, then add one aspirational leader.
Both, but for different reasons. Similar-sized products tell you whether you are winning your actual battles. The market leader sets the ceiling and shows which prompts are 'owned' and probably not worth chasing this quarter. Keep them in separate rows of your report so one does not hide the other.
Weekly for the trend line, monthly for the full write-up. LLM answers move faster than Google rankings, so a single snapshot goes stale within weeks. What you want is the direction of the gap over four to eight weeks, not a one-time score.
Often, yes, at the prompt level. LLMs weigh entity clarity, third-party mentions, and comparison content heavily, and a small product can win specific use-case prompts that a bigger competitor never targets. Focus on the prompts where the gap is narrow or where nobody is dominant, not on the ones the leader has owned for a year.
It works as a first look, but not as a benchmark. Answers vary between runs, between models, and between days, so a handful of manual queries can flip the result on you. To compare brands fairly you need the same prompts, run several times, across several engines, on a schedule. That is the whole reason tracking tools exist.