On this page
Your product either shows up when someone asks ChatGPT for “the best tool for X”, or it does not. But “shows up” is slippery: the same question asked twice can produce two different answers, and you sell across dozens of buyer questions, not one. An AI visibility score exists to turn that mess into a single number you can track week over week. This post explains what the number actually measures, how it is calculated, and where it can mislead you.
What an AI visibility score actually measures
A visibility score is the percentage of AI answer opportunities in which your brand appears. In its simplest form:
Visibility score = (answers that mention your brand / total answers checked) x 100
The important word is “opportunities”. You define a set of prompts your buyers realistically ask, run each prompt through one or more AI assistants, and check every response for a mention of your brand. If you run 50 prompts through 4 providers and get mentioned in 38 of the 200 answers, your raw visibility is 19%.
That is the whole idea. Everything else in this post is about the details that make the number honest instead of misleading: which prompts, which providers, how many repetitions, and how mentions are detected and weighted.
The four inputs that define the score
Two tools can both report “visibility score” and produce completely different numbers for the same brand. The difference always comes down to four inputs.
1. The prompt set. The score only reflects the questions you chose to measure. A prompt set full of “what is [your product]” questions will flatter you; a set of generic “best [category] tools” questions will humble you. A good set mixes branded, category, comparison, and problem-based prompts. Getting this right matters more than any formula detail, and there is a real tradeoff between coverage and cost, which the step-by-step guide to measuring AI share of voice walks through.
2. The providers. ChatGPT, Claude, Perplexity, and Gemini have different training data, different retrieval behavior, and different opinions about your category. It is common to have 40% visibility on Perplexity, which retrieves live web pages, and near zero on a model that relies more on training data. A score built on one provider hides this. A cross-provider score averages it, which is fine as long as you can also see the per-provider breakdown.
3. The number of samples. LLM answers are non-deterministic. Ask the same model the same question five times and your brand might appear in three of the five responses. A single run per prompt is a coin flip dressed up as a metric. Serious scores use multi-sampling: each prompt runs several times per provider, and the mention rate across samples becomes the data point.
4. The time window. Daily numbers are noisy even with sampling. Most dashboards compute the score over a rolling window, typically 7 or 30 days, so that one odd polling run does not swing the trend line.
From raw mention rate to a weighted score
The basic mention percentage treats every appearance equally. In practice, being the first recommendation in a list of eight is not the same as being a grudging afterthought in sentence twelve. Most scoring systems layer weights on top of the raw rate:
| Factor | What it captures | Typical effect on the score |
|---|---|---|
| Mention presence | Were you in the answer at all | The base of the score |
| Position | Rank within a list of recommendations | First mention worth more than fifth |
| Sentiment | Recommended, neutral, or criticized | Negative mentions discounted or flagged |
| Provider weight | Where your buyers actually ask | Optionally weight ChatGPT higher than others |
A weighted score might look like this: a first-position positive mention counts as 1.0, a mid-list mention as 0.6, a passing neutral reference as 0.3, and a negative mention as 0 (but surfaced as an alert, because “avoid [your product], it lacks X” is information you want immediately).
Whether weighting is worth it depends on your stage. Early on, presence is the battle: you mostly want the binary “am I in the answer” rate going up. Once you appear regularly, position and sentiment become the levers, because moving from “also worth considering” to “the top pick” changes buyer behavior far more than getting one extra mention somewhere.
Detection itself is less trivial than it sounds. Answers refer to products with abbreviations, misspellings, and partial names, so naive string matching both misses real mentions and false-positives on similarly named tools. Tools like AskAiRank combine pattern matching with LLM-based entity extraction to catch fuzzy variants; if you build your own checker, budget time for this or your score will drift for reasons that have nothing to do with your actual visibility.
Absolute score vs share of voice
Your visibility score in isolation answers “how often do I appear”. It does not answer “am I winning”. For that you need the competitive frame: on the same prompt set, how often do your competitors appear, and what share of all brand mentions belongs to you?
That second number is your AI share of voice, and it is often more actionable than the absolute score. A visibility score of 25% sounds weak until you learn the category leader is at 31% and everyone else is under 10%. It sounds very different if the leader is at 80%. The mechanics of computing it are straightforward once you already parse mentions, and we walk through them in how to measure AI share of voice.
Competitive framing also protects you from a subtle failure mode: category-wide shifts. If a model update makes an assistant recommend fewer tools per answer, everyone’s absolute score drops and yours says nothing about your performance. Share of voice stays interpretable, which is why tracking competitors in AI answers should be part of the setup from day one, not an upgrade for later.
How to read the number without fooling yourself
A few rules that keep the score useful:
- Trend beats level. The absolute number depends on your prompt set difficulty. The direction over 30-90 days is the signal.
- Never compare scores across tools. Different prompt sets, providers, and weights make cross-tool comparison meaningless. Compare within one measurement system.
- Freeze the prompt set before comparing periods. If you add 20 hard prompts mid-month, your score will drop for reasons unrelated to reality. When you change the set, treat it as a new baseline.
- Look at the per-provider split before reacting. A 5-point overall drop caused entirely by one provider’s model update is a different situation than erosion across all four.
- Pair the score with citation data. Visibility tells you that you appeared; citations tell you which pages earned it, which is what you can actually act on.
Calculate a rough score yourself this week
You do not need any tooling to get a first baseline:
- Write down 20 prompts your buyers would ask: category questions, “best X for Y” questions, comparisons against your top competitor, and problem-phrased questions.
- Run each prompt in ChatGPT and Perplexity, twice each, in fresh conversations. That is 80 answers.
- For each answer, record: mentioned yes/no, position if it is a list, and which competitors appeared.
- Your score is mentions divided by 80. Your share of voice is your mentions divided by all brand mentions recorded.
This takes an afternoon and gives you a real baseline plus an honest feel for why sampling matters, because you will see the same prompt produce different answers back to back. The manual version stops scaling almost immediately though: 20 prompts across 4 providers with proper sampling on a schedule is thousands of runs per month. That is the point where automated tracking earns its keep, and where AskAiRank does exactly this loop on a schedule: polling your prompts across ChatGPT, Claude, Perplexity, and Gemini, parsing mentions, and turning the results into a score and trend you can check in five minutes a week.
Either way, start with the manual baseline. A founder who has hand-scored 80 answers reads every dashboard, theirs or anyone else’s, with much better judgment.
Frequently asked questions
Answers about how the AI visibility score is calculated and what moves it.
Not directly. Each tool uses its own prompt set, provider mix, sample counts, and weighting, so a 40 in one tool and a 40 in another can mean very different things. Compare your own score over time within one tool, and compare against competitors measured by that same tool, but never compare raw numbers across products.
LLM answers are non-deterministic and the models themselves get updated. Providers ship new model versions, retrieval sources shift, and competitors publish content that displaces you. A score built on repeated sampling smooths the random part, but real drift from model updates and competitor activity is genuine signal, not noise.
There is no universal benchmark because the score depends on how competitive your prompt set is. A more useful frame: on prompts where you are a legitimate answer, you want to be mentioned in a meaningful share of runs, and you want your trend line moving up relative to the 2-3 competitors you track. Position relative to alternatives matters more than the absolute number.
As a rough floor, 20-50 prompts sampled repeatedly across multiple providers gives you a usable trend line. A single run of 5 prompts is a spot check, not a metric. The math is simple: more prompts reduce topic bias, more samples per prompt reduce randomness, and more days of data reduce day-to-day noise.
It is a leading indicator, not a revenue metric. A rising score means more buyers see your name at the moment they ask for recommendations, which tends to show up later as direct traffic and 'heard about you from ChatGPT' signups. Treat it like rankings in SEO: necessary context, but you still confirm impact through signup sources and revenue.