Methodology
How AI visibility tracking actually works
Mentionify measures what AI assistants really tell your buyers — by asking them the way a buyer would, and scoring every answer with one calibrated judge.
- 3 assistants asked
- 1 judge, 1 rubric
- Any schedule you choose
From a buyer’s question
to a number you can move
Four steps, run end to end on every topic you track. Nothing in between is a shortcut.
- 01
You write the questions
Topics in Mentionify are real buyer questions — “what's the best CRM for a small agency?” — not keyword strings. Assistants answer questions, so questions are what we ask.
- 02
We ask the real assistants
Each question goes to ChatGPT, Gemini, and Claude the way a consumer asks it: no system prompt, no output caps, web search available but never forced. The model decides whether to look things up.
- 03
One judge reads every answer
Whoever wrote the answer, the same judge model scores it against the same five-dimension rubric — the only way a 70 on Gemini and a 70 on ChatGPT can mean the same thing.
- 04
The numbers land on a schedule
Runs fire on the cadence your plan buys — daily on every paid tier — and every tile, trend, and leaderboard aggregates the same settled runs, so no two numbers can disagree.
We ask the model your customer actually talks to
Every provider’s API default is a developer baseline, not what a free consumer gets. Leaving it unset is not neutrality — it silently pins the wrong rung.
| What is set | Typical API default | Mentionify |
|---|---|---|
| Which model | Whatever the API defaults to — often a developer-tier or preview model | The model a free-plan user is actually served in the consumer app |
| Thinking / reasoning effort | The API's default effort, one or two rungs above the free chat | Pinned to the consumer rung, because effort changes behaviour — not just cost |
| System prompt | A steering prompt (“answer as an analyst…”), which changes who gets named | None. Your question, exactly as written |
| Web search | Forced on (or off) for consistency | Available, never forced — the model decides, and we record what it chose |
| Output limits | Token caps that truncate long answers | No caps. A truncated answer is a different answer |
The rung changes behaviour, not just cost. When we pinned two providers to their consumer tier, their web-search counts halved — and a grounded answer’s cost is 72–80% web search, so anyone measuring tokens alone never sees it. A model that searches less answers more from memory, and that is a different answer about your brand.
The model scores components.
We do the arithmetic.
Five dimensions, published with their weights. Your brand and every competitor are scored on exactly this — the same rubric, the same caps, the same maths.
Coverage
25%How much of the answer is about you — a bare name-drop, one clause, a paragraph, or the whole subject.
Placement
25%Where you land: the first sentence, the middle, or a footnote nobody reads.
Prominence
20%How you are presented — the heading and top pick, a top-three list item, or a parenthetical aside.
Frequency
15%How many times the answer comes back to you. Once is not the same as six times.
Framing
15%How you are described: the direct answer, an active recommendation, a balanced peer, or a contrast case.
AI Visibility Score=presence×prominence
One division, rounded once — never a rate multiplied by an average. Competitors use the same denominator, so the leaderboard compares like with like.
Why no model keeps its own headline. Measured across 47 stored answers, a judge asked for a single overall number scored itself about 6 points higher than its own rubric’s blend — up to 18 points on one answer. Whoever keeps a one-shot number is systematically flattered, so nobody keeps one here: every headline is recomputed from the five components.
A competitor is verified, not guessed
A company is its registrable domain — and a domain has to be sourced before it is allowed to decide anything.
- 01
Propose
A grounded search proposes the company's website. Nothing is decided here — a proposal is a hypothesis, and models invent plausible domains constantly.
- 02
Fetch
We fetch the proposed site ourselves. This is the one step in the whole pipeline that asks no model anything.
- 03
Require proof
The page has to actually name the company — transliteration included — or the binding is refused. Unreachable settles nothing: it is retried next run, never remembered as “no website”.
Models invent websites — we caught it. Asked which site belonged to a company it had just named, the judge produced two domains that do not resolve at all. An unverified domain can never key a leaderboard row or merge two companies here; it is kept for the logo and decides nothing.
And a rival has to sit at your layer
An answer to “where can I buy X in my city?” necessarily names both the shops and the brands they stock. Left alone, a judge will hand a retailer the products on its own shelves as competitors — we watched it list, as a rival to one shop, a brand that shop sells fifteen products of.
So every answer commits to your market layer before it names anyone, and each brand it mentions is labelled by its relation to you: competitor, merchandise, channel, or supplier. Only true competitors reach your leaderboard — and nothing is lost, because every other named brand still appears in the answer’s entity list.
What lands on your dashboard
Every metric below is aggregated over all settled runs in the period you select — one pipeline, so no two tiles can disagree.
Presence rate
How often you appear at all, per model and overall — the base of every other number.
AI Visibility Score
Presence × prominence in one figure, 0–100, computed on our servers from the rubric above.
Competitor leaderboard
Who else gets named, ranked on the same denominator as you — so the comparison is like for like.
Cited sources
The pages the models actually read before answering. That list is your AEO to-do list.
Per-model breakdown
ChatGPT, Gemini, and Claude side by side. They disagree more than you would expect.
Trends and deltas
Period-over-period movement on every metric, with markers where your topic set or plan changed.
Want to see it on your own brand first? Run one free check — no account, one question, two models, results in about a minute.
Methodology questions
How do I track my brand's presence in AI search?
Ask the assistants the questions your buyers ask, on a schedule, and analyse the answers. That is what Mentionify automates: you add your topics as real questions, we ask ChatGPT, Gemini, and Claude daily, and every answer is scored for whether you are named, how prominently, and who else appears. There is no console inside ChatGPT to read — the answers themselves are the data.
Which AI models does Mentionify ask?
Gemini, ChatGPT and Claude — all three on every paid plan; each project picks its own mix, and a check spends 1, 2 or 3 credits by model. Each is asked at its consumer tier with web search available but never forced, so the answer resembles what a real person would have seen.
How is the AI Visibility Score calculated?
It is presence × prominence: how often you appear across the period's answers, weighted by how visible you are when you do. The judge scores five dimensions per answer — coverage 25%, placement 25%, prominence 20%, frequency 15%, framing 15% — and our server does the arithmetic and applies the calibration caps. No model ever hands us a headline number we keep.
Why do you use one judge model for every provider's answers?
Because comparability is the point. If each model graded its own answers, or each provider were scored on a different rubric, the numbers could not honestly be placed on the same chart. One judge, one rubric, one calibration — applied to your brand and to every competitor symmetrically.
How often do checks run?
On the cadence your plan buys — daily on every paid tier, monthly on the free Snapshot plan. Nobody presses a Run button: the schedule fires, results settle, and the dashboard updates. That is deliberate, because a number you check only when you remember to is not a trend.
Do the answers you record match what a real user sees?
As closely as an API allows. Same consumer-tier model rung, same raw question, no system prompt, no output caps, and web search left to the model's own judgement. What cannot be replicated is personalisation — chat history, account context, and locale shift answers in ways no API exposes — so we measure the unpersonalised baseline consistently rather than pretending to be one specific user.
Measure what AI says
about you.
Start free, or track your whole question set daily across every assistant.




