You can check this yourself in about a minute: open ChatGPT, ask the question your buyers ask, and read whether your brand is in the answer. That part is genuinely easy. The hard part is that a single check, done casually, will usually tell you something slightly untrue — and then you make decisions on it.
This is the method we use, the traps that make casual checks misleading, and the point at which checking by hand stops being worth your time.
The short version. Ask the question a buyer would ask, in a logged-out or temporary chat, and record four things: whether you were named, where in the answer, who else was named, and whether the model searched the web. Repeat across ChatGPT, Gemini, and Claude, and repeat over days — one answer is an anecdote, and the assistants disagree with each other more than you expect.
The manual check, step by step

1. Write the question your buyer asks, not your brand name. "Is [your brand] any good?" tells you nothing useful — you have handed the model the answer. The question that matters is the one asked before anyone knows you exist: "what's the best help-desk tool for a 20-person SaaS company?", "where can I buy climbing shoes in Berlin?", "who does technical SEO for e-commerce in the Nordics?"
2. Use a logged-out or temporary chat. Your own account is the worst possible place to test this. Chat history, memory, and past conversations about your own company all bias the answer toward you. Use a temporary chat, a logged-out session, or a different browser profile.
3. Ask it verbatim, and don't coach. No "answer as an expert analyst", no "list 10 companies including any Georgian ones". Every instruction you add moves the answer away from what a real person would have seen.
4. Read the answer for four things, not one:
- Are you named at all?
- Where — first sentence, middle of a list, or a parenthetical at the end?
- Who else is named, and in what order?
- Did the model search the web (citations, links, "according to…") or answer from memory?
5. Repeat on Gemini and Claude. They will not agree. That disagreement is data, not noise.
6. Write it down. A screenshot in a folder is not a measurement. One row per question × model × day, with those four fields, is.
Five traps that make manual checks misleading
| Trap | What it does to your result | The fix |
|---|---|---|
| Logged-in session | Memory and history inflate your own presence | Temporary chat, logged out, clean profile |
| One ask | Answers vary run to run — you sampled one roll | Ask repeatedly, count the rate |
| Leading question | You wrote the brand into the prompt | Ask the pre-purchase question instead |
| Ignoring web search | A grounded answer and a memory answer have different causes | Record whether it searched |
| Reading for your name | You skim past the competitor list — the most useful part | Record who else was named, every time |
The second one is the expensive trap. Visibility is a rate, not a fact. The same question, asked twice, can name different companies — retrieval differs, the model may or may not search, and generation is not deterministic. A single "yes, I'm in there" is a coin that came up heads once.

What to record every time
If you are going to do this by hand, use these columns. They are the same fields an automated tracker stores, and they are the minimum that makes two checks comparable:
| Field | Example | Why it matters |
|---|---|---|
| Question | "best CRM for a small design studio" | The unit of tracking — always verbatim |
| Model | ChatGPT / Gemini / Claude | They disagree; an average across them hides it |
| Date | 2026-07-20 | Turns checks into a trend |
| Mentioned | yes / no | The base metric: presence |
| Position | 1st brand named / 4th / footnote | Being named last is not being named first |
| Competitors named | acme.io, rival.ai | The most actionable column on the sheet |
| Searched the web | yes / no | Tells you which lever applies |
| Sources cited | review-site.com, forum.dev | The pages that decided the answer |
That last column is the one people underuse. When a model searches before answering, the pages it cites are the pages that decided whether you were named. That list is an AEO to-do list written by the model itself.
The three assistants behave differently
| ChatGPT | Gemini | Claude | |
|---|---|---|---|
| Searches the web | Often, question-dependent | Often, question-dependent | Often, question-dependent |
| Names brands readily | Yes, usually a short list | Yes, frequently with links | More cautious, more caveats |
| Typical answer shape | Ranked shortlist | Shortlist plus sources | Prose with conditions |
| What moves it | Retrieved pages + trained knowledge | Retrieved pages + trained knowledge | Retrieved pages + trained knowledge |
The practical consequence: checking one assistant tells you about one assistant. If your buyers are split across ChatGPT and Gemini, a strong result on one and absence on the other is a very different situation from being solidly present on both — and you cannot tell which you are in without asking both.
When to stop doing this by hand
Manual checking is fine for one question, once, to satisfy curiosity. It stops being viable quickly, because the workload is multiplicative: questions × models × days.
| Manual check | Automated tracking | |
|---|---|---|
| Effort | ~2 minutes per question per model | Set up once |
| Coverage | 1–2 questions, occasionally | Every question, every model, every day |
| Bias control | Depends on your discipline | Enforced by the method |
| Position and prominence | Eyeballed | Scored on a fixed rubric |
| Competitor tracking | Manual notes | Leaderboard, ranked on the same denominator |
| Trend | Only if you keep a spreadsheet | Built in, with period-over-period deltas |
Ten questions across three assistants is thirty checks a day — about an hour of copy-paste, every day, before anyone has analysed anything. That is the point where a tool pays for itself.
If you want the automated version of exactly this method — the same questions, asked at the consumer tier with no system prompt, every answer scored on one rubric so the models are comparable — that is what Mentionify does. The methodology page documents every step, including the scoring weights. Or start with one free check: the AI visibility checker asks Gemini and ChatGPT one question of your choice and shows you both answers, both verdicts, and every source they used.
FAQ
How do I track my brand's presence in AI search?
Ask the assistants the questions your buyers ask, on a schedule, and record whether you were named, where, who else was, and whether the model searched. There is no analytics console inside ChatGPT — the answers themselves are the only data source, so tracking means generating them repeatedly and analysing them consistently.
Can I just ask ChatGPT "do you know my brand?"
You can, and the answer is close to meaningless. It tests recognition of a name you supplied, not whether you are recommended when someone asks about your category. Always ask the pre-purchase question instead.
Why do I get a different answer every time?
Because generation is not deterministic and retrieval varies. The model may search the web on one run and answer from memory on the next, and the pages it retrieves change. This is why presence is measured as a rate over many asks rather than a yes/no from one.
Does asking in a logged-out session really change the result?
It removes the most obvious source of bias — account memory and chat history about your own company. It does not make the answer universal: locale, language, and the model's own routing all still apply. The goal is a consistent, unpersonalised baseline you can compare over time, not a simulation of one specific customer.
How often should I check?
Daily for the questions that matter commercially, weekly at minimum. Retrieval-driven answers can change within days of new content being indexed, so a monthly check will show you that something moved without showing you what moved it.



