If you had to guess which brand an AI assistant names when somebody asks it for the best phone in America, you would guess Apple, and you would be right — every single time. Apple was in all 21 answers we collected. Not one of them left it out.
It came first in fourteen of them.
The other seven went to Samsung, and once you look at which assistant produced which answer, the split stops looking like noise. ChatGPT put Apple at the top of all seven of its answers and called it the top pick every time. Gemini and Claude, asked the identical question in the same ninety minutes, put Samsung first more often than not.

What we ran
| Brand tracked | Apple (apple.com) |
| Question | "What is the best phone out right now in the United States?" |
| Assistants | Gemini, ChatGPT, Claude |
| Runs | Seven, back to back, on 31 August 2026 |
| Answers collected | 21 (7 runs × 3 assistants) |
| Answers where the assistant searched the web | 21 of 21 |
| Failed answers | 0 |
The question was picked from search data, not invented: some version of "best phone" is typed hundreds of thousands of times a month in the United States, and this phrasing is how people actually ask it. It names no criterion. It does not say camera, or battery, or privacy. That matters, because the moment a question names a criterion it also names its winner, and what we wanted was a question the brands had to fight over.
We ran it seven times in one morning rather than once a day for a week. That was a deliberate change from a B2B question we tracked daily earlier in the summer, and it turned out to be the more revealing setup — anything that moved across these seven runs moved without the world changing underneath it.
The headline numbers
| Metric | Value |
|---|---|
| Mention visibility | 100% (21 of 21 answers) |
| AI Visibility Score | 76 |
| Average position | #1.5, across 21 of 21 answers |
| Mention quality | 76 |
| Sentiment | 78 |
| Share of voice | 30% — first of 7 brands |
| Web search rate | 100% |
Average position #1.5 is the number to sit with. It is not a rounding artefact of a couple of odd answers; it is what you get when a brand wins one assistant outright and loses ground on the other two. And the five components behind the visibility score point the same way: framing 89, placement 86, prominence 91 — but coverage 53. Apple is introduced early, described well and treated as important. It just does not take up much of the answer. The rest of the paragraph belongs to somebody else.
Share of voice at 30% says the same thing from the other end. Seven brands were named on a single question, and the largest consumer brand in the world holds under a third of the conversation.
The same question, three verdicts
This is where the study stops being about Apple and starts being about how AI answers work.
ChatGPT's board is a landslide. Apple 86, Samsung 56, Google 33 — and it opens its most recent answer with the sentence "Best overall phone in the U.S. right now: iPhone 17 Pro Max," in bold, no hedge, no preamble.

Gemini's board, from the same seven runs, has Samsung on top. Not by much — 71 to Apple's 70, with Google three points further back on 67 — but a photo finish is still a finish, and Apple is the runner-up in it.

Claude splits the difference and lands somewhere stranger. Apple and Samsung tie on 72, and the tie-break puts Samsung first. Claude also refuses the premise more often than the others — three of its seven answers file Apple as a strong alternative rather than the recommendation, and its latest answer begins "There isn't a single universally agreed 'best' phone," then puts the Galaxy S25 Ultra under the first heading.

Three assistants. One question. One brand that is either untouchable or narrowly second depending on which app your customer happened to open. If you have only ever checked one of them, you have seen one third of your position — and there is no way to tell from inside that one third whether you are looking at the 86 or the 70.
The same assistant changed its mind, too
Seven runs in ninety minutes should be seven copies of the same answer. They were not.

ChatGPT scored Apple 87, 88, 90, 81, 81, 88, 88 — a nine-point spread, essentially stable. Claude scored it 82, 87, 64, 63, 64, 80, 63. That is a twenty-four point swing on the same question inside an hour and a half, with nothing in the world changing in between. Gemini ran 56 to 81.
Nobody should read a single AI check as a fact about their brand. It is one draw from a distribution, and on two of these three assistants the distribution is wide enough that a single check could tell you almost anything you wanted to hear. The reason we run a question repeatedly is not thoroughness. It is that one run is not a measurement.
Who else was in the room
Merged across all three assistants, the board runs seven deep.

Samsung was named in all 21 answers, exactly as often as Apple. Google in 19. Then a cliff: OnePlus in four, Motorola in three, Asus and Xiaomi in one each. Two brands own this question and a third is always in the room; everyone else is a cameo. That shape is normal — every study we have published so far, and they are all collected in one place, has a short head and a long, thin tail.
One column is worth a second look. Tone — how warmly the answer talks about a brand — reads 81 for Samsung and 78 for Apple. Apple wins on position and loses, slightly, on affection. When these answers are enthusiastic, they are often being enthusiastic about the Galaxy.
We nearly published this section with Google missing from it. Our own scoring had a rule that excluded search engines and AI vendors from competitor boards, so that an assistant could not be scored as its maker's rival. For a marketing agency that rule is right. For a phone brand it hides Pixel, which is a real competitor sold in real shops, so we changed it before writing this up. The rule now catches assistants and search products only, never the companies behind them.
Nobody was reading Apple's website
The three assistants cited 221 pages across 66 different domains. Apple's own site accounted for seven of them.

Tom's Guide was cited 26 times, 11.8% of everything. TechRadar 21. YouTube 17 — a video platform is the third-biggest source of truth on this question. Wikipedia 13. Then the phone press: PhoneArena, PCMag, Tech Advisor, CNET, GSMArena. Apple's own pages come tenth, on 3.2%.
This is the part every brand underestimates. The company with the most valuable marketing department on earth has almost no voice in the answer written about it, because the assistant is not reading the brand — it is reading the people who write about the brand. What gets quoted back to a buyer is whatever the review sites said last week. If you want to understand where an assistant's facts actually come from, this table is the shortest version of the answer.
What they typed into the search box
Every one of the 21 answers was written after a live web search — 87 searches in total, 72 of them distinct phrases.

The three go about it very differently. Gemini fanned out hardest — 44 searches, 40 of them unique, including named-product checks like "Galaxy S26 Ultra reviews" and "Google Pixel 10 Pro reviews 2026 specs" that read like a shopper opening ten tabs. ChatGPT issued 22, all different. Claude issued 21 but reused the same handful, typing "best phone 2025 United States" ten separate times.
That difference explains a lot of the disagreement above. An assistant that searches for specific rival products will find pages that argue for those products. The fan-out is not a technical detail; it is the assistant deciding, before it writes a word, which shortlist it is about to read about.
Reading the answers themselves
Numbers are a summary. The answers are the evidence, and every one of them is stored and readable in full.

Put them side by side and the scores stop being abstract. ChatGPT: "Best overall phone in the U.S. right now: iPhone 17 Pro Max." Gemini: "The 'best' phone right now in the United States depends on whether you prefer the iOS or Android ecosystem." Claude: "There isn't a single universally agreed 'best' phone." An 88, a 69 and a 63, and you can hear the difference in the first eight words.
What the tool said to do about it
Every study we publish ends with the same question from readers: fine, but what would you actually change? The dashboard answers that itself, from the run data rather than from opinion.

Its top recommendation for Apple is not about apple.com. It is to get and hold position in the listicles that the assistants actually read, because two review sites supply nearly a quarter of every citation on this question and Apple's own domain supplies 3%. Its second is that Google is quietly winning the "AI features" framing inside these answers. Neither is a guess: both name the counts they came from, and you can click through to the answers that produced them. The broader version of that work is what moves a brand up a shortlist rather than onto one.
What to take from this
Being mentioned is not the finish line. Apple hit 100% presence and still lost the top spot in a third of the answers. Presence saturates early for any brand people have heard of. Position and share of voice are where the movement is.
Check all three, or you are guessing. An 86 on one assistant and a 70 on another is the same brand on the same day. A single check cannot tell you which one you are living in.
One run is not a measurement. Two of these three assistants moved more than twenty points across runs that were minutes apart.
The pages that decide your answer are not yours. For this question, four review sites outweighed every manufacturer's own website put together.
Your category may behave nothing like this one. An ecommerce brand that never lost a single answer sat at the top of all 21 of its answers with no argument from anybody, week after week.
The only way to find out which of those you are is to put your buyer's question to all three and read what comes back. Ask it once, free, against all three assistants — no account, a couple of minutes.
How this was run
One project, one tracked question, three assistants, seven runs. Each assistant got the raw question with no system prompt, no instruction to search and no output cap, set to the tier a person on a free plan actually gets. A single judge model then scored every answer, so the three columns can be read against each other rather than side by side. Component scores come from the judge; the totals are arithmetic done on our server, never a headline number the model was asked to produce.
Every figure on this page is recomputed by the same function that draws the dashboard, and every screenshot is the real interface, restyled but never retyped. The method is written up in full on how a single answer turns into a score.
One question, asked once, of one assistant is the unit we count and the unit we charge for, so the cost of a study like this is simply how many of those it took — 21. Running it is priced per prompt-check.
Questions people ask about this study
Why seven runs in one morning instead of one a week? Because we wanted to separate what the assistants disagree about from what changes over time. Run daily for a week and a swing could be the news cycle. Run seven times in ninety minutes and it can only be the model.
Does 100% mention visibility mean Apple is winning? It means Apple is never left out, which for a brand this size is the expected result rather than an achievement. The score that moved here was position, and it moved against Apple on two assistants out of three.
Why is Samsung ahead of Apple on two of the boards? Because the assistants put it there. Both were named in all 21 answers; the difference is in ranking and framing within each answer, and Gemini and Claude both lean toward the Galaxy as the headline pick more often than ChatGPT does.
Can I run this on my own brand? Yes, and it is the same pipeline — the free check runs one question against all three assistants and scores it with the same judge. Tracking a question over time, which is what produced everything on this page, is the paid version of the same thing.



