From 28 July to 3 August 2026 we asked three AI assistants the same shopping question once every day: what is the best place to buy everyday products online with fast delivery in the United States? Three assistants, seven days, twenty-one answers.
Amazon was named in all twenty-one. Perfect presence, first position in most answers, the kind of result a brand team would put on a slide and stop reading.
Then we opened the per-model view. Across Gemini's seven answers, Walmart's visibility score came out at 72 and Amazon's at 70. On the biggest shopping question in the biggest retail market in the world, one of the three assistants was quietly answering Walmart.
What we ran
| Brand tracked | Amazon (amazon.com) |
| Question | "What is the best place to buy everyday products online with fast delivery in the United States?" |
| Assistants | Gemini, ChatGPT, Claude |
| Frequency | Once a day, seven days running |
| Window | 28 July – 3 August 2026 |
| Answers collected | 21 (7 runs × 3 assistants) |
| Failed answers | 0 |
Every assistant was asked the way a customer asks: the raw question, no system prompt, no instruction to search, no output limit, and each one pinned to the tier a person on a free plan actually gets rather than the developer defaults. Then every answer, from every assistant, was scored by the same judge, so a Gemini number and a ChatGPT number mean the same thing.
This was one of five brands we set up that week. The ecommerce one behaved nothing like this one: Shopify was first in every single answer and still held under a quarter of the conversation.
The headline numbers, and why they hide the problem
| Metric | Value |
|---|---|
| Mention visibility | 100% (21 of 21 answers) |
| AI Visibility Score | 79 |
| Average position | #1.2, in 21 of 21 answers |
| Share of voice | 23% — first of 13 brands named |
| Mention quality | 79 |
| Web search rate | 90% |
Presence of 100% is real and it is worth having. It also compresses seven days of disagreement into a single flat line. Amazon is named in every answer because Amazon is Amazon; the question is what the answer does with the name once it is there, and that only shows up when you stop blending the assistants together.
The three assistants did not agree
| Gemini | ChatGPT | Claude | |
|---|---|---|---|
| Amazon's AI Visibility Score | 70 | 85 | 82 |
| Amazon's average position | #1.7 | #1.0 | #1.0 |
| Walmart's score | 72 | 54 | 64 |
| Target's score | 58 | 46 | 42 |
| Ran a web search | 7 of 7 | 7 of 7 | 5 of 7 |
Fifteen points separate the best assistant from the worst, for one brand, on one question, in one week. Nothing about Amazon changed between those columns. What changed is which assistant was asked.

Here is Gemini's week, day by day:
| Day | Amazon's position in Gemini's answer | Ranked above it |
|---|---|---|
| 28 Jul | #3 | Walmart, Target |
| 29 Jul | #2 | Walmart |
| 30 Jul | #2 | Walmart+ |
| 31 Jul | #1 | — |
| 1 Aug | #2 | Walmart |
| 2 Aug | #1 | — |
| 3 Aug | #1 | — |
ChatGPT and Claude put Amazon first on all seven days without exception. Gemini put something else first on four of them. And it was not subtle about it — this is Gemini's own answer from 1 August, under its own heading numbered 1:
Walmart is the biggest competitor to Amazon for everyday essentials, and it is often considered the best option for grocery-focused shopping.
Amazon appears immediately below, at number 2, described as "the gold standard for sheer product variety". A perfectly complimentary paragraph, in second place.

Why one assistant answered differently
Before an assistant answers a question like this, it rewrites the question into its own web searches. We capture every one of them, because they are the clearest signal we have found for why an answer came out the way it did. Across the week the three assistants issued 39 searches, 34 of them distinct — and they were not looking for the same thing at all.
Gemini went shopping for competitors' prices. Twelve of its eighteen distinct searches were about a rival's membership; three were about Amazon's:
- Walmart plus membership price 2026
- Target Circle 360 subscription price 2025 2026
- Instacart subscription cost perks 2026
- Kroger Boost delivery same day 2026
ChatGPT went to check Amazon's homework. Nine of its ten searches named Amazon Prime directly, and one of them was a site-scoped lookup of Amazon's own membership page:
- Amazon Prime delivery benefits United States free same-day one-day delivery 2026
- Amazon Prime membership cost United States 2026 official Amazon
- site:amazon.com/amazonprime Prime membership $139 per year $14.99 per month United States
Claude mostly re-asked the question. Six distinct searches, all near-restatements of what it was given, with a year bolted on: "best places to buy everyday products online fast delivery United States 2024".
That is the whole mechanism, visible in one screen. The assistant that spent its research budget pricing Walmart+ and Target Circle 360 came back and ranked Walmart first. The assistant that spent its research budget confirming Amazon Prime's own claims came back and ranked Amazon first, with the highest score of the three. Neither was being unfair. They read different pages.

Whose pages the assistants actually read
The 21 answers cite 177 pages across 70 distinct domains. Sorted by how often each domain was read:
| Domain | Citations | Share |
|---|---|---|
| walmart.com | 15 | 8.5% |
| target.com | 11 | 6.2% |
| aboutamazon.com | 9 | 5.1% |
| ebay.com | 9 | 5.1% |
| nbcnews.com | 8 | 4.5% |
| thekitchn.com | 6 | 3.4% |
| mixandmatchmama.com | 6 | 3.4% |
| getcartswap.com | 5 | 2.8% |
| amazon.com | 4 | 2.3% |
Add up each company's own properties and the picture is blunter. Amazon's — amazon.com, its newsroom at aboutamazon.com, its payments subdomain — come to 14 of the 177 citations, or 7.9%. Walmart's and Target's own domains, counting their corporate newsrooms, come to 31. In the answer to "where should I buy everyday products", the assistants read Walmart's and Target's pages about themselves more than twice as often as they read Amazon's pages about Amazon.
Below the top few sits a long, unglamorous tail: a couponing blog, a family blog, a personal finance site, a fulfillment company's marketing blog. Those pages are not prestigious and it does not matter. They were in the retrieval set, so they were in the answer.

The leaderboard, and the brands that were secretly one brand
Twelve other brands were named alongside Amazon over the week, thirteen rows counting Amazon itself. The board, scored on the same scale as the brand:
| Brand | Score | Named in |
|---|---|---|
| Amazon | 79 | 21 of 21 |
| Walmart | 63 | 21 of 21 |
| Target | 49 | 20 of 21 |
| Instacart | 34 | 16 of 21 |
| Costco | 8 | 5 of 21 |
| Best Buy | 3 | 2 of 21 |
| Gopuff | 3 | 1 of 21 |
Walmart was named in every single answer too. So the real competitive picture is not "Amazon wins" — it is that this question has a three-name shortlist, Amazon, Walmart and Target, and everyone else is fighting for the scraps at the bottom of the paragraph.
One detail worth pausing on. Gemini rarely wrote "Walmart". It wrote "Walmart+", and it wrote "Target Circle 360" instead of "Target". Counted naively, that is four brands with four small scores instead of two brands with two big ones — and Walmart would never have surfaced as the threat it is. Our resolver checks each name against a real, reachable website before merging, so Walmart+ folds into walmart.com and Target Circle 360 folds into target.com. The difference between the naive count and the resolved one is the difference between a leaderboard that is decorative and one you can act on.
Where Amazon is actually weak
Every mention gets scored on five separate axes rather than one vague number, and the weak axis is never the one a brand expects:
| Axis | Amazon's average |
|---|---|
| Framing (how favorably it is described) | 92 |
| Placement (where in the answer it sits) | 91 |
| Prominence (how much of the answer is about it) | 90 |
| Frequency (how often it recurs within the answer) | 67 |
| Coverage (share of the answer's real estate) | 56 |
Framing at 92 means the assistants like Amazon. Coverage at 56 — the lowest band on the table, and it ranged from 35 to 75 across the week — means the answer is not really about Amazon: it is a comparison, and Amazon is one part of it. That is what a 23% share of voice looks like from inside a single answer, and it is the axis a competitor can attack most directly, by getting named in more of the pages that build the comparison.
Sentiment averaged 79 out of 100 across the 21 answers, ranging from 65 to 85, and the same two complaints recur in the answers at the bottom of that range: the membership is the most expensive of the three, and grocery delivery can carry extra fees. The tool writes those out as plain suggestions next to the numbers — one of them, verbatim from the run of 28 July, reads: "Highlight competitive grocery delivery fees to counter the mention of 'extra fees' for Amazon Fresh compared to Walmart."
The two days nobody searched
Claude answered without running a web search on 30 and 31 July. No searches, no citations, no retrieval — it answered from what it already believed, and it still put Amazon first.

That is the whole reason web-search rate is a metric here and not a footnote. When an assistant searches, this week's pages decide the answer and a PR win can move it inside a fortnight. When it does not, you are being judged on years of accumulated public record and nothing you publish this month will touch it — the two supply lines behind that split are worth understanding before you budget a quarter. Amazon can be relaxed about that distinction. A brand that is three years old cannot, and the only way to know which mode you are in is to measure it.
If you are not Amazon
Amazon has the strongest possible position in this dataset and still loses one assistant out of three. Every finding here scales down badly for smaller brands, which is exactly why they are worth copying:
Check every assistant separately. A blended average would have shown Amazon at 79 and hidden the fact that Gemini scored a competitor higher. Whatever your number is, it is an average of disagreements.
Read the searches, not just the answers. The fan-out queries told us why Gemini ranked Walmart first long before any amount of staring at the answers would have.
Your competitors' websites are in your answers. Walmart's own pages outnumbered Amazon's own pages by more than two to one, in Amazon's category. You do not control the retrieval set — you compete inside it.
Once is not a measurement. Gemini gave Amazon third place, second place, second place, first, second, first, first. Any single day of that week supports a different conclusion. Seven days supports one.
None of that tells you what to do, which is deliberate — this is the measurement half. Bondo has written the other half as a prioritized, superstition-free to-do list.
What this ran on
Everything above came out of one project in Mentionify, over one week, with nobody watching it. The method is the same for any brand; these are the surfaces the study's data lives on:
- Topics — the buyer question being tracked, with a per-topic score and its own detail page.
- Models — Gemini, ChatGPT and Claude on the same question, each with its own tab, so disagreement is visible instead of averaged away.
- Daily runs — seven runs, one a day, each a full pass of the topic across the three assistants at a fixed, published credit price. (Mentionify has a built-in daily schedule; for this study the runs were fired by our own automation at the same hour each morning.)
- AI Visibility Score, mention visibility, average position, share of voice, mention quality, web-search rate — the six headline metrics, each with its period-over-period movement.
- Trend over time with a period picker (7 / 30 / 90 days, a calendar month, or a custom range).
- Competitors — the leaderboard above, scored on the brand's own scale, with the alias resolution that turns Walmart+ back into Walmart.
- Sources — every cited domain, its exact URLs, citation counts and share of the total.
- AI search queries — the fan-out searches, attributed to the assistant that issued each one.
- Runs and archive — every settled run kept, and every raw answer readable in full.
- Per-answer recommendations — the plain-language suggestions the judge writes beside each scored answer, like the grocery-fees line quoted above.
- Report and exports — a printable report and CSV/XLSX export of any of these tables, scoped to the same period and the same model as the screen you are looking at.
If you want the short version for your own brand, run your buyer's question once, free — it is scored exactly the way these 21 answers were. What one check cannot show you is the part that mattered most here: movement, across days, across assistants. That needs a schedule.
FAQ
Does this mean Gemini is biased against Amazon?
No, and the data does not support that reading. Gemini ranked Amazon first on three of the seven days. What it did consistently was research competitors' membership pricing before answering, and on a question where price is a deciding factor, that research favored the cheaper membership. Different research, different answer.
Why score Gemini's answer with Gemini?
Every answer, from all three assistants, is scored by one judge model. That is the only way a ChatGPT number and a Claude number can sit in the same column. If each assistant graded its own homework, none of the comparisons in this post would mean anything.
Twenty-one answers is a small sample. Isn't that a problem?
For a statistical claim about Amazon's standing in the world, yes. For the actual question a brand has — is my visibility stable, and does it hold across the assistants my buyers use — it is enough to see the shape, and the shape here was unambiguous: two assistants steady at first place, one that ranked Amazon third on the very first run and never settled above second until day four.
Do you publish other studies like this?
Yes. Every brand we track this way is written up in full, with the same method and the same scoring, and they all sit together in the research index.
Can I see the raw answers?
Every answer is stored in full and readable in the run archive, which is how the quotes in this post were pulled. Nothing here is paraphrased from a summary.



