Case Study

Amazon vs Walmart in AI Answers: A 7-Day Study

We put one shopping question to ChatGPT, Gemini and Claude every day for a week. Amazon was named in all 21 answers — Gemini still scored Walmart higher.

Mentionify dashboard for amazon.com showing 100% mention visibility, an AI Visibility Score of 79, average position #1.2 across 21 of 21 answers, 23% share of voice and a 90% web-search rate

From 28 July to 3 August 2026 we asked three AI assistants the same shopping question once every day: what is the best place to buy everyday products online with fast delivery in the United States? Three assistants, seven days, twenty-one answers.

Amazon was named in all twenty-one. Perfect presence, first position in most answers, the kind of result a brand team would put on a slide and stop reading.

Then we opened the per-model view. Across Gemini's seven answers, Walmart's visibility score came out at 72 and Amazon's at 70. On the biggest shopping question in the biggest retail market in the world, one of the three assistants was quietly answering Walmart.

What we ran

Brand trackedAmazon (amazon.com)
Question"What is the best place to buy everyday products online with fast delivery in the United States?"
AssistantsGemini, ChatGPT, Claude
FrequencyOnce a day, seven days running
Window28 July – 3 August 2026
Answers collected21 (7 runs × 3 assistants)
Failed answers0

Every assistant was asked the way a customer asks: the raw question, no system prompt, no instruction to search, no output limit, and each one pinned to the tier a person on a free plan actually gets rather than the developer defaults. Then every answer, from every assistant, was scored by the same judge, so a Gemini number and a ChatGPT number mean the same thing.

This was one of five brands we set up that week. The ecommerce one behaved nothing like this one: Shopify was first in every single answer and still held under a quarter of the conversation.

The headline numbers, and why they hide the problem

MetricValue
Mention visibility100% (21 of 21 answers)
AI Visibility Score79
Average position#1.2, in 21 of 21 answers
Share of voice23% — first of 13 brands named
Mention quality79
Web search rate90%

Presence of 100% is real and it is worth having. It also compresses seven days of disagreement into a single flat line. Amazon is named in every answer because Amazon is Amazon; the question is what the answer does with the name once it is there, and that only shows up when you stop blending the assistants together.

The three assistants did not agree

GeminiChatGPTClaude
Amazon's AI Visibility Score708582
Amazon's average position#1.7#1.0#1.0
Walmart's score725464
Target's score584642
Ran a web search7 of 77 of 75 of 7

Fifteen points separate the best assistant from the worst, for one brand, on one question, in one week. Nothing about Amazon changed between those columns. What changed is which assistant was asked.

Mentionify metric cards filtered to Gemini only: 100% mention visibility, AI Visibility Score 70, average position #1.7 in 7 of 7 answers, 22% share of voice, 70% mention quality, 100% web search
The same dashboard with the model filter set to Gemini: score 70, average position #1.7. The blended view never showed either number.

Here is Gemini's week, day by day:

DayAmazon's position in Gemini's answerRanked above it
28 Jul#3Walmart, Target
29 Jul#2Walmart
30 Jul#2Walmart+
31 Jul#1
1 Aug#2Walmart
2 Aug#1
3 Aug#1

ChatGPT and Claude put Amazon first on all seven days without exception. Gemini put something else first on four of them. And it was not subtle about it — this is Gemini's own answer from 1 August, under its own heading numbered 1:

Walmart is the biggest competitor to Amazon for everyday essentials, and it is often considered the best option for grocery-focused shopping.

Amazon appears immediately below, at number 2, described as "the gold standard for sheer product variety". A perfectly complimentary paragraph, in second place.

Mentionify competitor leaderboard for the shopping question: Amazon 79 marked You with 21 mentions, Walmart 63 with 21, Target 49 with 20, Instacart 34 with 16, Costco 8, and eight more retailers below
Blended across the three assistants, Amazon leads by 16 points. The per-model view is where it stops leading.

Why one assistant answered differently

Before an assistant answers a question like this, it rewrites the question into its own web searches. We capture every one of them, because they are the clearest signal we have found for why an answer came out the way it did. Across the week the three assistants issued 39 searches, 34 of them distinct — and they were not looking for the same thing at all.

Gemini went shopping for competitors' prices. Twelve of its eighteen distinct searches were about a rival's membership; three were about Amazon's:

  • Walmart plus membership price 2026
  • Target Circle 360 subscription price 2025 2026
  • Instacart subscription cost perks 2026
  • Kroger Boost delivery same day 2026

ChatGPT went to check Amazon's homework. Nine of its ten searches named Amazon Prime directly, and one of them was a site-scoped lookup of Amazon's own membership page:

  • Amazon Prime delivery benefits United States free same-day one-day delivery 2026
  • Amazon Prime membership cost United States 2026 official Amazon
  • site:amazon.com/amazonprime Prime membership $139 per year $14.99 per month United States

Claude mostly re-asked the question. Six distinct searches, all near-restatements of what it was given, with a year bolted on: "best places to buy everyday products online fast delivery United States 2024".

That is the whole mechanism, visible in one screen. The assistant that spent its research budget pricing Walmart+ and Target Circle 360 came back and ranked Walmart first. The assistant that spent its research budget confirming Amazon Prime's own claims came back and ranked Amazon first, with the highest score of the three. Neither was being unfair. They read different pages.

Mentionify AI search queries table showing Gemini searching Walmart plus and Target Circle 360 membership prices, Claude re-asking the buyer question, and OpenAI searching Amazon Prime delivery benefits
The searches themselves, as recorded. Gemini pricing the rivals' memberships; OpenAI checking Amazon Prime; Claude repeating the question back.

Whose pages the assistants actually read

The 21 answers cite 177 pages across 70 distinct domains. Sorted by how often each domain was read:

DomainCitationsShare
walmart.com158.5%
target.com116.2%
aboutamazon.com95.1%
ebay.com95.1%
nbcnews.com84.5%
thekitchn.com63.4%
mixandmatchmama.com63.4%
getcartswap.com52.8%
amazon.com42.3%

Add up each company's own properties and the picture is blunter. Amazon's — amazon.com, its newsroom at aboutamazon.com, its payments subdomain — come to 14 of the 177 citations, or 7.9%. Walmart's and Target's own domains, counting their corporate newsrooms, come to 31. In the answer to "where should I buy everyday products", the assistants read Walmart's and Target's pages about themselves more than twice as often as they read Amazon's pages about Amazon.

Below the top few sits a long, unglamorous tail: a couponing blog, a family blog, a personal finance site, a fulfillment company's marketing blog. Those pages are not prestigious and it does not matter. They were in the retrieval set, so they were in the answer.

Mentionify cited sources table: walmart.com 15 citations at 8.5%, target.com 11 at 6.2%, aboutamazon.com 9 at 5.1%, ebay.com 9, nbcnews.com 8 and three smaller blogs below
Amazon-owned pages: 14 citations. Walmart-owned and Target-owned pages: 31.

The leaderboard, and the brands that were secretly one brand

Twelve other brands were named alongside Amazon over the week, thirteen rows counting Amazon itself. The board, scored on the same scale as the brand:

BrandScoreNamed in
Amazon7921 of 21
Walmart6321 of 21
Target4920 of 21
Instacart3416 of 21
Costco85 of 21
Best Buy32 of 21
Gopuff31 of 21

Walmart was named in every single answer too. So the real competitive picture is not "Amazon wins" — it is that this question has a three-name shortlist, Amazon, Walmart and Target, and everyone else is fighting for the scraps at the bottom of the paragraph.

One detail worth pausing on. Gemini rarely wrote "Walmart". It wrote "Walmart+", and it wrote "Target Circle 360" instead of "Target". Counted naively, that is four brands with four small scores instead of two brands with two big ones — and Walmart would never have surfaced as the threat it is. Our resolver checks each name against a real, reachable website before merging, so Walmart+ folds into walmart.com and Target Circle 360 folds into target.com. The difference between the naive count and the resolved one is the difference between a leaderboard that is decorative and one you can act on.

Where Amazon is actually weak

Every mention gets scored on five separate axes rather than one vague number, and the weak axis is never the one a brand expects:

AxisAmazon's average
Framing (how favorably it is described)92
Placement (where in the answer it sits)91
Prominence (how much of the answer is about it)90
Frequency (how often it recurs within the answer)67
Coverage (share of the answer's real estate)56

Framing at 92 means the assistants like Amazon. Coverage at 56 — the lowest band on the table, and it ranged from 35 to 75 across the week — means the answer is not really about Amazon: it is a comparison, and Amazon is one part of it. That is what a 23% share of voice looks like from inside a single answer, and it is the axis a competitor can attack most directly, by getting named in more of the pages that build the comparison.

Sentiment averaged 79 out of 100 across the 21 answers, ranging from 65 to 85, and the same two complaints recur in the answers at the bottom of that range: the membership is the most expensive of the three, and grocery delivery can carry extra fees. The tool writes those out as plain suggestions next to the numbers — one of them, verbatim from the run of 28 July, reads: "Highlight competitive grocery delivery fees to counter the mention of 'extra fees' for Amazon Fresh compared to Walmart."

The two days nobody searched

Claude answered without running a web search on 30 and 31 July. No searches, no citations, no retrieval — it answered from what it already believed, and it still put Amazon first.

Mentionify run history showing four dated daily batches, each with three per-model scores and a badge marking whether the assistant ran a web search or answered from knowledge
31 July, first card: Claude scores 80 and the badge reads Knowledge, not Web Search. That is an answer composed with no page behind it.

That is the whole reason web-search rate is a metric here and not a footnote. When an assistant searches, this week's pages decide the answer and a PR win can move it inside a fortnight. When it does not, you are being judged on years of accumulated public record and nothing you publish this month will touch it — the two supply lines behind that split are worth understanding before you budget a quarter. Amazon can be relaxed about that distinction. A brand that is three years old cannot, and the only way to know which mode you are in is to measure it.

If you are not Amazon

Amazon has the strongest possible position in this dataset and still loses one assistant out of three. Every finding here scales down badly for smaller brands, which is exactly why they are worth copying:

Check every assistant separately. A blended average would have shown Amazon at 79 and hidden the fact that Gemini scored a competitor higher. Whatever your number is, it is an average of disagreements.

Read the searches, not just the answers. The fan-out queries told us why Gemini ranked Walmart first long before any amount of staring at the answers would have.

Your competitors' websites are in your answers. Walmart's own pages outnumbered Amazon's own pages by more than two to one, in Amazon's category. You do not control the retrieval set — you compete inside it.

Once is not a measurement. Gemini gave Amazon third place, second place, second place, first, second, first, first. Any single day of that week supports a different conclusion. Seven days supports one.

None of that tells you what to do, which is deliberate — this is the measurement half. Bondo has written the other half as a prioritized, superstition-free to-do list.

What this ran on

Everything above came out of one project in Mentionify, over one week, with nobody watching it. The method is the same for any brand; these are the surfaces the study's data lives on:

  • Topics — the buyer question being tracked, with a per-topic score and its own detail page.
  • Models — Gemini, ChatGPT and Claude on the same question, each with its own tab, so disagreement is visible instead of averaged away.
  • Daily runs — seven runs, one a day, each a full pass of the topic across the three assistants at a fixed, published credit price. (Mentionify has a built-in daily schedule; for this study the runs were fired by our own automation at the same hour each morning.)
  • AI Visibility Score, mention visibility, average position, share of voice, mention quality, web-search rate — the six headline metrics, each with its period-over-period movement.
  • Trend over time with a period picker (7 / 30 / 90 days, a calendar month, or a custom range).
  • Competitors — the leaderboard above, scored on the brand's own scale, with the alias resolution that turns Walmart+ back into Walmart.
  • Sources — every cited domain, its exact URLs, citation counts and share of the total.
  • AI search queries — the fan-out searches, attributed to the assistant that issued each one.
  • Runs and archive — every settled run kept, and every raw answer readable in full.
  • Per-answer recommendations — the plain-language suggestions the judge writes beside each scored answer, like the grocery-fees line quoted above.
  • Report and exports — a printable report and CSV/XLSX export of any of these tables, scoped to the same period and the same model as the screen you are looking at.

If you want the short version for your own brand, run your buyer's question once, free — it is scored exactly the way these 21 answers were. What one check cannot show you is the part that mattered most here: movement, across days, across assistants. That needs a schedule.

FAQ

Does this mean Gemini is biased against Amazon?

No, and the data does not support that reading. Gemini ranked Amazon first on three of the seven days. What it did consistently was research competitors' membership pricing before answering, and on a question where price is a deciding factor, that research favored the cheaper membership. Different research, different answer.

Why score Gemini's answer with Gemini?

Every answer, from all three assistants, is scored by one judge model. That is the only way a ChatGPT number and a Claude number can sit in the same column. If each assistant graded its own homework, none of the comparisons in this post would mean anything.

Twenty-one answers is a small sample. Isn't that a problem?

For a statistical claim about Amazon's standing in the world, yes. For the actual question a brand has — is my visibility stable, and does it hold across the assistants my buyers use — it is enough to see the shape, and the shape here was unambiguous: two assistants steady at first place, one that ranked Amazon third on the very first run and never settled above second until day four.

Do you publish other studies like this?

Yes. Every brand we track this way is written up in full, with the same method and the same scoring, and they all sit together in the research index.

Can I see the raw answers?

Every answer is stored in full and readable in the run archive, which is how the quotes in this post were pulled. Nothing here is paraphrased from a summary.

Track what AI answers,
every day.

Your buyers' questions, asked to every assistant, scored on one rubric.