There is a kind of result that looks finished. From 28 July to 3 August 2026 we put one question to three AI assistants once a day — what is the best ecommerce platform for a growing online store in the United States? — and collected 21 answers. Shopify was named in every one of them, in first position in every one of them, and recommended outright as the top pick in every one of them. Twenty one out of twenty one, three times over.
A record like that seems to close the subject. It does not, and the reason it does not is the number sitting next to it: share of voice, 23%. Of every brand name those 21 answers produced, fewer than a quarter belonged to Shopify. Being first is not the same as being most of the answer, and in AI answers the difference between the two is where the competition actually happens.
The setup
| Brand tracked | Shopify (shopify.com) |
| Question | "What is the best ecommerce platform for a growing online store in the United States?" |
| Assistants | Gemini, ChatGPT, Claude |
| Frequency | Once a day, seven days running |
| Window | 28 July – 3 August 2026 |
| Answers collected | 21 (7 runs × 3 assistants) |
| Answers where the assistant ran a web search | 21 of 21 |
Each assistant was asked the way a buyer asks — the plain question, no system prompt, no instruction to search, each model set to the tier a person on a free plan actually gets. Every answer was then scored by one judge model, so the three columns can be compared without an asterisk.
The same week we tracked Amazon on a retail question, and that one behaved completely differently: three assistants, and one of them ranked Walmart above Amazon.
What a perfect record looks like
| Metric | Value |
|---|---|
| Mention visibility | 100% (21 of 21 answers) |
| AI Visibility Score | 85 |
| Average position | #1.0, in 21 of 21 answers |
| Mention quality | 85 |
| Sentiment | 78 |
| Share of voice | 23% |
| Web search rate | 100% |
An average position of exactly 1.0 across 21 answers is unusual enough that it is worth stating plainly: across a full week, three different assistants, every one of which went and read the live web before answering, not one of them put another platform ahead of Shopify. ChatGPT's answer of 1 August opens on the word "Best overall" and the sentence ends in Shopify.
The three assistants still did not score it identically:
| Gemini | ChatGPT | Claude | |
|---|---|---|---|
| Shopify's AI Visibility Score | 83 | 91 | 81 |
| BigCommerce's score | 60 | 48 | 62 |
| WooCommerce's score | 56 | 47 | 52 |
Ten points between the friendliest assistant and the least friendly, with the ranking unchanged underneath. This is what a stable position looks like when you break it apart: the order holds, the generosity varies.
The eleven other names
Here is the part the headline metric cannot show you. The full board for the week:
| Brand | Score | Named in |
|---|---|---|
| Shopify | 85 | 21 of 21 |
| BigCommerce | 57 | 21 of 21 |
| WooCommerce | 51 | 21 of 21 |
| Adobe Commerce (shown as Adobe) | 17 | 10 of 21 |
| Wix | 8 | 5 of 21 |
| Squarespace | 5 | 3 of 21 |
| Commercetools | 3 | 2 of 21 |
| Big Cartel | 3 | 2 of 21 |
| Ecwid | 3 | 2 of 21 |
| OpenCart | 3 | 2 of 21 |
| Square Online | 2 | 1 of 21 |
| Salesforce Commerce Cloud | 2 | 1 of 21 |
Three brands were named in all 21 answers, not one. BigCommerce and WooCommerce were as unavoidable in this question as Shopify was — they simply arrived second and third, described in the language of exceptions: for complex B2B catalogs, for WordPress-heavy sites, for people who want total control. The fourth name, Adobe Commerce, made fewer than half the answers. After that the board falls off a cliff: eight platforms sharing single-digit scores, each appearing once or twice all week.
That shape is the real finding, and it is the same shape we keep seeing on questions like this. An AI answer to a "best X" question is not a ranked list of everyone in the market. It is a shortlist of about three names, with a scattering of qualified alternatives underneath. Shopify's problem is not losing first place. It is that first place comes with two permanent co-tenants — and if your brand is not on a shortlist like this one yet, getting onto it is its own body of work.

The one place the order did move was the second slot. On 1 and 2 August, ChatGPT put WooCommerce above BigCommerce, then the following day it went back. Nothing about the two products changed overnight. The pages ChatGPT happened to read that morning did.
The five axes underneath the score
A mention is not one thing, so it is not scored as one thing. Shopify's week, broken out:
| Axis | Score |
|---|---|
| Framing (how favorably it is described) | 97 |
| Placement (where in the answer it sits) | 96 |
| Prominence (how much of the answer is about it) | 96 |
| Frequency (how often it recurs within the answer) | 76 |
| Coverage (share of the answer's real estate) | 63 |
Four of the five axes are close to their ceiling. The fifth, coverage, sits at 63 — it never rose above 75 in any answer — and coverage is simply the share-of-voice problem seen from inside a single answer. Every assistant leads with Shopify, describes it warmly and at length, and then gives the rest of the reply to the two situations where you would choose something else.
This is the axis worth arguing with, and the judge's own notes across the week kept pointing at the same two arguments the alternatives were winning on. Transaction fees, which BigCommerce uses as its whole pitch. And what one run called "app creep" — the sense that the platform is cheap until the apps arrive. Those objections were in the answers before they were in our notes; they are what the retrieved pages said.
Where the answers came from
The 21 answers cite 198 pages across 75 distinct domains. The most-read ones:
| Domain | Citations |
|---|---|
| shopify.com | 19 |
| bigcommerce.com | 14 |
| youtube.com | 11 |
| wise.com | 10 |
| mrpeasy.com | 6 |
| emailvendorselection.com | 6 |
| woocommerce.com | 6 |
| taxually.com | 6 |
| jetfuel.agency | 5 |
| elementor.com | 5 |
Two things stand out. The first is that Shopify's own properties are the most-read source in its own category — 26 citations counting the help center, changelog and app store, or 13% of everything the assistants read. That is not the normal result. On the Amazon question, in the same week, the brand's own pages accounted for 7.9% of citations and were outnumbered two to one by its competitors'. Shopify's documentation and pricing pages are doing real work: ChatGPT reached for shopify.com/pricing directly rather than taking a third party's word for what a Shopify plan costs.
The second is who else is in that list. There is no publisher of record here. Forbes appears three times. Zapier appears three times. The domains doing the heavy lifting are a money-transfer company's blog, a manufacturing-software blog, a tax-compliance blog, a marketing agency's comparison post, a page-builder's listicle. Sixty-one of the 75 domains were cited three times or fewer.

If you have been waiting for a mention in a famous publication to move your AI visibility, this table is the argument against waiting. The pages that decide these answers are mostly small, specific and unfashionable, and they are reachable. What decides whether an assistant reads any page at all — rather than answering from memory — is the retrieval step itself.
What the assistants searched for
We record the searches each assistant issues before answering, which is the closest thing to reading its mind. Forty-four searches across the week, 33 of them distinct, and the three assistants had visibly different habits.
Gemini rephrased the buyer's question fourteen different ways and searched each one, never naming a brand: "best ecommerce platform for scaling online store US", "top ecommerce platforms scaling business 2026 US". It went looking for the category consensus and reported it back.
ChatGPT named brands immediately and then went to verify specifics on official pages — Shopify's pricing plans, the app store's published app count, and a site-scoped lookup of Adobe Commerce's pricing. It is the only assistant that consistently checked a claim against the vendor's own page before repeating it.
Claude issued the most searches (19) but only 9 distinct ones, repeating "Shopify vs BigCommerce vs WooCommerce comparison scalability 2025" across days. It also searched for a specific publisher's roundup by name.

Those habits explain the citation table above completely. The assistant that verifies against official pages is the reason shopify.com is the most-cited domain; the assistant that searches for comparison articles is the reason the long tail of small comparison blogs exists at all.
What to take from a perfect scoreboard
Rank and share of voice answer different questions. Shopify's rank has no room left to improve. Its share of voice does, and that is the number that moves when a competitor gets added to one more comparison post.
Watch the co-tenants, not the leader. BigCommerce and WooCommerce appeared in 100% of answers. If you are competing in this category, those two brands are what you are actually being compared against, in the same paragraph, every time.
Your own documentation is a retrieval target. Shopify's pricing and help pages were read more than any third party's. Publishing the page that answers the specific factual question — what does it cost, what is included, how many apps — is what makes an assistant cite you instead of paraphrase someone else's guess about you.
The objections in the answers are the roadmap. Transaction fees and app costs came up week after week, in every assistant. That is a content brief, not a complaint.
Every one of those four points is readable off a dashboard rather than argued from intuition, which is the only reason this post has numbers in it at all. How the scoring works is written up separately.
What produced these numbers
One project, tracked for a week, running on its own — one of several public-brand studies we have published. The surfaces this study's data lives on:
- Topics — the tracked buyer question, with its own score and detail page.
- Models — the three assistants side by side, each with its own tab and its own baseline, so a per-model view compares like with like.
- Daily runs — seven runs, one a day, each at a fixed credit price. (Mentionify has a built-in daily schedule; this study's runs were fired by our own automation at the same hour each morning.)
- AI Visibility Score, mention visibility, average position, share of voice, mention quality, sentiment, web-search rate — the headline metrics, each with its own period-over-period delta.
- Trend over time across a rolling window, a calendar month, or an explicit date range.
- Competitors — the twelve-brand board above, every rival scored on the brand's own scale.
- Sources — every cited domain, the exact URLs behind it, and its share of total citations.
- AI search queries — the fan-out searches, attributed to the assistant that issued each.
- Runs and archive — every settled run preserved, every raw answer readable.
- Per-answer recommendations — the plain-language suggestions the judge writes beside each scored answer, which is where the transaction-fee and app-cost themes surfaced.
- Report and exports — a printable report and CSV/XLSX of any table, scoped to the same period and model as the screen.
To see one answer scored the same way for your own brand, a no-account check takes a couple of minutes. What it cannot show you is whether your position holds, which is the only thing this dataset really measures.
FAQ
Is 23% share of voice bad?
It was first place out of twelve brands, so no. It is the honest measure of a market where AI answers name three platforms as a matter of course. The useful thing about tracking it is not its absolute value but its direction: a competitor working its way into more of the cited comparison pages shows up here before it shows up anywhere else.
Why does the same question produce different orders on different days?
Because every one of these 21 answers was composed after a live web search, and the retrieval set changes. That is exactly what the second-place flip between BigCommerce and WooCommerce was. It is also why a single check tells you very little and a week of checks tells you the shape.
Do the assistants agree more on some questions than others?
Yes, and the contrast is stark. On this question, all three put the same brand first for seven straight days. On the retail question we ran the same week, one assistant scored a competitor higher than the brand we were tracking. Agreement is a property of the question, not a constant — which is why the per-model view is not a nice-to-have.
How often do I need to ask, then?
Often enough to see a rate rather than a screenshot. This study used seven consecutive days because that was enough for the pattern to stop moving; weekly is a reasonable floor for most categories, and the plans are priced per check so the cadence is yours to pick.
Where do the quotes and numbers come from?
Every answer is stored whole. The scores are recomputed from the same stored records, by the same scoring function that draws the dashboard, so a number in this post and a number on the screen cannot drift apart.



