How people will actually shop in 2030

Pilea TeamAugust 18, 20268 min read

By Baptiste Vanpoperinghe, co-founder and CEO, Pilea

& Amaury Grange, co-founder and COO/CFO, Pilea

In 1995, the consensus was that the internet would kill the paper catalogue. It did not. Catalogues repositioned toward luxury, B2B and premium, and a number of them are still profitable thirty years later.

Between 2000 and 2010, the consensus was that e-commerce would kill the physical store. Store networks contracted and changed function, but they did not disappear.

Between 2010 and 2020, the consensus was that Google Shopping and Google Pay would disintermediate merchant sites. Almost nobody completes a purchase inside Google Shopping today.

Three times, the same prediction. Three times, a division of territory rather than a replacement.

I want to be careful with that pattern, "It did not happen the last three times" is not evidence that it will not happen this time. Every technology that did displace an incumbent was preceded by several that did not.

So what follows is a hypothesis, set out alongside the ones that compete with it. I have tried to give each of them its strongest form rather than its most convenient one, and to say what would move me.

The evidence, before the interpretation

Two datasets from the last eighteen months matter more than the rest, and they point in different directions.

The traffic shift is real and large. Adobe, working from more than a trillion visits to US retail sites, reports that traffic from AI sources to US retail sites grew 393% year over year in the first quarter of 2026, following a 693% year-over-year increase during the November to December 2025 holiday period.

More striking than the volume is the quality inversion. In March 2025, traffic arriving from AI sources converted 38% worse than non-AI channels. In March 2026, it converted 42% better, which Adobe reports as a record. Those shoppers spend 48% longer on site and view 13% more pages.

The transaction shift has, so far, gone the other way. In October 2025, Walmart announced a partnership with OpenAI allowing customers to buy through ChatGPT's Instant Checkout. Target, Instacart, PayPal, Etsy, Shopify and Salesforce followed with comparable arrangements. In March 2026, Walmart ended it.

The reported reasons are worth sitting with, because they are not the ones most people predicted. The problem was not consumer reluctance to buy inside a chat interface, and it was not conversational quality. According to reporting on the wind-down, Instant Checkout sourced its product data by scraping partner retailers' sites, which meant it could see product options but could not reliably verify whether an item was in stock or what the real delivery time would be. Customers reported wrong items appearing in their carts. Conversion ran below what Walmart achieves through its own channels. OpenAI's framing was that instant checkout is moving into apps, and that ChatGPT is concentrating on discovery before handing shoppers to retailers.

I should be honest about the epistemic status of that second dataset. Adobe's numbers are published research with a stated methodology. The Walmart account is trade and business press citing unnamed sources, and it describes one partnership over five months. It is a data point, not a law.

Four ways this could go

Here are the hypotheses I think are live. I hold the second one, but not by a wide margin.

H1. Full disintermediation. Agents eventually own discovery and the transaction. The current failures are early-implementation failures, not structural ones. This is the strongest case against my view, and the honest version of it is compelling: scraping was always a bridge technology, and once retailers expose structured inventory, pricing and fulfilment through protocols rather than HTML, the accuracy problem that killed Instant Checkout largely evaporates. Payment rails are being built. Returns are hard but not harder than things that have been solved before. Under H1, the merchant site becomes a fulfilment backend with a marketing department attached.

H2. The split. Commerce separates into two commerces that behave differently. Constrained, repetitive, low-consideration purchasing becomes invisible and automated. Chosen, expensive or identity-linked purchasing becomes assistant-mediated for the comparison work and stays on brand surfaces for the transaction. Sites do not die; they lose the monopoly on discovery, which they have arguably already lost.

H3. Retailer assistants win the interface. The models become distribution rather than destination. Walmart's actual response is evidence for this one: rather than withdrawing, it is embedding its own assistant into ChatGPT and Gemini and keeping the customer data and the transaction. If that pattern generalises, the winner is neither the model nor the classic storefront but the retailer-owned agent riding on someone else's user base.

H4. Plateau. AI referral becomes a meaningful channel, roughly the size of paid search, and then stops. Growth rates in the hundreds of percent are what small bases do. Under H4 the interesting question in 2030 is not agents at all, and most of what is currently being written about this, including this essay, ages badly.

I lean toward H2, with H3 as a near-neighbour that may turn out to be the same thing described from a different seat. But I would put meaningful probability on all four, and I do not think anyone reading the current evidence honestly can rule out H1.

Why I lean toward the split

Three reasons, each of which is contestable.

The first is the shape of the friction. The reasons a transaction is hard to move are unglamorous and slow-changing: tax calculation across jurisdictions, shipping and slot coordination against real inventory, payment processing and fraud liability, returns, and customer data ownership. None of these are user experience problems, which is why better conversational interfaces do not obviously solve them. The counterargument is that unglamorous infrastructure problems are exactly the kind that get solved quietly over five years.

The second is that the protocol layer being built right now assumes it. OpenAI's Agentic Commerce Protocol, Google's Universal Commerce Protocol and Shopify's MCP implementation differ substantially and agree that the merchant remains merchant of record. That is a meaningful signal about where the people building this expect value to sit. It is also, I should note, a signal that could be read as a negotiating position rather than a conviction.

The third is about the nature of the purchase itself. The former chief product officer of a leading digital experience analytics company made a point I keep returning to: visual and browsing-led shopping is structurally more resistant to agents than text search is, because the customer often does not know what they want until they see it. Delegation requires you to specify the outcome in advance. A board advisor at a marketplace software company draws a similar line at around $30 for repeat purchases, above which the buyer wants to be involved. Both observations feel right to me, and both are the sort of thing that has been said about consumer behaviour many times before and been wrong.

The claim I am more confident about than the thesis

Here is the part I would defend harder than any of the four scenarios, because it does not depend on which one wins.

There is now a substantial industry devoted to getting brands mentioned in AI-generated answers. It is a real problem. Adobe published data in April 2026 scoring how much of a retail page is machine-readable: US retail homepages averaged 75%, category pages 74%, and individual product pages 66%. Roughly a third of product page content, on average, is invisible to a model.

That work is not what we do, and I want to be clear about the boundary. Our interest starts one step later. Assume the visibility problem is solved. The model recommends you, and a customer or an agent arrives. Does what they find hold up?

Under H1, that question is the whole game, because an agent transacting on your behalf has no tolerance for inconsistency. Under H2 and H3, it is the game at the handoff, which is where the money changes hands. Under H4 it matters exactly as much as it always did, which was already quite a lot and was already under-measured.

The Walmart episode is the clearest available illustration, and its interest to me is not that it validates a prediction. It is that the failure mode was mundane. Stock status and delivery times could not be trusted, and wrong items ended up in carts. That is not a frontier problem. That is a data consistency problem of a kind that exists on most large storefronts today and has been tolerated for years because humans route around it.

That last point may be the durable one. Humans are error-correcting. They improvise past broken interfaces constantly, which is precisely why so many broken interfaces survive in production for years without anyone noticing. A shopper who sees a French navigation label in an Arabic interface shrugs. A shopper who finds a colour shown on the listing page is unavailable on the product page goes back and picks another. Systems reading the same surfaces do not shrug and do not go back. Whether or not agents take over the transaction, the share of traffic arriving through a non-improvising intermediary is going up, and defects that were absorbed become defects that propagate.

Questions I would want answered, in any scenario

Not recommendations. These are the things I find myself asking in conversations with digital teams, and where I have noticed that nobody in the room has a number.

What is our completion rate when a system, rather than a person, tries to get from search to checkout on our own site? Almost nobody has run this. It produces a number in an afternoon.

Which internal system is accountable for the accuracy of stock and delivery data at the moment it leaves our perimeter, now that it leaves through several doors rather than one?

Where do our defects concentrate: on pages, or at the seams between systems that are each individually correct?

Which of our surfaces is the customer actually on, and does our testing effort match that distribution?

If someone can answer those four, the strategic question of which scenario arrives matters considerably less, which is roughly the point.

Pilea is an e-commerce experience quality platform. It continuously inspects enterprise storefronts for visual, functional, content, localisation and merchandising defects, and ranks them by estimated business impact. Pilea analyses roughly 35,000 enterprise retail pages per month.

See what's quietly costing you conversions