What AI Shopping Agents Actually See on Your Storefront

Most of the writing about agentic commerce is about protocols: who signs the payment mandate, which standard wins, what share of transactions will be agent-initiated by 2030. That debate matters, and it is being had competently elsewhere.
This article is about something narrower and, for anyone running a large storefront this quarter, more immediately actionable. Before an agent can buy anything on your site, it has to be able to use your site. And the way an agent uses a storefront is close to the way a keyboard user does, which means the defects that block agents are not new, exotic, AI-era problems. They are the defects that have been sitting in your storefront for years, invisible to your monitoring, and newly expensive.
We publish a lot of real defect evidence. In our twelve-defect teardown, every one of the twelve was live on the production site of a retailer doing between €200M and €25B in online revenue, and not one produced a technical alert. Rereading that list with agents in mind is uncomfortable. At least five of the twelve would stop an AI shopping agent cold, and two of them would make it confidently tell a customer something false.
The short version
An AI shopping agent does not look at your storefront. It reads it. It parses the DOM or a rendered text extraction of the page, identifies interactive elements by their role and accessible name, and acts on them programmatically. It presses Enter rather than clicking a dropdown link. It trusts the first price it can parse. It rarely opens modals. It does not infer that the greyed-out swatch means out of stock, unless something in the markup says so.
That gives you a workable definition. Agentic commerce readiness is not a new integration project. It is the machine-readability and behavioural correctness of the storefront you already have. The protocol work sits on top of that. If the page underneath is wrong, a perfectly implemented checkout API just lets an agent complete a purchase of the wrong thing, faster.
How an AI shopping agent actually navigates a storefront
There is more than one architecture in production today, and the differences matter less than people assume.
Text and DOM extraction. The agent fetches the page, renders it, and flattens it into a structured text representation. Headings, links, form fields, buttons, ARIA roles, accessible names. This is the dominant mode, because it is cheap and reliable.
Structured feeds. Where a merchant exposes a product feed or a commerce API, the agent may prefer it for catalogue data, then return to the storefront for availability, delivery and checkout. The emerging protocol work, including the Agentic Commerce Protocol and Google's AP2, is largely about formalising this path.
Vision. Some agents screenshot and reason over pixels when the DOM is unusable. It is slower and more expensive, so it is usually the fallback rather than the default.
Keyboard-first interaction. Whichever mode it uses to read, when an agent acts it does so through programmatic events: focus, keypress, form submission. It does not move a mouse across your viewport and click a coordinate.
Put those four together and a pattern falls out. Agents are strongest exactly where traditional analytics is strongest, on the happy path through well-structured pages, and they fail exactly where your existing defects live: in keyboard handlers, in modals, in state transitions, in data that disagrees with itself between two systems.

The defects that stop agents, with real examples
Every example below is one we have found live on an enterprise storefront. All are anonymised and no retailer is named.
1. The search bar where the Enter key does nothing
We found this on a US direct-to-consumer travel goods brand doing over €200M online. A customer types a query, presses Enter, and nothing happens. The results appear only if you click the small "View All Results" link inside the dropdown.
For a human, this is an annoyance that costs a fraction of a second and some goodwill. For an agent, it is a dead end at the first step of every single shopping journey. Submitting a search form is the canonical agent action. An agent that presses Enter and receives no navigation event has no reason to hunt for an undocumented link inside a dropdown it may not have rendered. It reports that the site has no results for the query, and it moves to the next retailer.
Nothing caught this because the search endpoint was healthy and returning results correctly. The dropdown rendered. The failure was a missing keyboard event handler, which is not an error condition, and click-based test suites pass happily on a site where the keyboard is broken.
2. "Sort by price, low to high" that does not sort by price
Same retailer. Applying the ascending price sort returned a $157 item ahead of a $25 item, on almost every category page.
A human scanning a grid notices the disorder within a second or two and works around it. An agent asked to find the cheapest option in a category does the rational thing: it applies the price sort and takes the top results. It has no independent reason to distrust the order, so it does not verify it. The output is a confident, specific, wrong recommendation delivered to the customer in the agent's voice rather than yours.
This is the category of defect that worries us most in an agentic world. It does not block a purchase. It produces a false one.
3. Out-of-stock sizes that stay selectable inside the size guide
European sporting goods retailer, over €200M online. Sizes correctly shown as unavailable on the product page were selectable through the size guide overlay, because the modal read availability from a different source than the main selector. The customer only learned the size was out of stock at checkout.
Modals are where agent journeys go to die. They are separate components, frequently owned by a different team or a third party, and they are rarely in anyone's test matrix. Crawlers and regression suites rarely open them, which is precisely why defects accumulate inside them. An agent that does open one and acts on what it finds inherits whatever inconsistency lives there.
4. The product that disappears when you close the delivery-slot selector
Large European grocery group. A customer adds a product, is shown a delivery or collection slot selector, and closes it without choosing. The product is silently dropped from the basket.
Agents abandon interactions constantly. They open a thing, determine it is not what they need, and back out. That is normal agent behaviour and it is exactly the path that test scripts never walk, because a test script exists to complete the flow. The state you reach by not doing the thing is rarely enumerated, and it is where agents spend a surprising amount of their time.
5. Free shipping at €80 that does not trigger at exactly €80
Large French apparel retailer group. A basket of exactly €80.00 did not trigger the free-shipping confirmation. €80.01 did.
Agents optimise to thresholds. That is one of the most useful things they do for a customer: build the basket that clears the free-shipping bar with the least spend. An agent that lands on exactly €80.00 and sees no confirmation concludes the threshold is wrong, or that the offer does not apply, and either adds an item nobody wanted or reports your shipping as more expensive than it is.
Boundary conditions are the single most reproducible category of defect we find, because thresholds are where business rules and code meet and where nobody writes the equality case.
6. Prices and promotions that disagree with themselves
Two from the same set. A product priced at $118 in the catalogue displayed at $98 on the storefront, with no promotion, no strikethrough, no discount badge. And a promotional expiry date rendering on the storefront as the year 2300.
Humans are excellent at absurdity detection. A date in the year 2300 reads as a bug instantly. An agent has no such instinct and no baseline for what your prices should be. It extracts the number in the price node and treats it as the price. If your catalogue, your storefront and your feed disagree, the agent will surface one of them to the customer, and it will not necessarily be the one you would have chosen.
7. Untranslated strings and locale-dependent failures
A marketplace operating in eight markets and three languages had French source strings surviving in its English and Arabic interfaces: navigation categories, product type filters, call-to-action buttons. Separately, a grocery retailer blocked account registration entirely for users whose device locale was set to US English, because of a date format mismatch in the birth date field.
Filter labels and taxonomy are exactly what an agent uses to navigate a catalogue. An agent working in English that encounters French filter labels has to guess, and it may simply not find the category. The registration case is worse: a whole class of user, agents included, blocked at the door by an axis nobody tests.
The five structural patterns, read through an agent's eyes
Sorting the full set of what we find across roughly 100,000 enterprise retail pages per month, five patterns account for nearly all of it. Each one gets more expensive when the visitor is an agent.
Boundary conditions. Thresholds, exact values, first and last items, the edge of a range. Business rules are written in round numbers and tested in round numbers. Agents aim directly at the boundary, because optimising to it is the job.
State transitions rather than states. The defect is not on a page. It is in what happens between two pages, or in the state you reach by abandoning an interaction. Agents abandon constantly.
Locale as an axis nobody tests. Device language, date format, currency, character set, right to left. Agents operate across locales by default, and often in a language that is not the one QA ran on.
Derived and third-party surfaces. Modals, size guides, cross-sell modules, promotional themes. Fixes applied to the main surface do not propagate to the fork, and agents cannot tell which surface is authoritative.
Data consistency between systems that are each internally correct. Listing page and product page disagree. Filter and selector disagree. Catalogue, feed and storefront disagree. No single system is broken, and the agent picks whichever one it parsed.
None of these five are detectable by asking a system whether it is healthy. All five are detectable by asking whether someone trying to buy something can actually do it. That has always been true. Agents just remove the human capacity to compensate for it.
What actually changes when the shopper is an agent
Three things, and only three, but they compound.
No compensation. Every storefront in production is propped up by human tolerance. Customers re-sort, re-search, scroll past the broken module, click the link when Enter does not work. That entire buffer disappears. A defect that costs a human four seconds costs an agent the transaction.
No signal. You will not see this in analytics. A customer who cannot reach result 60 does not generate an event describing what they failed to find, and an agent that gives up on your site does not fill in an exit survey. Agent traffic that fails silently looks like agent traffic that never arrived.
Errors propagate with your name on them. This is the genuinely new risk. When an agent reads your broken price sort and tells a customer your cheapest jacket is €157, the customer does not experience a sorting bug. They experience your brand being expensive. The defect has been laundered into a recommendation, and you never see the conversation.
How to test agentic commerce readiness
A practical sequence, roughly in order of what it costs you.
Run your top ten journeys keyboard-only. No mouse. Search, filter, sort, select a variant, add to basket, apply a code, reach checkout. Anything that requires a click to work is invisible to a large share of agents and to every keyboard user you have.
Check that roles and names are real. Buttons that are
divs, inputs with no accessible name, swatches whose availability lives only in a CSS class. If the state is not in the markup, the agent cannot read it.Test your thresholds at the threshold. Exactly €80.00, not €79.99 and €80.01. Then the first and last item of every paginated set.
Walk the abandonment paths. Open every modal and overlay and close it without completing. Check what happened to the basket.
Diff your surfaces. Catalogue against storefront, storefront against product feed, listing page against product page, main selector against modal. Agents may read any of them.
Re-run everything in a second locale. Different device language, different date format, different currency. Include one right-to-left market if you operate one.
Read your own pages as text. Strip the CSS and look at what is left. That flattened text is close to what the agent gets. If you cannot shop from it, neither can the agent.
None of this is agent-specific work. It is experience quality work that agents make impossible to defer. Steps one, two, five and six will improve conversion for humans this quarter, independently of whether a single agent ever touches your site.
The honest part
We can tell you the mechanism. We are more careful about the euros.
Our impact model assigns each of roughly 32 defect categories an estimated conversion impact drawn from external benchmarks, multiplied by page traffic and order value. It produces a number. It is a modelled number, and one enterprise customer pushed back on our methodology in a way that was partly correct.
The one controlled measurement we have is human, not agentic: Tikamoon, a European furniture brand with over €120M in revenue, A/B tested a set of Pilea-identified fixes across more than 700,000 users over 28 days and measured a 6 percent increase in sales, driven by both conversion rate and average order value. One test, one site, one category.
On agent traffic specifically we have no controlled measurement to offer, and neither does anyone else at the time of writing. The share of enterprise retail revenue that is agent-initiated today is small, and forecasts for where it goes vary by an order of magnitude depending on who is selling what. We are not going to add a number to that pile.
What we will say is narrower and, we think, defensible. The defects that block agents are already live on your storefront, they are already costing you human conversions, and fixing them is not a bet on a forecast. It is the same work either way. The agentic scenario just removes your last excuse for leaving it undone.
Frequently asked questions
Can AI shopping agents use my e-commerce site today? Partly. Most agents can read a well-structured catalogue and product page. They fail on keyboard-only interactions that are click-dependent, on content inside modals, on state transitions reached by abandoning a flow, and wherever two systems on your site disagree about price, availability or variant data.
What is agentic commerce readiness? The degree to which a storefront can be read and operated correctly by a software agent acting for a customer. In practice it is three things: machine-readable structure (roles, accessible names, state in the markup), behavioural correctness (keyboard submission, sorting that sorts, thresholds that trigger), and data consistency across catalogue, feed and storefront.
Do I need to implement an agentic commerce protocol first? Protocols such as ACP and AP2 standardise how an agent transacts once it has decided what to buy. They do not fix a broken price sort or a modal that reports the wrong stock. Discovery and evaluation still happen on your storefront, so storefront quality comes first in practice whatever your protocol roadmap looks like.
Will accessibility work make my site agent-ready? It is the single highest-overlap investment available, and it is the right place to start. Semantic roles, accessible names, keyboard operability and focus management are exactly what agents parse. It will not cover the other half, which is data consistency between systems and the correctness of business rules at their boundaries.
How do I know if agents are already failing on my site? You largely cannot, from analytics. A failed agent journey produces no distinctive event. Server logs and user-agent analysis give you a partial view of arrival, not of outcome. The only reliable method available today is to run the journeys yourself, keyboard-only and text-only, and see where they break.
About Pilea. Pilea is an e-commerce experience quality platform. It continuously inspects enterprise storefronts for visual, functional, content, localisation and merchandising defects, and ranks them by estimated business impact. Pilea analyses roughly 100,000 enterprise retail pages per month and classifies findings against a taxonomy of approximately 32 defect categories. All defects in this article are anonymised, and no retailer is identified.
See what's quietly costing you conversions
More from the blog

The Black Friday Countdown: Why a Proactive UX/UI Audit Is the Best ROI Move of Q4 2026
Black Friday 2026 is 11 weeks out. Why a proactive UX/UI audit before peak traffic hits is the highest-ROI line item on the Q4 checklist, with real defect examples — forms that crash under load, CTA buttons broken on mobile — that only appear during the spike.

The Hidden Risk of Enterprise Shopify Replatforming
68% of conversion-impacting issues surface after go-live. Here’s why enterprise Shopify replatforming requires continuous monitoring to protect storefront performance and conversion.

How people will actually shop in 2030
AI is reshaping how people discover products, but will it replace the storefront? Four scenarios for commerce in 2030, what the evidence says, and the one thing that matters whichever one wins.