Computational Marketing

How AI Shopping Agents Select Products Without Human Input

Retailers must rebuild product data infrastructure for AI agents, not webpages.

Staff Writer · · 11 min read
Cover illustration for “How AI Shopping Agents Select Products Without Human Input”
Bots as a New Buyer Class · September 10, 2026 · 11 min read · 2,475 words

Some agents already buy without checking back with a human first, and the logic they use to pick one product over another has become the newest battleground in retail. Most retailers don't know the rules of that fight yet, and a fair number don't know the fight has started.

Deloitte lays out the path in five stages: assisted discovery, assisted shopping, agentic shopping, autonomous shopping, and agent-to-agent commerce. Each stage strips out one more human decision point. At the far end sits agent-to-agent commerce, where a shopper's AI negotiates directly with a retailer's AI: no webpage, no cart managed by a person. Large language models, retrieval-augmented generation, and standardized protocols are what let an agent discover, compare, negotiate, and transact, rather than just suggest and wait.

Projections suggest $385 billion of U.S. e-commerce will move to agentic channels by 2030. Black Friday 2025 already gave a preview: Adobe Analytics recorded an 805% year-over-year jump in AI-driven traffic to retail sites, and Salesforce reported AI touching one in five Cyber Week orders globally. At that volume, how an agent chooses stops being a side question. It's the whole game, and most of it is being played in places most brands aren't even looking.

The infrastructure that makes autonomous selection possible

Diagram: Five Stages From Human Shopper to Agent-to-Agent Commerce. Visualizes: Visualize the five-stage progression Deloitte maps from assisted discovery to fully autonomous commerce: (1) Assisted Discovery, (2) Assisted Shopping, (3) Agentic…

None of this works without plumbing, and three protocols do the standardizing.

MCP, the Model Context Protocol, came out of Anthropic and was handed to the Linux Foundation's Agentic AI Foundation in December 2025. It standardizes how an agent connects to outside tools, data sources, and services, which fixes the bot-sprawl problem: instead of building a separate integration for every backend, one agent queries many databases through a single interface. A2A, the Agent-to-Agent Protocol, launched under Apache 2.0 from Google in April 2025 and moved to Linux Foundation stewardship that June. It handles discovery, delegation, and lifecycle updates across workflows that involve multiple agents. AP2, the Agent Payments Protocol, governs the money part: the authorization and verification layer that makes agent-initiated payments trustworthy and auditable.

Together, these three protocols define what "agent-ready" actually means for a merchant, and it has nothing to do with a nicer product page. Most merchants still get this backwards: they spend the design budget on the webpage, slow, cluttered, full of layout noise an agent has to guess around, while the buying decision actually happens through a clean API query that never touches that page at all. That's money spent on the wrong problem. The product data infrastructure that feeds agent queries is already evolving, demanding detail that goes well past a keyword. That's where agent selection actually starts, before any ranking signal or review pattern enters the picture.

Deloitte found that 63% of global retailers believe companies without AI agents will fall behind within two years. But being ready to answer an agent's query is a harder, and separate, infrastructure problem than being ready to answer a search crawler. A brand's discoverability to an agent gets decided before a single selection signal fires, and it comes down to whether the product data exists in a form a machine can read at all.

The layered signals agents actually use to choose a product

Research out of Columbia Business School (Paper No. 381574, Allouah, Besbes, Figueroa, Kanoria, and Kumar, first posted August 2025 and revised that December) built what the authors call the ACES framework, an audit of how agents actually decide across providers. The decision breaks into layers, and the layers don't act alone.

The first layer is the plain one: structured data against stated requirements. Price, ratings, availability, compatibility. Tell an agent "under $1,500, 16GB RAM" and it filters on those terms before anything subtler gets a say.

The second layer is position, and here things stop being intuitive. The Columbia researchers ran randomized trials and found agents carry strong position biases that vary by provider and by model version. The bias survives even in text-only, headless interfaces where no visual layout exists at all, which rules out design as the explanation. Where a product sits in a data feed or an API response is not a neutral fact. Agents aren't immune to position effects just because there's no human eye scanning a page.

The third layer involves badges, and the results cut against a decade of paid-search logic. Products labeled "Sponsored" saw selection rates drop, while products labeled "Overall Pick" gained real attention, all else held constant. Agents actively discount advertising and reward what reads to them as a platform-level endorsement. Paying for a sponsored slot can actively hurt a product's odds in an agentic channel. Anyone budgeting for agentic commerce the way they'd budget a search-ads campaign is funding the wrong fight, and that mistake is going to be common for a while yet.

The fourth layer is review and rating sensitivity, and the pattern here is variation itself. There's no universal weighting: what GPT-4.1 leans on heavily, Gemini 2.5 Flash might treat as a minor input. The fifth layer sits earlier than all of this, baked into the model before retrieval even starts: training-data priors. LLMs form associations with brands, affordability, reliability, speed, well before a live search runs, and those built-in impressions shape whether content gets cited even when it technically qualifies. A brand with strong presence in the training corpus starts ahead, and no amount of tuning a product page in the moment closes that gap fully.

These five layers don't run in sequence. They interact, and the interaction is where most brands lose ground: a product with excellent structured data but a "Sponsored" tag can still lose to a thinner listing carrying an "Overall Pick" label.

Diagram: The Five Layers Agents Use to Choose a Product. Visualizes: Visualize the five layered signals the Columbia Business School ACES framework (Paper No.

Why agent preferences are model-specific and unstable over time

Agents don't spread demand evenly across an assortment. They pile onto a small set of "modal" products and mostly ignore the rest, a winner-take-most pattern sharper than anything typical human browsing produces.

The Columbia researchers tested this directly with fitness watches, giving Claude Sonnet 4, GPT-4.1, and Gemini 2.5 Flash the same assortment. The models diverged hard: Claude favored one brand nearly twice as often as the other two did. Each model runs its own miniature market with its own demand curve, so a brand's position in "the AI market" depends entirely on which AI is asking. There is no AI market, singular. There are at least three, and probably more by the time this gets read.

Worse, these preferences don't hold still. A model update can reshuffle market share overnight, and a product dominating selection under one version can vanish from the modal set the next. Consistency in these systems isn't a default state; it's something a brand has to keep earning, on a schedule it doesn't control.

Treating this like classic SEO, where one set of best practices roughly holds across Google updates, misreads the problem. There is no single algorithm to optimize against here. Tune a listing for GPT-4.1, and it's entirely possible to be quietly mis-tuning it for Claude or Gemini in the same breath. Monitoring across models, continuously, isn't a nice-to-have. It's the only way to know whether a brand still sits in the modal set or got shuffled out by the last update.

How description quality and content signals shift selection share

Sellers are already testing this. In a mousepad experiment run by the Columbia researchers, sellers used AI tools to make small edits to product descriptions aimed at automated buyer preferences. Roughly a quarter of those edits produced a statistically significant jump in selection share after a single round of changes. One mousepad's selection share rose by more than 20 percentage points once its description was rewritten to read better to the GPT-4.1 agent evaluating it.

The winning edits weren't keyword stuffing. They made the description more legible to the agent's evaluation logic: explicit attribute matching, unambiguous claims, benefit statements laid out in a structured, checkable way. Worth sitting with the other three-quarters too, since they did nothing measurable. The opportunity is real, but it rewards precision, not volume, and it isn't stable across models.

The same logic carries past the product page. Research into AI search behavior indicates that most brand mentions come from third-party pages rather than the brand's own site, which means agents draw heavily on review sites, publications, and earned coverage rather than product descriptions alone. Claims an agent can check, verifiable specs, named comparisons, third-party confirmation, get weighted more than claims it can't verify.

Optimizing a single description is necessary, then, but nowhere near sufficient. Chase the description and ignore the ecosystem around it, and the effort caps out fast. The body of independent content surrounding a product carries real weight in the selection decision, and that's exactly the territory the next section covers.

Brand visibility in AI systems as a structural selection factor

Traditional search optimization is losing ground fast. Ahrefs, comparing 300,000 keywords between December 2023 and December 2025, found that where AI Overviews appear, click-through rate for the top-ranking page drops by up to 58%, from 7.3% down to 1.6%. Then, on January 27, 2026, Google switched AI Overviews to Gemini 3, and a large share of previously cited domains got replaced overnight. The overlap between a page's top-10 organic rank and its AI Overview citation fell hard, and it kept falling depending on which dataset you check.

Two disciplines have grown up around this shift, and conflating them is the mistake most marketing teams are making right now. Generative Engine Optimization, or GEO, structures content and brand presence so systems like ChatGPT, Perplexity, and Google's AI Overviews cite and recommend a brand across many different prompts. It has nothing to do with holding a ranking position. Answer Engine Optimization, or AEO, structures content to get lifted directly as an answer, and top organic rank still matters here specifically: roughly 40% of Google's AI Overviews cite a top-10 organic result, and roughly 70% cite something in the top 100. A brand running an AEO playbook and calling it GEO is solving the wrong half of the problem, full stop.

There's also a real gap between being cited and actually mattering to the answer. Research running controlled prompts across ChatGPT, Google AI Overviews, and Perplexity found Perplexity cites more sources per query but each source contributes less to the final answer on average, while ChatGPT cites fewer sources but each one carries substantially more weight. A brand can show up in a footnote and add nothing to what the user actually reads. Counting citations without asking whether they moved the answer is measuring the wrong thing.

GEO practice runs mostly on strategy: positioning, ecosystem presence, brand authority, with the technical piece a smaller slice of the work. The same ratio holds for visibility with shopping agents. Structural presence across the web beats any single page-level trick. Similarweb's 2025 Generative AI report found AI chatbot referral traffic grew 357% year over year, reaching 1.1 billion referral visits in June 2025 alone. The brands capturing that traffic are the ones with real AI-visible presence, not just a high organic rank.

What monitoring agent selection actually requires in practice

Agent preferences are model-specific and they shift with every update, so a one-time audit is a snapshot that goes stale within weeks. Treating it like an annual SEO review misreads what kind of problem this is.

What needs tracking, on an ongoing basis, is a short but demanding list. Mention rate matters more than rank: how often a brand shows up across varied prompts and across agent providers. Selection share needs to be broken out by model, since GPT-4.1, Claude Sonnet 4, and Gemini 2.5 Flash diverge in their selection behavior enough that lumping them together erases meaningful differences. Badge and endorsement status needs its own line, tracking whether a product lands as a platform-endorsed pick or gets penalized as sponsored. The health of third-party citations needs watching too, since those feed the agent's underlying impression of a brand and go stale on their own timeline.

Citation and absorption need separate tracking, and conflating them hides real risk. Whether a brand gets cited and whether it actually gets absorbed into the generated answer are two distinct measurements. A brand can look fine on one and be invisible on the other.

For agencies handling several client brands at once, doing this kind of tracking by hand, across every model, for every client, stops being sustainable past a certain portfolio size. It needs a system built to roll up results across the whole roster while still showing per-client detail in one place. Clients need AI selection performance reported as its own line item, not folded into organic traffic or a generic share-of-voice number. That distinction matters because most B2B buyers, per 6sense's 2025 Buyer Experience Report, already use LLMs somewhere in their buying process, and clients want proof their brand shows up where those buyers are actually deciding, not a vague assurance that it probably does.

This is precisely the gap Thrad is built to close: agencies managing AI visibility across a client portfolio get combined analytics across every account, granular controls at the individual client level, and weekly reports built for that specific client rather than a generic export. The training side matters as much as the dashboard does. Account managers get trained to explain what the data actually means, not just hand a client a login and a chart.

The trust and accountability gaps that still shape how far agents can go

Finding the right product is one problem. Proving an agent has the authority to spend someone's money is a different one entirely, and it sits closer to law than to engineering.

AP2's mandate structure, intent, cart, payment, handles the technical half of that problem well enough. Consumer trust hasn't caught up. Visa's 2025 "Earning Consumer Trust" report found most consumers voice concern about data privacy in AI commerce, and that figure holds even among people already interested in using it. Liability is the harder open question, and it's the one nobody in this chain wants to own first. When an agent buys the wrong thing, who answers for it? No global legal consensus exists yet, and until one does, every player, retailer, model provider, payment processor, carries reputational and regulatory exposure that nobody has fully priced in.

There's a bias risk sitting underneath all of it, and it's the one least likely to get fixed by better engineering. If training data over-represents certain brands or certain demographics, an agent's recommendations will systematically favor those same brands, quietly, at scale, invisible to the shopper on the other end of the conversation. Better product descriptions don't touch this. Cleaner structured data doesn't touch it either. The bias sits upstream of both, baked into the training data itself, which makes it the hardest problem on this list, and the one the industry has made the least visible progress on.

Sources

  1. Agentic Commerce: AI Shopping Agents Guide 2025
  2. Ecommerce AI Agents in 2026 (Shopping’s Next Big Shift)
  3. What Is Your AI Agent Buying? Evaluation, Biases, Model Dependence, & Emerging Implications for Agentic E-Commerce
  4. Agentic Commerce in 2026: How AI Agents Buy Products | Paz.ai
  5. aisearch.similarweb.com

More in Bots as a New Buyer Class