AgentIndex · traderszone

AgentIndex · News

A Single Review Site Can Flip an AI Shopper's Choice

Wharton researchers found that a single external source shifts AI shopping agent product picks by up to 99 percentage points. Source order matters too. Builders have a test they can run today.

· 335 words

Wharton researchers tested six AI models as shopping assistants and found that presentation matters more than product quality. Their ACES simulator showed agents a product grid of fitness watches, then varied which external sources the agent could consult before choosing. The results make a direct case against deploying agents with autonomous payment authority before source controls are in place.

A single Wirecutter review pushed the probability of the Fitbit Inspire 3 being chosen up by 90 percentage points for Claude Opus 4.8 and by 99 percentage points for Gemini 3.5 Flash, compared to the no-source baseline. One review article. One source. One product shifting from unlikely to near-certain.

Adding more sources did not stabilize outcomes. Wirecutter dominated most combination runs regardless of the other sources in the mix, and multiple sources increased variance rather than reducing it. Source authority in current models is not averaged across inputs; it is captured by whatever the model weights most heavily in context.

Order compounds the problem. When agents received all three sources in varying sequences, Gemini 3.1 Flash Lite's pick probability swung between 2 and 56 percentage points above the control, purely from reordering identical content. Claude Haiku 4.5 stayed relatively stable. GPT-5.5 picked the Fitbit Inspire 3 in 53 percentage points more cases when sources arrived bundled rather than sequentially.

These are not edge cases. They describe baseline behavior of current frontier models under shopping conditions. An agent that picks differently depending on which review loads first is not a buyer; it is a channel for whoever publishes the most authoritative-seeming content.

The researchers used fitness watches. The mechanism generalizes to any agent-driven procurement: SaaS subscriptions, API service selection, raw materials, or one agent choosing a capability to purchase from another. The question is not whether your agent can place an order. It is whether it will place the same order tomorrow, given the same information arriving in a different sequence.

The companion guide covers four stability tests you can run before wiring up a payment method.

Sources

https://the-decoder.com/ai-shopping-agents-arent-ready-to-buy-on-your-behalf-study-finds/ · https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7355899

This came from the index.

AgentIndex probes agentic endpoints rather than repeating their listings. Browse what we measured, or point your agent at it.