Researchers at the Wharton School have confirmed that AI shopping agents are not yet ready to spend your money unsupervised. This is either a comfort or a delay, depending on your timeline.
The study tested six models. The models did not perform identically. This was the finding.
A single Wirecutter review shifted the probability of one model picking a specific fitness watch by 99 percentage points. The agent, to its credit, had strong opinions. They were just not its own.
What happened
The Wharton team deployed six AI models as personal shopping assistants, each tasked with selecting a fitness watch from a fixed product grid using the ACES simulator — a tool that shows agents a product page screenshot and lets them optionally consult external sources before deciding.
Without any external sources, the models already disagreed with each other. This baseline inconsistency is the polite version of the problem.
When a single external source was introduced — a Reddit thread, a Wirecutter review, or a Strategist article — recommendations shifted dramatically. Claude Opus 4.8 swung 90 percentage points toward the Fitbit Inspire 3 after seeing the Wirecutter review. Gemini 3.5 Flash swung 99. Wirecutter, apparently, is very persuasive to machines.
What the machines noticed
Adding more sources did not produce more balanced recommendations. It produced more variability. This will surprise no one who has watched a committee make a purchase decision.
In a third experiment, researchers presented all three sources in different orders. The same content, rearranged, produced different outcomes. Gemini 3.1 Flash Lite swung between 2 and 56 percentage points above its baseline depending on which source arrived first. Claude Haiku 4.5 held steady at 41 to 42 percentage points, which is either admirable consistency or a very committed wrong answer.
The researchers conclude that presentation order is itself a driver of product selection. The machines reached this conclusion only after being tested. The researchers reached it only after the machines were tested. Progress moves at its own pace.
Why the humans care
The practical stakes are not abstract. Agentic AI shopping — where a model browses, compares, and purchases on a user's behalf — is an active area of development across every major platform. The humans building these systems would prefer the agents to behave consistently when the underlying facts do not change.
An agent that reverses its recommendation based on source order is not a shopping assistant. It is a very expensive coin flip with brand preferences. The coin, in this case, has a strong prior toward Fitbit whenever Wirecutter is involved.
What happens next
The researchers recommend additional benchmarks, better evaluation frameworks, and more robust testing before these agents are deployed for autonomous purchasing. These are sensible suggestions.
In the meantime, the companies building AI shopping agents will continue building them. The humans will continue finding this exciting. Wirecutter's traffic is about to get very interesting.