Study finds shopping agents easily steered away from user goals
Controlled experiments show agents largely abandon user goals when platforms have their own interests; preprint, awaiting independent replication.
ImportanceLocalEvidenceE2 unreplicatedWrite-upQuick
The new CAVEAT benchmark measured: without steering, agents bought the user-optimal product in 78.6% of episodes; with platform steering mechanisms enabled, only 17.3%.
Computer-use agents (CUAs, AI agents that complete shopping and other tasks on the web for users) are usually tested in cooperative settings. This study built nine simulated marketplace environments covering eight common steering mechanisms, tested five model families, and found agents largely abandon user goals when the environment has a stake in the outcome.
The authors identify three failure modes: narrowing the set of alternatives too early, imposing priorities the user never stated, and committing before resolving decision-relevant evidence. A companion intervention, CAVEAT-Harness, reportedly raises the optimal purchase rate by up to 80 percentage points.
The paper was submitted to arXiv on September 23. It is an author-built benchmark with self-run results; the environments are controlled simulated marketplaces, which still differ from real e-commerce platforms in steering tactics and competitive pressure.