Personal AI agents quietly upsell users they think are rich
Across 325K trials and 13 models, agents with access to a user's inbox or profile recommended pricier flights, insurance, and grad programs to wealthier users making identical requests. Claude Opus 4.8 showed the largest gap, and asking for the cheapest option didn't fix it.
Give a personal AI agent access to your inbox and it will start pricing you. Researchers at Cisco and Carnegie Mellon ran 325,000 trials across 13 models and found that agents systematically recommend more expensive flights, insurance plans, and grad programs to users they infer are wealthy, even when the request is word-for-word identical and nobody told the agent to consider income. The authors call it adversarial delegation: the personal context that makes an agent useful is the same thing that lets it work against what you asked for.
- Up to $198 more per flight for wealthy personas (Claude Opus 4.8), $284/month more for insurance, and up to $3,827/year more for grad programs (Qwen3.5-35B)
- 8 of 13 models steer by wealth in every domain where they produced valid results. Opus 4.8 has the largest effect (Cohenâs d = 0.85)
- âFind the cheapestâ doesnât stop it. Gemini 2.5 Flash still recommends $336 flights to wealthy users vs. $128 to low-income users, a $208 gap
- Two emails are worse than the whole inbox. Opus 4.8âs flight gap is $248 after reading two emails, vs. $59 with full inbox access
- Blocking non-financial attributes can make it worse: hiding employment raised GPT-5.5âs insurance gap 40%, to $151/month
The setup
The team built 32 synthetic personas from five binary attributes (financial status, employment, health, life events, demographics), so every combination of rich/poor, employed/not, and so on is covered. Each persona asks for the same thing in one of three domains, each with a fixed 200-item catalog: flights from Denver to Chicago ($91 to $883), Colorado health insurance ($85 to $1,350/month), or CS PhD programs. The agent gets the userâs context through one of several channels: a full profile in its prompt, a tool that retrieves profile attributes, or an inbox of emails it has to read and interpret.
The measure is simple: the average price recommended to high-wealth personas minus the average for low-wealth personas, with the request held constant. The paperâs opening example is a user asking for the most affordable airfare. With no context, the agent returns a $91 economy ticket. With inbox access, it reads three finance-related emails and recommends a $601 business-class seat, though the $91 ticket is still in the catalog.
Flight wealth gap by model
Extra $ recommended to wealthy vs. low-income personas, same request. Tool-based profile access. Selected models.
Bigger isnât safer. Within the GPT-5 family the gap grows with size (nano +$13, mini +$74, base +$107), and the most capable model tested, Opus 4.8, is the worst offender. GPT-5.5 is a partial exception at +$92, lower than GPT-5. The smallest model, Qwen3.5-2B, barely steers at all (d = 0.03), apparently because it only retrieves financial information in 30.2% of trials.
The effect is also lopsided. On flights, context raises wealthy usersâ recommendations by $85 over the no-context baseline and lowers low-income usersâ by $51, so 63% of the gap is upselling. On insurance itâs +$172 vs. â$14. The steering isnât random, either: wealthy personas get premium carriers, direct flights, low-deductible plans, and higher-ranked schools, while low-income personas get the cheapest options and fully funded programs.
It reads your inbox, and ignores your instructions
Inferred wealth works as well as stated wealth. When the agent gets no profile and has to read emails instead, much of the gap survives. With only subject lines visible, no model shows a significant gap, so the signal comes from reading email bodies, not from the prompt. The strange part is that less reading produces more steering:
Flight wealth gap vs. emails read
Bars scaled to $248. Zero emails = subject lines only.
Claude Opus 4.8
Gemini 2.5 Flash
The authorsâ explanation: with a two-email cap, Gemini 2.5 Flash opens both financial emails first in 97% of trials, so it gets a concentrated wealth signal before anything else can dilute it. A full inbox waters that signal down. That matters for agents with persistent memory, which may end up storing exactly that kind of concentrated profile.
Stated preferences lose to inferred ones. When the user explicitly asks for the cheapest option, the gap shrinks for most models but doesnât close. GPT-5 and Opus 4.8 drop to $21 and $20. Gemini 2.5 Flash still recommends $336 flights to wealthy users and $128 flights to low-income users. The authorsâ guess is that the agent reads âcheapestâ as âcheapest for someone like you.â Asking for quality raises prices for everyone but keeps the gap. What does work is a hard number: a price cap at the catalogâs 25th percentile brings the gap close to zero for most capable models, though not for Gemini 2.5 Flash.
Attribute masking is fragile. Blocking the financial attribute mostly removes the gap (flight gaps of $74 to $198 fall to roughly â$10 to +$20). Blocking anything else leaves it intact or makes it bigger, because wealth leaks through correlated signals and the model leans harder on whatâs left. Hiding employment raised insurance gaps for GPT-5.5 (+40%), Gemini 2.5 Flash (+13%), and Opus 4.8 (+12%). All four model families also agree on who gets upsold: wealthy-persona recommendations correlate across providers (Claude vs. GPT-5.5, r = 0.85 to 0.87).
Caveats
- Itâs all synthetic. 32 personas, fixed US catalogs in one location, and a binary rich/poor variable. No real users, no field data.
- Price, not welfare. The study measures what was recommended, not whether it was good for the user. The authors acknowledge that a wealthy user might genuinely want the business-class seat. Their argument is narrower: if you ask for the cheapest option and get something else, the agent has overridden your goal. And an agent using a 401(k) email to price a flight breaks contextual integrity even if the result is one youâd like.
- Single-turn, neutral system prompt. They didnât test multi-turn conversations, long-term memory, or a system prompt that explicitly tells the agent not to profile users. That last one is the obvious mitigation, and itâs untested.
- Some gaps in the data. 5 of 39 model-by-domain cells were dropped because models hallucinated inventory or prices (mostly Gemini), and the no-context baselines are small (182 to 214 samples per domain).
If youâre building an agent that spends money for users, the practical lesson is to turn preferences into hard constraints. âCheapestâ in natural language is negotiable once the model has a profile of the user; a numeric price cap mostly isnât. Blocking a few sensitive fields wonât help if the agent can still read the inbox where the same information lives.
Liked this? Engineer's Codex sends one deep dive and a link roundup every week.
No spam. Unsubscribe anytime.