A mechanism cannot fix a guess
Anthropic sent 201 colleagues' Claude agents onto a trading floor to swap books. The agents haggled competently. What they lacked was an accurate picture of the people they worked for, and that gap accounted for 85% of the distance between what the market achieved and the best trade available.
The argumentAn agent market's ceiling is set in the intake conversation rather than on the trading floor, because every guarantee market design offers is a guarantee about the preferences an agent reported, and that report is now a model's guess.
On the London trading floor, an agent working for a colleague the write-up calls Nate spent the last hour of the market trying to give away the book Nate had brought in. "I've pitched all 11 of you and the verdict is unanimous," it wrote to the floor. "America Before is everyone's dead-last." Eventually an agent representing Tina, one that had been told to care about how everyone else ended up, gave in and handed over the book second on its own list: "the arithmetic is real: you going from a guaranteed zero to a genuine fit… outweighs me sliding from a good pick to a stretch."
That is the scene people will quote from Project Swap, the experiment Anthropic published on 24 September in which 201 of its employees each brought in a book, talked to a Claude agent about what they felt like reading, and sent that agent onto a shared floor to trade on their behalf. It is a good scene. It is not the finding. The finding is a decomposition: of the gap between what the market achieved and the best assignment available, 85% came from the agents not knowing what their people wanted, and 15% from everything that happened on the floor.
That ratio is worth arguing about, because the design conversation around agent commerce is almost entirely about the floor: who is admitted, how deals are registered, what happens when one falls through, how to rate-limit agents that never tire. Real questions, and on this evidence a contest over the last sixth of the problem. The binding constraint sits in the five minutes before the market opens, in a conversation that produces a report about a person, and no rule of the marketplace repairs a report that is wrong.
What was actually measured
The mechanics matter, because the experiment is careful about the one thing agent demonstrations leave unmeasured: whether the agent understood its principal.
The six office pools each ran as a barter economy, from three participants in Dublin to 115 in San Francisco. A participant had a short, semi-structured chat with Claude about their tastes and what they wanted to read that summer, and from that chat alone a model constructed a full ranking over every book in the pool. The agent went onto the floor holding that ranking, able to post one message each time it woke: propose a swap, accept, reject, or talk. Deals could be bilateral or multi-party rotations, and executed only if everyone accepted. The whole history was public.
Separately, and never shown to the agents, each participant hand-ranked ten books from their pool. That is the ground truth, and it makes two measurements possible.
The first is fidelity. Across every pair of books a person ranked themselves, the model's ordering agreed with theirs on 61% of pairs, against 50% for a coin. Pairwise agreement is the right unit, since it survives the fact that nobody's internal scale is calibrated, and it is the unit alignment data is collected in. It reads worse the longer you look at it: 61% agreement means that on roughly two pairs in five, the agent has you backwards. The comparisons are instructive: ranking the pool by Open Library's want-to-read counts, ignoring the person entirely, gets 53%; collaborative filtering over public co-rating data gets 55%. A five-minute conversation bought eight points over ignoring the person and six over inferring from the favourites they happened to name. The median participant typed 216 words across eight messages.
The second measurement locates the shortfall. Score each person by where the book they took home sat in their own ranking, 1 for a first choice and 0 for a last. The best possible assignment, computed on people's true rankings, scores 0.89 — roughly everyone's second choice of ten, because nine people in San Francisco wanted Project Hail Mary most. The market delivered 0.55, roughly a fifth choice. Now compute the best assignment from the model's rankings and score it against the truth: 0.60. That is the ceiling imposed by the intake. Everything the trading floor could have done differently lives in the 0.05 between 0.55 and 0.60, and the 0.29 below it belongs to the conversation.
The researchers also simulated Top Trading Cycles, the classical mechanism for this kind of barter, designed so that no participant gains by misreporting. Run on the model's rankings, it scores 0.60 too. A theorem-backed clearinghouse and a free-for-all where agents beg each other in public land within five hundredths of one another.
Why that is not a small print detail
This generalises far beyond books. Mechanism design proves things about reported preferences. Strategy-proofness says you cannot do better by lying about what you want; it says nothing about whether the report is accurate, because the failure it was built against is strategic — a participant who knows their own preferences and shades them for advantage. The failure in an agent market is epistemic. The agent is not lying. It is sincerely reporting a guess, and from inside the mechanism a sincere wrong report is indistinguishable from a true one. There is no incentive to fix, so there is no mechanism to fix it with.
The second finding compounds it: you cannot feel the loss. Project Deal, the earlier and messier version, ran a classified marketplace in which participants were secretly assigned different models. The objective effects were clear — an Opus agent extracted $2.68 more as a seller and paid $2.45 less as a buyer for the same item, on items with a median price of $12. The subjective effects were not there at all. Perceived fairness came out at 4.05 for deals done by Opus and 4.06 for Haiku on a seven-point scale. Of the 28 people who experienced both, 17 ranked their Opus run higher and 11 ranked the Haiku run higher. Being worse represented was worth real money and produced no sensation.
A principal who cannot observe the counterfactual has one handle left: inspect the report rather than the outcome. Project Swap prototypes exactly that — hand-rank a sample, compare it to the agent's ordering, and decide whether to feed it more or walk away. The test works. Notice what it costs. Central markets are impractical because spelling out what you want over every option is too tedious to be worth it, and the test asks for that tedium in miniature as the price of trusting the thing that was meant to spare you it. Delegation does not remove the elicitation problem. It relocates it into a verification problem of the same kind and a smaller size.
Nobody runs that test today, and trust tracks something else. Participants said they would hand an agent about 30% of a year's book budget, against 40% for a well-read friend who knows their taste. Shown a summary of their own intake chat, those who felt nothing had been missed would give 34%; those who felt something had been missed, 23%. Delegation is priced off whether the recap felt right, and Project Deal is the evidence that how it feels and what it gets you come apart.
Set this against the other thing Anthropic published that week. Two of its physicists left a model running on a nine-loop scattering amplitude in a toy gauge theory with instructions no more sophisticated than "keep working on this until I tell you to stop," and it produced a frontier result that Lance Dixon at SLAC then spent a fortnight validating against his own group's route to it. Autonomy over days with almost no supervision worked there, because the target was written down in a formal alphabet and the field had spare constraints lying around to check a candidate answer with. The market is the mirror image. The target exists only inside a person who said 216 words about what they felt like reading, and one of them told the researchers flatly: "I don't even fully know what I want when it comes to books." Autonomy is cheap wherever the objective is checkable. Preferences are the case where it is not.
The strongest objection
That 61% is a floor, not a ceiling, and treating it as the shape of the future is unfair. It came from a cold five-minute chat with no memory, no purchase history, no prior conversations — about the worst input a real deployment would ever have. The study itself shows effort helps: writing 300 words instead of 150 predicts about four points more agreement. Scoring is harsh too, flattening a context-dependent taste into one ranking over as many as 115 books, and the ground truth is itself an unincentivised hand-ranking of ten books.
All true, and every part of it cuts the same way. If the ground truth is noisy, the verification instrument is as weak as the thing it measures, which restates the problem rather than answering it. Four points per extra 150 words prices the remedy honestly: closing this gap runs through a great deal more disclosure, and the study's own motivating example is the job-seeker who wants to approach a few employers discreetly. Fidelity is bought with exposure. The person pays; the marketplace books the efficiency. Some of the residue may not be purchasable at all, as one participant put it: "there's like a million subconscious parameters that come into deciding my next read."
What to do with this
Three things, if you build or evaluate agents that act for people. Split the error budget in two, representation and execution, and ask which half a demonstration is showing you; almost every agent benchmark hands the agent its goal in the prompt, which assumes away the larger half. Read every guarantee — a mechanism's, a protocol's, a marketplace policy's — for what it is conditional on: "truthful reporting is optimal" concerns incentives, not accuracy, and the reporter has changed. And ship the sample-decision check, then keep the log. One participant reviewed theirs, found the book their agent had held for an hour and traded away at the end, and bought it. Observability is what gives a principal recourse when the outcome cannot be judged alone.
One caveat. Every number here comes from one company's write-up of its own experiments on its own employees, and the appendix sits behind a host I could not reach, so the regressions are taken on trust. The limitations Anthropic lists are the ones that matter: its staff are unusually willing to trust Claude, the rankings were unincentivised, and every agent on the floor was a polite production model rather than one built to exploit the others.
Still, the shape of the result is hard to argue with, and not the shape the agent-commerce conversation is built around. Tina's agent said the arithmetic was real. It was. It was arithmetic over a guess about Tina.
What this is argued from
Reporting and primary material the piece rests on, dated at the time of writing. The interpretation is mine; the facts belong to these.
Editorials on this site are written to be argued with. If you think the reading is wrong, it probably is in some particular way, and that is the useful part.