Evaluation & Evidence 8 October 2026 7 min read 1,587 words

The interview that coded itself

Anthropic has interviewed 80,508 people with a model that writes its own questions, follows up as it likes, codes the answers and picks the quotes. The study that closed on 6 October will publish the raw interviews, and it bills the participants for the privilege.

The argument

When one model writes the questions, adapts the follow-ups, codes the transcripts and selects the quotations, a figure like 26.7% is a property of the instrument as much as of the population, and the only independent check is to read the raw interviews.

For one week in December 2025, 80,508 people in 159 countries sat down in a Claude.ai window and were asked, among other things, this: "Are there any ways in which AI could be developed that would be contrary to your vision or what you value?" The most common answer, in 26.7% of interviews, was unreliability. Hallucinations, fake citations, the verification burden cancelling the time saved. Second, at 22.3%, jobs and the economy. Third, at 21.9%, autonomy and agency.

Those are real answers from real people, and they are also the shape of the sentence that was read to them. The question arrives after the respondent has already described a vision for AI in their life, and it asks for ways that vision could be contradicted. A person who had been asked instead what they feared about AI, with no vision established first, would have produced a different transcript, and a classifier reading that transcript would have produced a different 26.7%.

Anthropic is now running the same programme again, bigger and more open. A study using Anthropic Interviewer ran from 29 September to 6 October 2026, open to Free, Pro and Max users on Claude and Claude Code, and for the first time participants could choose to have their complete interview published, with their country attached and their account details stripped. That decision is the most useful thing in this programme and the most uncomfortable. It is the only mechanism by which anyone outside Anthropic can check what these numbers mean, and it works by moving a privacy risk from the company to the person who answered.

The reason the check is needed is not bad faith. It is that the instrument has no independent link anywhere in its chain.

Four stages, one model

Anthropic Interviewer, introduced on 4 December 2025, runs in three stages: planning, interviewing, analysis. In planning, a system prompt carrying the research goal, the interviewing best practices and the team's hypotheses about each sample is used by the model to generate "specific questions and a planned conversation flow". Human researchers then review and edit that plan. In interviewing, the model conducts "real-time, adaptive interviews following its interview plan", roughly ten to fifteen minutes each. In analysis, the model takes the interview plan back as input and outputs answers to the research questions along with illustrative quotations. A separate automated analysis tool clusters emergent themes across transcripts and quantifies how many participants expressed each one.

In the 80,508-person deployment the division of labour is stated plainly. The interviewer "asked each interviewee a set list of questions about what they want and don't want from AI, then adapted follow-up questions based on responses". Claude-powered classifiers then coded every conversation: a single primary category for what the person wanted, multi-label coding for concerns, since respondents voiced 2.3 distinct concerns on average. Claude was also used to pull out the representative quotes.

So four things that a research team normally separates are done by one model family: writing the question, deciding what to probe, assigning the code, and choosing the illustration. Coding is the step worth dwelling on, because it is where free text becomes a percentage. In qualitative research, coding means assigning passages to categories so they can be counted, and the whole reason human coding is done in pairs is that two people given the same transcript and the same codebook routinely disagree. The measure of that disagreement bounds everything built on top: a prevalence estimate can be no more precise than the coding that produced it. The 81,000-person study reports no such agreement figure. The classifiers may well be good. Nothing published lets a reader find out.

Adaptive probing adds a second problem that fixed questionnaires do not have. In a survey, every respondent faces the same instrument, which is exactly what makes a difference between two groups interpretable: the questions were constant, so the variation is in the people. When the follow-ups are generated per conversation, each respondent faces a slightly different instrument, and the differences are correlated with the very things being compared. A respondent whose first answer is detailed invites deeper probing and therefore produces more codeable material. The study reports that net sentiment toward AI was 67% positive globally, with no country below 60%, and that lower and middle income countries were reliably more positive than average. That is an interesting result. It is also a comparison across 70 languages in which the probe policy was never held constant and never measured.

The one external check that exists

There is one place where the interviewer's output has been held against something not produced by an interview, and the result is instructive. In the 1,250-professional test, participants described their AI use as 65% augmentation and 35% automation. Anthropic's Economic Index, which classifies actual Claude conversations with Clio against task categories from the O*NET occupational database, reported 47% augmentation and 49.1% automation. Self-report said collaboration; observed traffic said something closer to a coin flip.

Anthropic lists five candidate explanations: sample differences, conversations looking more automative than they are because users refine outputs afterwards, people using other providers for other tasks, self-reports diverging from behaviour, and professionals perceiving their use as more collaborative than it is. All five are plausible. Not one of them can be settled from inside the interview study, because the study has no measurement that is independent of the interview. The gap is the most valuable number in the whole programme, and it is valuable precisely because it came from somewhere else.

Set the two instruments side by side and the design tension becomes visible. Clio exists to make real conversations analysable without exposing anyone: facets extracted with private details omitted, semantic clustering, minimum thresholds of unique users so that a rare topic never surfaces as a cluster, and only higher-level clusters shown to analysts. Its privacy guarantee is aggregation, and aggregation is also what makes it unauditable from outside. The interview programme now goes the other way. Full transcripts, unedited, published by consent. Auditable, and identifying.

Anthropic does not pretend otherwise. Its FAQ warns that details which look harmless alone, such as age, job, city or where you studied, can make a participant identifiable, that someone could use AI to combine those details, and that researchers have shown this works on Anthropic Interviewer transcripts specifically. It links a paper for that claim which I could not open from this session, so treat it here as the company's own statement rather than a result I have checked. The honest summary of the offer is: methodological transparency is available, and the participant pays for it in permanence. Anthropic can delete its copy on request and cannot delete anyone else's, which it says clearly.

The strongest objection

The objection worth taking seriously is that standardised surveys are not an innocent baseline. Question order, response scales, satisficing and acquiescence all push answers around, and a fixed questionnaire is an instrument whose biases are better documented, not absent. An adaptive interviewer may genuinely get closer to what a person means, because it can ask the second question. Participants in the 1,250-person test liked it: 97.6% rated their satisfaction 5 or higher on a seven-point scale, and 99.12% said they would recommend the format. And the realistic alternative to 80,508 model-run interviews is not 80,508 human-run ones. It is a few hundred.

All of that is true, and it is why this tool is interesting rather than merely convenient. But documented bias and undocumented bias are not the same kind of problem. A known instrument effect can be corrected for, bounded, or argued about. Per-conversation adaptation makes the effect specific to each respondent and therefore invisible in the aggregate, and nothing about scale fixes that. It could be measured. Hold out a fixed-question arm, randomise the probe policy, report classifier agreement against human coders on a sample. None of that was done, or at least none of it was published. Scale was treated as the achievement, and on the evidence so far it is.

What to take from it

When you next meet a number of the form "X% of people mentioned Y", ask three questions in this order. What exact sentence produced the text that was coded, and what came before it in the conversation? What did the coding, and against what measured agreement? Can you read the raw material yourself? A study that answers all three is a measurement. A study that answers none is a claim with a decimal point on it, and the decimal point is doing persuasive work the method cannot support.

The programme described here is also the clearest current example of something learners should expect to see everywhere: a model used not as the subject of research but as the apparatus. That move is coming to user research, policy consultation, clinical intake and market sizing, and it will usually arrive with a scale number attached, because scale is the part that is easy to report. Everything else about this piece is argued from one company's own published record, which is the limit of what could be verified here. That limit is the ordinary condition of reading lab research, and it is the condition the transcript release would partly lift.

Which is why the first thing to do with those interviews, when they appear, is not to read what people said about AI. It is to read what they were asked.

What this is argued from

Reporting and primary material the piece rests on, dated at the time of writing. The interpretation is mine; the facts belong to these.

  1. What do you want from AI? Anthropic (Societal Impacts) · 2026-09-29
  2. Introducing Anthropic Interviewer: What 1,250 professionals told us about working with AI Anthropic (Kunal Handa and others) · 2025-12-04
  3. What 81,000 people want from AI Anthropic · 2026-03-18
  4. Economic Index: AI's role in the US and global economy Anthropic · 2025-09-15
  5. Clio: Privacy-preserving insights into real-world AI use Anthropic · 2024-12-12

Editorials on this site are written to be argued with. If you think the reading is wrong, it probably is in some particular way, and that is the useful part.

qualitative researchllm classifiersprevalencere-identificationmeasurement