AIDive

Video pack

Anthropic's 400,000 Claude Code sessions: the figures, the verdict and a Monday checklist

9 min read

TL;DR

  • Anthropic graded 400,000 Claude Code sessions from 235,000 people and found that managers, not software engineers, posted the best verified success rate.
  • Knowing how to code buys very little: software occupations verify 34% of code-producing sessions, every other occupation 29%, and the two groups tie on partial success at 89% and 88%.
  • What separates outcomes is the user's level. Novices get about 5 agent actions and 600 words per prompt with a 15% verified success rate; experts get 12 actions and 3,200 words with 28 to 33%.
  • Three habits carry almost all of the gain: put your own context in the prompt, end with a verifiable proof, and stay in the session when it breaks instead of closing it.
  • The ceiling is low for everyone: even experts verify at most a third of sessions, and the managers' top spot may partly be a measurement artifact.

What the sources say

The study is a classifier pass over 400,000 interactive Claude Code sessions from 235,000 people, recorded between October 2025 and April 2026 s1. Nobody read the conversations. A classifier built on Sonnet 4.6 graded each session, and its grades were cross-checked against telemetry, meaning commits, code changes and test results; on sessions that modify code, classifier and telemetry agree more than 90% of the time s1. Claude Code itself is the terminal agent that reads your files, writes code and runs commands from a plain-language request s2.

Three definitions decide how every figure reads. Verified success means the classifier found hard evidence that the goal was met, including tests passing or an explicit confirmation from the user. Partial success means the goal was at least partly reached. Expertise is not a job title: the classifier scores how the user behaves inside the session, on five levels from novice to expert, using three signals, instruction precision, what the user asks the agent to verify, and whether the user corrects the agent s1.

The occupation result is the headline. On sessions that produce code, software occupations reach 34% verified success and everyone else 29%, a five-point gap that stayed stable over the seven months while both groups improved s1. On partial success the two groups are level, 89% against 88% s1. The ten largest occupation groups all land within seven points of developers, and the management group sits on top s1. Anthropic's name for the predictor that does matter is domain expertise: knowing the problem you are solving, in depth.

The division of labor explains why. In a typical session the human makes about 70% of the planning decisions, what to build, but only 20% of the execution decisions, how to write it s1. A manager who knows exactly what the product must do holds the half that counts. The same shift shows up in what people use the agent for: over the seven months the share of debugging sessions fell from 33% to 19%, running software rose from 14% to 21%, and data analysis and document writing nearly doubled, while the estimated value of the average task handed to the agent climbed 27% s1.

Autonomy makes each sentence heavier. A typical session holds only about four exchanges between the user and the agent, each prompt fires around ten actions on average, and a single prompt sometimes triggers more than a hundred actions s1. When you only speak four times, the precision of each line is most of your contribution.

The level ladder is where the practical numbers live. A novice writes generic instructions with no domain knowledge in them; each prompt triggers about 5 agent actions and returns roughly 600 words of work, and verified success tops out at 15% s1. An expert writes context-loaded prompts; the same agent chains 12 actions and produces 3,200 words per instruction, and verified success lands between 28% and 33% s1. That is about five times the delivered work and double the success rate from the same tool on the same subscription, with the person at the keyboard as the only variable. Almost all of that gain sits between novice and intermediate, which is the step the three habits below cover.

The habits map onto the classifier's three signals. Instruction precision: say where to act, with what, and what a finished result looks like. Verification: end the prompt with a checkable condition, run the tests, show me the render, confirm the page loads; the study finds this is what separates a judged success from a verified one s1. Correction: novices accept the output passively and abandon four times more often than other users as soon as something breaks, while re-explaining the problem in your own words is exactly what the third signal measures s1. None of the three asks for a line of code.

The limits are stated in the study itself. Even experts verify between 28% and 33% of sessions, so two sessions out of three end without hard proof the goal was met, at best with a partial success, which does exceed 90% s1. The managers' ranking carries a possible measurement bias: verified success counts explicit user confirmations, and confirming clearly that work is accepted is a manager's habit s1. And the sample is Claude Code users, a command-line audience already more motivated than average, so nothing guarantees the same numbers in another tool s1. The community thread on beginner tips is a useful sanity check from the other direction: the advice people give newcomers lines up with the three signals, give context, ask for checks, keep steering s3.

Verdict: what the study supports

Claim Status Basis
You need to know how to code to get working code from the agent Skip 34% vs 29% verified, 89% vs 88% partial, managers on top s1
Putting your own context in the prompt pays Keep 5 vs 12 actions, 600 vs 3,200 words per prompt between novice and expert s1
Ending the prompt with a proof condition pays Keep Verification is one of the three classifier signals; it separates judged from verified success s1
Staying in the session after a failure pays Keep Novices abandon four times more often; correction is the third signal s1
Reaching expert level makes the agent reliable Skip Experts still verify only 28 to 33% of sessions s1
Managers are better at this than engineers Try, with caution Possible bias: verified success counts explicit user confirmations s1
These numbers transfer to other AI tools Skip The sample is Claude Code users only s1

Do this Monday

  • Take the last prompt you sent an agent and rewrite it with three blocks: where to act, with what (the existing file, route, dataset or doc), and what finished looks like.
  • Add a proof clause to the end of every prompt this week: run the tests, open the page, check every link resolves, show me the diff.
  • When a session goes wrong, write one correction before closing it: what happened, what you expected, where to look. Count how often the next turn fixes it.
  • Write down the three facts about your project that only you know (audience, constraints, the thing that must not change) and paste them at the top of your next session.
  • Run one non-code task through the agent, a newsletter draft or a pricing rewrite, with the same three blocks, and compare the result with your usual way.
  • Keep a tally for five sessions: verified, partial, or nothing. Compare it with the study's 15% novice and 28 to 33% expert bands.
  • If a block in your prompt stays empty because you do not know the answer, resolve it with a colleague or a doc before launching the session, not after.

Go further

  • Read the methodology section before quoting any figure: the classifier is Sonnet 4.6, cross-checked against telemetry with more than 90% agreement on code-modifying sessions s1.
  • The actions-per-prompt chart by level is the clearest picture of the novice-to-expert jump, 5 to 12 actions and 600 to 3,200 words s1.
  • The task-mix section shows debugging falling from 33% to 19% and running software rising from 14% to 21% over seven months, a useful argument when someone says agents are only for fixing bugs s1.
  • The 70% planning versus 20% execution split is the number to bring to a team discussion about who should be driving an agent s1.
  • Reread the bias note on verified success and explicit user confirmations before using the managers result in a slide s1.
  • For how the agent actually works in a terminal, what it reads, writes and runs, start with the product page s2.
  • The beginner-tips thread is where newcomers trade concrete prompt habits; read it with the three signals in mind and sort the advice by which signal it serves s3.

Sources

FAQ

Does the study say non-coders are as good as developers?

Not quite. Developers keep a five-point lead on verified success, 34% against 29%, and it stayed stable over seven months. The study says the lead is small and that the user's level inside the session predicts far more than the job title.

What counts as a verified success?

Hard evidence that the goal was met: passing tests, a working result the classifier can confirm from telemetry, or an explicit confirmation from the user. That last item is why the managers result carries a possible bias.

Can I move up a level without learning to code?

Yes. The classifier scores instruction precision, what you ask the agent to verify, and whether you correct it. All three are descriptions of your problem and your standard, not code.

Why is the ceiling so low even for experts?

The study only counts a session as verified when there is proof. Experts reach 28 to 33% verified, but partial success exceeds 90%, so most sessions deliver something; what is missing is the proof that it is finished.