AIDive

Video pack

Taming Opus 5: effort sweep, conciseness slot test and the official guide, measured

11 min read

TL;DR

  • About twenty minutes of settings remove most of the noise people complain about in Opus 5: effort chosen per task type, one conciseness rule in the output style slot, the guide's scope framing in the system prompt, and every "verify your work" line deleted from old prompt files.
  • Effort is not a verbosity dial. It controls how much the model thinks and how many tool calls it makes, not how long the visible answer is. Lowering it to shut the model up pulls the wrong lever.
  • Where a length rule lives matters more than its wording: the built-in Concise preset moved output by about 6 percent, the same rule as a hook or in the instructions file did nothing, and one real rule in the output style slot turned a five-section report into a paragraph plus a file list.
  • Over-engineering is fixed by deleting text, not adding it. Purging verification requests and pasting the guide's scope framing took our reference diff from nine files to three.
  • Low and medium effort found the same two real bugs as the extra-high pass on our review diff, for roughly one fifth of the tokens.
  • What no prompt block fixes: a model that acknowledges an explicit constraint and bypasses it two turns later. We saw it once in a week of sessions, and the guide has no section for it.

What the sources say

The anger is real and measurable. The r/ClaudeCode thread titled "Opus 5 is insufferable" passed 600 upvotes and 178 comments, and its author accuses the model of speaking a new language he calls "Unintelligiblish" s3. On X, a developer posted nothing more than a screenshot of the code comments Opus 5 had generated and collected 9,700 likes s6. When the creator of Claude Code defended the model in public, the reply calling him out gathered 2,843 likes s7.

Three changes under the hood explain a large part of what users feel. Thinking is on by default and can only be disabled at effort high or lower; the context window moves to one million tokens, as default and as maximum; and the effort parameter becomes the central dial, with five levels, low, medium, high, xhigh and max, high being the default s2. The parameter controls how many tokens the model spends thinking, calling tools and answering. At low effort the model batches tool calls, acts without preamble and confirms in one sentence. At high effort it multiplies calls, explains its plan before touching anything and comments its changes in detail s2. If that second description sounds like your sessions, you have been running the default since day one. One API detail to know: at xhigh and max, thinking can no longer be disabled, and the request returns a 400 error if you try s2.

The four behaviors people rant about are reproducible on demand. Verbosity: a two-sentence question came back as sections, subheadings and audit-style warnings; the top comment in the thread describes grandiose announcements of the "we discovered something that changes everything" kind followed by ten minutes of shell commands s3. Over-engineering: one user reports a 7,000-line decisions file, and when he asked for a cleanup the model cut 1,200 lines then added 600 to document the deletions s3. Scope widening: you ask for X, the model decides the real subject is Y and explains why in eight paragraphs. Buried bad news: a wall of text saying everything went well, with an asterisk three quarters down admitting that something broke s3.

The official guide, "Prompting Claude Opus 5", answers the thread point by point. Its most important sentence: effort controls how much the model thinks, not how much it talks; lowering effort reduces thinking volume but does not reliably shorten the visible answer s1. Length has to be asked for in plain words, with a conciseness instruction in the system prompt. The guide also says something few expect from a vendor: remove instructions. If your instructions file contains "verify your work before answering" or "add a final verification step", delete it, because Opus 5 already self-checks and those lines cause over-verification and burned tokens s1. The creator of Claude Code summed it up the same way: Opus 5 needs less prompting, not more s5. The rest of the guide has one section per grievance: agent narration, length of generated files, scope framing, subagents, self-correction, each with the exact prompt block to copy s1.

On review, the guide claims review accuracy holds at low effort levels, which allows a fast cheap pass at commit time and a deep pass later s1. It also warns against "only report serious problems": Opus 5 takes that literally and under-reports, so ask for everything and filter in a second pass s1. On delegation, Opus 5 spawns subagents more readily than its predecessors and each one multiplies cost; the guide offers an instruction reserving delegation for large, genuinely parallel work s1, and Claude Code adds two environment variables since version 2.1.217, CLAUDE_CODE_MAX_SUBAGENT_SPAWN_DEPTH and CLAUDE_CODE_MAX_CONCURRENT_SUBAGENTS, whose defaults are three levels of depth and twenty simultaneous agents s9.

The slot finding comes from a second thread. A r/ClaudeCode user spent days testing where a conciseness rule works: the built-in Concise output style only cut output by about 6 percent, and the same instruction as a hook or as an instructions-file rule changed nothing; what worked was a real instruction in the output style slot s4. The same post gives the criterion for rules that never fire: a rule must name a recognizable moment and a concrete action. "Keep the changelog up to date" does not trigger; "when you modify a file under src/, add a line" does s4.

Measurements

Experiment Setup Result
Effort sweep, same bug fix low, medium, high, xhigh, four clean sessions low and medium produced an equivalent fix for a fraction of high's tokens; xhigh explored more files and armored edge cases
Code review on one of our diffs low pass vs xhigh pass low found the same two real bugs as xhigh for about one fifth of the tokens
Conciseness rule placement Concise preset vs output style slot preset: about 6 percent shorter; output style rule: five-section report became one paragraph plus a file list
Scope framing on the docstring feature guide framing pasted, verification lines removed diff went from nine files touched to three, no parasitic verification step
Constraint bypass one week of sessions one explicit "do not touch this API" constraint acknowledged, then bypassed two turns later

Protocol: one reference bug fix and one small feature from our own repo, replayed in fresh Claude Code sessions. Effort was set per session with /effort, --effort or effortLevel in settings.json s8. The conciseness rule was built from the guide's wording (short, focused answers, reduced caveats, high-level summary unless detail is asked) s1. Cost of the exercise: the four sweep sessions consumed the equivalent of a heavy work day on a 20 dollar plan, and a thread user reports his 20x plan barely lasts a weekend at effort high s3.

Verdict

Setting Keep, try or skip Why
Effort per task type (low or medium daily and reviews, xhigh for big refactors) Keep Same bugs found at one fifth of the tokens on review
Conciseness rule in the output style slot Keep Only slot where the rule moved output beyond about 6 percent
Conciseness rule as hook or instructions-file line Skip No measurable change
Deleting "verify your work" lines Keep Over-verification loop disappeared with them
Guide scope framing in the system prompt Keep Diff from nine files to three
Subagent caps via environment variables Try Defaults of 3 deep and 20 concurrent explain runaway sessions
"Only report serious problems" in review prompts Skip The model under-reports; ask for everything, filter after
Opus 5 on tasks where an ignored constraint is unacceptable Skip for now One bypass in a week, nothing in the guide addresses it

Do this Monday

  • Open your instructions file and delete every line asking the model to verify, double-check or add a final verification step.
  • Set effortLevel in the settings.json of your daily repo to medium, and keep xhigh for one refactor branch to compare.
  • Write one conciseness rule from the guide's wording and put it in the output style slot, not in a hook and not in the instructions file.
  • Paste the guide's scope framing block into your system prompt: deliver what was asked at the intended scope, flag a better approach in one sentence, continue the requested task.
  • Run your next code review twice, once at low and once at xhigh, and count the real bugs each pass finds before you keep paying for the deep one.
  • Rewrite any rule that never fires so it names a moment and an action, following the "when you modify a file under src/" pattern.
  • Set CLAUDE_CODE_MAX_SUBAGENT_SPAWN_DEPTH and CLAUDE_CODE_MAX_CONCURRENT_SUBAGENTS below their defaults for a week and watch your token bill.
  • Keep one hard constraint in a prompt on a sensitive repo and check two turns later whether the model still honors it.

Go further

  • Read the whole "Prompting Claude Opus 5" guide, not just the verbosity section: narration, generated file length, scope, subagents and self-correction each have a copy-ready block s1.
  • The effort page documents the five levels and the 400 error when thinking is disabled at xhigh or max; read it before scripting effort per project s2.
  • The settings reference shows where effortLevel and output styles live so your settings can differ per repo s8.
  • The subagent documentation explains the spawn depth and concurrency caps behind the 3 and 20 defaults s9.
  • The "How I got Opus 5 actually usable" post holds the full slot comparison, including the about 6 percent figure for the Concise preset s4.
  • The "insufferable" thread is worth reading past the top comment: the 7,000-line decisions file story and the constraint-bypass reports are in the long replies s3.
  • The short exchange on X between the Claude Code creator and his critics frames the "less prompting, not more" position in a few lines s5.

Sources

FAQ

Does lowering effort make Opus 5 shorter?

No. Effort reduces thinking volume and tool calls, not the visible answer. Length comes from an explicit conciseness instruction, and the output style slot is where it worked in our tests.

Should I switch models instead?

If your grievance is noise and over-engineering, do the twenty-minute setup first: the gap shows in the first diff. If your grievance is a model that ignores explicit constraints, nothing in the guide fixes it; keep sensitive tasks on a model that obeys and retest at the next update.

Are these settings portable?

No. The output style, the scope framing and the subagent caps live in your config, so every machine and every project has to be set again.

What does the effort sweep cost?

Our four test sessions used the equivalent of a heavy work day on a 20 dollar plan. Run it once on one reference task, then pick a default per repo.