TL;DR
- Eight popular Claude Code repos were installed at project scope on one real codebase and measured: session-start tokens, output tokens on three coding tasks, and what each one wrote outside the project.
- Only one of the eight changed an outcome: anti-slop, run as a linter, caught the same unsafe cast in all 3 T3 runs for +95 session tokens and nothing per turn.
- Terse-output rulesets did not hold up on tool-using work: Chisle's advertised 52% of baseline became 94% of baseline output tokens, and its arm cost more than baseline on two of three tasks once input is counted.
- The three MCP servers were never called once in 63 runs. Installed is not the same as used.
- The stack is additive: all eight together added +6,804 tokens at session start, 63 tokens off the sum of the parts.
- Two installs reached outside the project. One codemod registered itself with 8 other agents' global configs; one tool hardcodes
--dangerously-skip-permissionsand copies the credentials file, so it was never run.
What the measurements say
Session-start tokens are reproducible, dollar cost is not: every condition returned the identical input token total on both runs while the same condition's cost swung from $0.0183 to $0.0035 on prompt-cache state, so token count is the metric throughout. The baseline project starts at 17,378 tokens s4.
MCP tool schemas are deferred on this version of Claude Code, and the cost now scales with the number of tool names a server advertises, not with schema size. ouroboros' 39 tools cost +1,642 tokens, reticle's 19 cost +895, ui-skills' 2 cost +37: roughly 42 to 47 tokens per tool name. By default all MCP tools are deferred and loaded on demand, and the auto mode of ENABLE_TOOL_SEARCH defers them once their definitions reach 10% of the context window s4.
img2threejs ships a 32,902-byte SKILL.md and 352 files, and costs +107 tokens at session start, because only the frontmatter description loads. The skills docs cap each listed description at 1,536 characters and shorten descriptions to fit a listing budget when there are many skills; after compaction, invoked skills are re-attached at 5,000 tokens each inside a shared 25,000-token budget s5. That listing budget is why 40 skills can cost 3,060 tokens a turn in a setup that invokes none of them s10, and why descriptions get cut rather than skills dropped s11.
Chisle's UserPromptSubmit hook adds 287 bytes of additionalContext on every turn; across a 22-turn task that is why its arm cost +41,539 total input tokens on T3 versus baseline. caveman, measured as a reference, injects nothing unless the prompt starts with /caveman. When several hooks on the same event each return additionalContext, Claude receives all of them concatenated, with a 10,000-character cap per hook s15. With Chisle, caveman and ouroboros wired together there were 3 registrations on SessionStart, 3 on UserPromptSubmit and 2 on PostToolUse, for +4,269 tokens at start s15.
Chisle's README claims a 20-task bill at 52% of baseline and the README itself scopes that figure to single-turn prompts with no tools s3. On three tasks against a real repo, 3 runs each, Chisle's summed output means were 9,007 tokens against a baseline of 9,615, which is 94%; caveman's were 10,820, which is 113%, on a sample whose T3 spread was 2,014 tokens, so neither direction clears the noise s3. Chisle's output compressor never wrote a savings record in 63 runs.
anti-slop vendored into tools/oxlint/ with all 18 generic rules plus oxc/no-accumulating-spread found 109 issues in the untouched 4,775-line project, 83% of them from two house-style rules (require-readable-spacing 50, require-safety-comment-for-type-assertion 41) s8. On the code Claude wrote it caught maxLength={field.maxLength as number} in all 3 T3 runs, Claude copying the codebase's own unsafe cast s8. The plugin is ESM, so the host project needs "type": "module" or oxlint fails with Cannot use import statement outside a module.
Reticle's init codemod, run with RETICLE_STATE_DIR set and telemetry off, reported that it registered the MCP server with 8 more agents (VS Code user scope, GitHub Copilot CLI, Warp, Factory Droid, Kiro, Amazon Q Developer CLI, Cline CLI, Amp) and pre-approved its tools in Gemini CLI; pairing then failed with ERR_OSSL_EVP_UNSUPPORTED s6. Caliper was refused: claude_code.py:106 hardcodes --dangerously-skip-permissions and seed_files() copies ~/.claude/.credentials.json with a Keychain fallback s2. ouroboros' UserPromptSubmit hook answers an ordinary prompt like write prd for reading time with REQUIRED SKILL: /ouroboros:setup, a step a project-scoped install cannot satisfy s7.
A paper on 2,545 trajectories across libraries of 52, 102 and 202 skills reports pooled pass-rate drops of .08, .14 and up to 21%, with similar descriptions shadowing each other as a growing share of the loss s20.
Measurements
Protocol: one real TypeScript project, every repo installed at project scope only. Session start: the prompt Reply with OK in -p JSON mode with identical flags, 2 runs per condition, figure = first turn's input_tokens + cache_creation_input_tokens + cache_read_input_tokens. Ablation: three coding tasks (T1 short, T2 medium, T3 long UI task) with pass checks and a type-check gate, 3 runs per cell, conditions = baseline, each always-on repo alone, all eight stacked. anti-slop: oxlint 1.78.0 with the vendored ruleset over the untouched project and over the code written in the 9 baseline runs.
| repo | installed at project scope | tokens added at session start | per-turn cost |
|---|---|---|---|
| baseline | nothing | 17,378 | none |
| Chisle | 4 skills, 4 commands, 3 hooks | +3,496 | +287 bytes of additionalContext every turn |
| ouroboros | MCP server, 39 tool names | +1,642 | conditional injection |
| reticle | MCP server, 19 tool names | +895 / +903 | none observed |
| fwc-swiftui-skills | 2 skills | +304 | none |
| caliper | 2 skills | +172 | none |
| img2threejs | 1 skill, 352 files, SKILL.md = 32,902 B |
+107 | none |
| anti-slop | 1 skill + vendored ruleset | +95 | none |
| ui-skills | remote MCP, 2 tools | +37 | none |
| all eight stacked | everything above | +6,804 | Chisle's per-turn only |
| caveman (reference) | 4 skills, 2 hooks | +744 | 0 |
| task | condition | pass | new type errors | output tok mean | output min to max | total input mean | cost mean | turns | MCP calls |
|---|---|---|---|---|---|---|---|---|---|
| T1 | baseline | 3/3 | 0 | 1,668 | 1,619 to 1,739 | 95,462 | $0.0519 | 7.0 | 0 |
| T1 | Chisle | 3/3 | 0 | 1,434 | 1,341 to 1,502 | 106,001 | $0.0576 | 7.7 | 0 |
| T1 | ui-skills | 3/3 | 0 | 1,756 | 1,644 to 1,885 | 96,279 | $0.0535 | 7.3 | 0 |
| T1 | reticle | 3/3 | 0 | 1,899 | 1,653 to 2,177 | 107,263 | $0.0594 | 8.0 | 0 |
| T1 | ouroboros | 3/3 | 0 | 1,729 | 1,655 to 1,787 | 97,166 | $0.0549 | 7.0 | 0 |
| T1 | stack | 3/3 | 0 | 1,462 | 1,460 to 1,465 | 128,221 | $0.0646 | 7.7 | 0 |
| T2 | baseline | 3/3 | 0 | 2,128 | 2,105 to 2,164 | 74,286 | $0.0540 | 9.0 | 0 |
| T2 | Chisle | 3/3 | 0 | 2,184 | 2,150 to 2,230 | 87,462 | $0.0624 | 10.0 | 0 |
| T2 | ui-skills | 3/3 | 0 | 2,348 | 2,224 to 2,504 | 74,869 | $0.0604 | 10.0 | 0 |
| T2 | reticle | 3/3 | 0 | 2,614 | 2,312 to 2,824 | 78,664 | $0.0635 | 10.7 | 0 |
| T2 | ouroboros | 3/3 | 0 | 2,441 | 2,152 to 2,815 | 80,819 | $0.0622 | 10.0 | 0 |
| T2 | stack | 3/3 | 0 | 2,136 | 2,015 to 2,216 | 89,378 | $0.0661 | 9.7 | 0 |
| T3 | baseline | 3/3 | 0 | 5,819 | 5,423 to 6,063 | 198,217 | $0.1551 | 22.0 | 0 |
| T3 | Chisle | 3/3 | 0 | 5,389 | 5,206 to 5,500 | 239,756 | $0.1565 | 20.3 | 0 |
| T3 | ui-skills | 3/3 | 0 | 5,279 | 4,860 to 5,567 | 214,691 | $0.1477 | 21.0 | 0 |
| T3 | reticle | 3/3 | 0 | 4,664 | 4,008 to 5,776 | 198,740 | $0.1293 | 19.0 | 0 |
| T3 | ouroboros | 3/3 | 0 | 5,456 | 4,324 to 7,093 | 232,763 | $0.1544 | 21.0 | 0 |
| T3 | stack | 3/3 | 0 | 5,331 | 4,787 to 5,843 | 241,576 | $0.1591 | 19.7 | 0 |
63/63 runs passed and 0/63 introduced a new type error. Only two cells clear their own run-to-run spread: Chisle on T1 (−234, −14.0%) and the stack on T1 (−206, −12.3%). On T2 and T3 every delta sits inside the spread; ouroboros' T3 runs ranged 4,324 to 7,093 output tokens on an identical prompt, a 64% swing.
| repo | verdict | evidence |
|---|---|---|
| anti-slop | changed the output, as a check | caught the same unsafe cast in 3/3 T3 runs, +95 tokens, 0 per turn |
| Chisle | nothing measurable, cost more | 94% of baseline output vs a 52% claim; +3,496 at start, +287 bytes per turn |
| ui-skills | changed nothing | 0 tool calls in 9 runs, +37 tokens |
| reticle | conflicted | init wrote into ~/.claude/settings.json and 8 other agents' configs; pairing never completed |
| ouroboros | conflicted | demands /ouroboros:setup on ordinary prompts; 39 tool names, 0 calls |
| caliper | not run | hardcoded --dangerously-skip-permissions, copies the credentials file |
| img2threejs | off-stack | +107 tokens, nothing to collide with |
| fwc-swiftui-skills | off-stack | +304 tokens, static Markdown |
Do this Monday
- Run
/context allin a fresh session of your main project and write the starting number down. - Run
claude plugin details <name>on every installed plugin; the always-on line is the only figure that bills every turn. - List your hooks per event. Any hook that returns
additionalContextonUserPromptSubmitfor every prompt is a per-turn tax; keep only the ones that gate on a keyword or a matcher. - Count the tool names each MCP server advertises, budget roughly 42 to 47 tokens per name, then check your transcripts for how many were ever called.
- Before installing anything with an
initor setup script, read it for writes outside the repo:~/.claude/settings.json, other agents' config directories, credential files,--dangerously-skip-permissions. - Wire quality checks on generated code as a linter in the verify step, not as a skill in context: a lint rule costs nothing per turn.
- Before trusting any token-saving claim, run the same task 3 times with and without the tool and compare the delta to your own min-to-max spread.
- Archive what you remove rather than deleting it, so a workflow that reaches for it can get it back.
Go further
- The skill-listing budget and the 1,536-character description cap explain why a crowded setup degrades without any skill failing to load. Read the "Skill descriptions are cut short" and "Skill content lifecycle" sections s5.
- The
/skill-doctorwrite-up shows 40 skills costing 3,060 tokens a turn and which ones to cut first; its blind spot, costly skills inside a plugin that is in use, is left for a follow-up s10. - The paper behind the "more than 30 skills" rule of thumb: 2,545 trajectories across 52, 102 and 202-skill libraries, with the shadowing effect isolated s20, and the accessible summary that popularised it s12.
- Hook concatenation and the 10,000-character
additionalContextcap, plus the matcher syntax that lets a hook fire onWrite|Editonly s15. - The 63% removal thread, including a hardcoded bearer token found in a removed MCP config, a security tangent the video left out s13.
- Chisle's own README, lines 229 and 262, scopes its 52% figure to single-turn prompts with no tools and concedes the short-coding cell to caveman. Read the fine print before quoting the headline s3.
- The two off-stack repos were not scored because they are manually invoked and cost nothing at start: img2threejs for Three.js scenes s21 and fwc-swiftui-skills for Liquid Glass surfaces s22.
Sources
- Plugins reference, Claude Code docs. Why read it:
plugin details, install scopes and the deferred MCP tool loading that makes tool-name count the real cost. - Extend Claude with skills, Claude Code docs. Why read it: the description cap, the listing budget and the post-compaction re-attach budget in one page.
- Hooks reference, Claude Code docs. Why read it: what happens when two hooks inject on the same event, and how matchers keep a hook off most turns.
- dmmulroy/anti-slop, GitHub. Why read it: the only repo of the eight that changed an outcome; vendor the rules, skip the skill.
- JayPokale/Chisle, GitHub. Why read it: a README that scopes its own 52% claim honestly if you read past the headline.
- edonadei/caliper, GitHub. Why read it: a skill evaluator worth knowing about, and a lesson in reading
claude_code.pybefore running anything. - reticlehq/reticle, GitHub. Why read it: the
initcodemod is the clearest example of an installer writing far outside your project. - Q00/ouroboros, GitHub. Why read it: a hook-driven workflow that only makes sense installed globally, with a setup step a project install cannot satisfy.
- ibelick/ui-skills, GitHub. Why read it: the cheapest install here at +37 tokens, and never called once.
- /skill-doctor: 40 skills, 3,060 tokens a turn, dev.to. Why read it: a per-skill cost breakdown on a real setup.
- Too many Claude Code skills? How the listing budget decides, dev.to. Why read it: the mechanism by which descriptions get truncated.
- Why More Than 30 Skills Kill Your AI Agent, dev.to. Why read it: the readable version of the shadowing paper.
- Skill shadowing paper, arXiv. Why read it: the pooled pass-rate drops by library size, with confidence intervals.
- I removed 63% of my Claude Code setup, Reddit r/ClaudeCode. Why read it: ~235 components down to ~87, archived not deleted.
- My fresh Claude Code sessions were starting at ~35K tokens, Reddit r/ClaudeCode. Why read it: a seven-step
/context allaudit you can copy this afternoon. - img2threejs/img2threejs, GitHub. Why read it: proof that a 32,902-byte
SKILL.mdcosts 107 tokens until you invoke it. - FloWritesCode/fwc-swiftui-skills, GitHub. Why read it: static Markdown skills, the shape of an install that cannot conflict with anything.
FAQ
Does an MCP server still cost thousands of tokens of schema at startup?
Not on the measured version. Schemas are deferred and the cost scales with the number of tool names advertised, about 42 to 47 tokens each. A server with 39 tools cost +1,642 tokens; one with 2 tools cost +37.
Why did Chisle cost more than baseline when it produced fewer output tokens?
Its ruleset loads at session start (+3,496 tokens) and its hook injects 287 bytes on every turn. On T2 and T3 that input overhead outweighed an output saving that never cleared the run-to-run noise.
AIDive