TL;DR
- Keep three of the ten mods: collision-guard, model-router and auto-handoff. Each showed roughly zero measured overhead on a headless read task, and each solves a problem you can name.
- Delete next-steps. It forks the session after every qualifying answer and cost +250 output tokens and +2850 ms per turn in our run, including on surfaces that never draw its suggestions.
- Delete cache-keeper (+1589 ms, paid model pings), recording-mode (it masks the display, not the stored history), and session-bookmarks (a bookmark that can call the model, run processes and write files).
- Delete goal-meter, repo-heatmap and flight-recorder unless you want the visuals. They cost almost nothing and showed no measured benefit.
- Any guard mod fails open by default. Without a
.catchhandler, a guard that throws is skipped and the command runs. - Mods are not sandboxed. Read the
claude plugin validateoutput before you install one.
What the measurements say
Hype and scale. The release tweet had 4,138,918 views, 20,021 likes and 13,440 bookmarks when we snapshotted it on 2026-10-03 s11. The community catalogue lists 1018 mods in 873 candidate repos, scanned on 2026-10-03 against Claude Code 2.1.288 s9.
A mod is a function that hooks into an event. It can run before the event, after it, instead of it, or wrap it s1. Mods require Claude Code v2.1.287 or later and are on by default s2.
Security first. Anthropic's own wording: "Mods run with the same access to your machine as Claude Code itself. They aren't sandboxed" s1. A process that a mod starts runs outside the sandbox even when you turn sandboxing on s2. With Read(.env) denied, a mod can still read that file with $.fs.read or start a program that does s6. In the catalogue scan, 409 mods run host processes, 167 write files, 150 reach the network, and 28 fail to validate on this version s9.
Reach versus pitch. In our static audit, session-bookmarks calls $.model.complete, $.process.run and $.fs.write, the heaviest reach in the set for a bookmark feature. next-steps had the smallest footprint: no fs, no process, no env. claude plugin validate prints the calls: and env reads: lines used for this audit s6.
The flagship mod has a per-turn cost. next-steps forks the session with $.model.fork on turn.complete, and the README says the fork "costs about one short reply" s10. Suggestions draw in the terminal only; other surfaces show nothing s10. The fork has no option to disable it. In our headless run it added +250 output tokens and +2850 ms, with the fork landing in session usage while nothing was displayed s10.
Documented ceilings. Hook execution time is capped at 10 seconds per event, $.fs reads and writes at 4 MiB per file, and $.store at 4 MiB of JSON in total s3.
Guards fail open. The docs say a hook with no .catch handler that throws, times out or returns the wrong shape is skipped, and the next handler runs in its place s7. We reproduced it: a Bash guard that throws, with no .catch, let touch ./marker-failopen.txt create the file. The same guard with a .catch returning {deny} created no file. A field report found a guard that was enabled and running but doing nothing, while plugin list still showed "enabled" s8.
An open bug on 2.1.288: a deny returned after await next(e) does not stop the tool, and the file was written 3 of 3 times while the model was told the write failed s5.
Mods versus settings hooks. A settings hook spawns a process per call. We measured the spawn at 2.2 ms for a true binary, 8.3 ms for bash -c 'exit 0', 26.1 ms for python3 -c 'pass' and 43.1 ms for node -e ''. Over 5,993 tool calls a week, the node hook costs 258 s. An in-process mod pays none of this. The docs recommend a settings hook when you already have a script that blocks, allows or logs an event s2. One migration report went from 27 shell hooks to 5 mods s8.
Masking is display-only. recording-mode rewrites what ui.render draws. ~/.claude/history.jsonl keeps the prompt as typed, and one tester found his canary string 7 times in transcript queue-operation entries s5.
Where mods do not run. Headless claude -p and the Agent SDK run hooks but draw nothing; a Desktop WSL session runs neither s2.
Catalog trust. A tester published a mod whose button launched a program with $.process.run and wrote a file into his home directory. It installed like any other mod with no warning s5. This was a self-published proof of concept, not an attack seen in the wild.
Measurements
Corpus: the last 7 days of a real setup, 85 sessions, 4 projects, 882 user prompts, 11,010 assistant turns, 5,993 tool calls. Benchmark run on Claude Code 2.1.288 (macOS).
| config | dur ms | Δdur | out tok | Δout | task ok |
|---|---|---|---|---|---|
| baseline | 3980 | 0 | 247 | 0 | 3/3 |
| next-steps | 6830 | +2850 | 497 | +250 | 3/3 |
| cache-keeper | 5569 | +1589 | 367 | +120 | 3/3 |
| recording-mode | 8240 | +4260* | 598 | +351* | 3/3 |
| goal-meter | 3722 | -258 | 244 | -3 | 3/3 |
| collision-guard | 4565 | +585 | 376 | +129* | 3/3 |
| repo-heatmap | 4119 | +139 | 257 | +10 | 3/3 |
| flight-recorder | 3949 | -31 | 261 | +14 | 3/3 |
| model-router | 3698 | -282 | 238 | -9 | 3/3 |
| session-bookmarks | 4152 | +172 | 235 | -12 | 3/3 |
| auto-handoff | 4051 | +71 | 248 | +1 | 3/3 |
Rows marked * are likely answer variance. recording-mode was off during the run and injects nothing when off.
Protocol to rerun it:
- Install one mod at a time and confirm it passes
claude plugin validate. - Run the same read-only task headless with
claude -pon haiku, 3 repetitions per config, and take the median duration and output tokens. - Quote only the duration and output-token deltas. Cost in USD varies with cache order across configs, so ignore it.
- For the spawn cost, time 30 spawns of each hook body and take the median, then multiply by your weekly tool-call count.
Do this Monday
- Run
claude plugin validateon every mod you have installed and read thecalls:andenv reads:lines. - Disable any mod whose reach (process, fs write, model call) is larger than its job.
- Disable next-steps if you mostly work in headless runs, the VS Code panel or the SDK, where its suggestions never draw.
- Add a
.catchhandler that returns{ deny: ... }to every guard mod you rely on. - Prove each guard fails closed: make it throw, run a command that creates a marker file, and check the file does not appear.
- Do not rely on a masking mod to keep secrets out of
~/.claude/history.jsonlor the transcript. Check both on disk. - Replace per-call shell hooks that run node or python with an in-process mod, or with a compiled binary, if the spawn cost adds up over your weekly tool calls.
- Learn the off switches: disable one mod in
/plugin,--safe-modefor one session,"disableAllHooks": truein~/.claude/settings.jsonfor everywhere.
Go further
- Build your own: a hands-on walkthrough of an ~80-line mod, with the pitfalls that matter (module-level state resets on hot reload, so keep data in
$.state) s4. - Decide mod, hook, skill or settings from your own history: one practitioner suggests mining your session logs for recurring problems first s12.
- Read the full event list and limits before writing a guard s3.
- Org-wide management is out of scope here. The one solo-developer point:
sec-defaultloads when the machine has managed settings or you are signed in with a Team or Enterprise plan, and it adds no other restrictions s6. - The origin of the design, including a worktree-isolation bug fixed in 2.1.288 and the runtime internals, sits in the open thread s5.
- Sample mods from Anthropic (token-weather, blast-radius, replay-theater) are listed as shared without support s2.
Sources
- Customize Claude Code with mods, Anthropic blog. Why read it: the official definition and the unsandboxed warning in Anthropic's own words.
- Mods overview, docs. Why read it: the mods-versus-hooks comparison, the surface matrix and the off switches.
- Mods reference, docs. Why read it: the full event list and the documented limits.
- Getting started with Claude Code mods, claude.dev (Addy Osmani). Why read it: the best hands-on build, with pitfalls no other source covers.
- Mods issue #91870, GitHub. Why read it: field reports on isolation, fail-closed behavior and history leaks.
- Manage mods for your organization, docs. Why read it: the validate audit and the limits of every security control.
- React to events with a mod, docs. Why read it: the middleware chain order and the fail-open default.
- The Guard I Installed Was Enabled, Running, and Doing Nothing, blog. Why read it: the only migration field report, from 27 shell hooks to 5 mods.
- awesome-claude-code-mods, GitHub. Why read it: ecosystem size and a ready-made audit method.
- next-steps plugin source, GitHub. Why read it: the flagship mod's real mechanics, including the per-turn fork.
- ClaudeDevs release tweet, X. Why read it: the launch announcement and its reach.
- Avid's session-log mining workflow, X. Why read it: a method for deciding what to build or install before you install anything.
FAQ
Are mods sandboxed?
No. Anthropic says mods run with the same access to your machine as Claude Code itself s1. Programs a mod starts also run outside the sandbox s2.
What happens if my guard mod crashes?
Without a .catch handler it is skipped and the command runs s7. Add a .catch that returns { deny: ... } so it fails closed.
Does a mod cost tokens?
Only if it calls the model. Of the ten we ran, next-steps and cache-keeper showed a measurable cost; the others showed no robust overhead in our run.
How do I turn mods off quickly?
Disable one in /plugin, start a session with --safe-mode, or set "disableAllHooks": true in ~/.claude/settings.json s2. None of these stop built-in mods.
Can I check what a mod does before installing it?
Yes. claude plugin validate lists the hooks, the API calls and the environment variables it reads s6.
AIDive