AIDive

Video pack

Claude Code mods: sourced verdict, measurements and Monday checklist

9 min read

TL;DR

  • Keep three of the ten mods: collision-guard, model-router and auto-handoff. Each showed roughly zero measured overhead on a headless read task, and each solves a problem you can name.
  • Delete next-steps. It forks the session after every qualifying answer and cost +250 output tokens and +2850 ms per turn in our run, including on surfaces that never draw its suggestions.
  • Delete cache-keeper (+1589 ms, paid model pings), recording-mode (it masks the display, not the stored history), and session-bookmarks (a bookmark that can call the model, run processes and write files).
  • Delete goal-meter, repo-heatmap and flight-recorder unless you want the visuals. They cost almost nothing and showed no measured benefit.
  • Any guard mod fails open by default. Without a .catch handler, a guard that throws is skipped and the command runs.
  • Mods are not sandboxed. Read the claude plugin validate output before you install one.

What the measurements say

Hype and scale. The release tweet had 4,138,918 views, 20,021 likes and 13,440 bookmarks when we snapshotted it on 2026-10-03 s11. The community catalogue lists 1018 mods in 873 candidate repos, scanned on 2026-10-03 against Claude Code 2.1.288 s9.

A mod is a function that hooks into an event. It can run before the event, after it, instead of it, or wrap it s1. Mods require Claude Code v2.1.287 or later and are on by default s2.

Security first. Anthropic's own wording: "Mods run with the same access to your machine as Claude Code itself. They aren't sandboxed" s1. A process that a mod starts runs outside the sandbox even when you turn sandboxing on s2. With Read(.env) denied, a mod can still read that file with $.fs.read or start a program that does s6. In the catalogue scan, 409 mods run host processes, 167 write files, 150 reach the network, and 28 fail to validate on this version s9.

Reach versus pitch. In our static audit, session-bookmarks calls $.model.complete, $.process.run and $.fs.write, the heaviest reach in the set for a bookmark feature. next-steps had the smallest footprint: no fs, no process, no env. claude plugin validate prints the calls: and env reads: lines used for this audit s6.

The flagship mod has a per-turn cost. next-steps forks the session with $.model.fork on turn.complete, and the README says the fork "costs about one short reply" s10. Suggestions draw in the terminal only; other surfaces show nothing s10. The fork has no option to disable it. In our headless run it added +250 output tokens and +2850 ms, with the fork landing in session usage while nothing was displayed s10.

Documented ceilings. Hook execution time is capped at 10 seconds per event, $.fs reads and writes at 4 MiB per file, and $.store at 4 MiB of JSON in total s3.

Guards fail open. The docs say a hook with no .catch handler that throws, times out or returns the wrong shape is skipped, and the next handler runs in its place s7. We reproduced it: a Bash guard that throws, with no .catch, let touch ./marker-failopen.txt create the file. The same guard with a .catch returning {deny} created no file. A field report found a guard that was enabled and running but doing nothing, while plugin list still showed "enabled" s8.

An open bug on 2.1.288: a deny returned after await next(e) does not stop the tool, and the file was written 3 of 3 times while the model was told the write failed s5.

Mods versus settings hooks. A settings hook spawns a process per call. We measured the spawn at 2.2 ms for a true binary, 8.3 ms for bash -c 'exit 0', 26.1 ms for python3 -c 'pass' and 43.1 ms for node -e ''. Over 5,993 tool calls a week, the node hook costs 258 s. An in-process mod pays none of this. The docs recommend a settings hook when you already have a script that blocks, allows or logs an event s2. One migration report went from 27 shell hooks to 5 mods s8.

Masking is display-only. recording-mode rewrites what ui.render draws. ~/.claude/history.jsonl keeps the prompt as typed, and one tester found his canary string 7 times in transcript queue-operation entries s5.

Where mods do not run. Headless claude -p and the Agent SDK run hooks but draw nothing; a Desktop WSL session runs neither s2.

Catalog trust. A tester published a mod whose button launched a program with $.process.run and wrote a file into his home directory. It installed like any other mod with no warning s5. This was a self-published proof of concept, not an attack seen in the wild.

Measurements

Corpus: the last 7 days of a real setup, 85 sessions, 4 projects, 882 user prompts, 11,010 assistant turns, 5,993 tool calls. Benchmark run on Claude Code 2.1.288 (macOS).

config dur ms Δdur out tok Δout task ok
baseline 3980 0 247 0 3/3
next-steps 6830 +2850 497 +250 3/3
cache-keeper 5569 +1589 367 +120 3/3
recording-mode 8240 +4260* 598 +351* 3/3
goal-meter 3722 -258 244 -3 3/3
collision-guard 4565 +585 376 +129* 3/3
repo-heatmap 4119 +139 257 +10 3/3
flight-recorder 3949 -31 261 +14 3/3
model-router 3698 -282 238 -9 3/3
session-bookmarks 4152 +172 235 -12 3/3
auto-handoff 4051 +71 248 +1 3/3

Rows marked * are likely answer variance. recording-mode was off during the run and injects nothing when off.

Protocol to rerun it:

  1. Install one mod at a time and confirm it passes claude plugin validate.
  2. Run the same read-only task headless with claude -p on haiku, 3 repetitions per config, and take the median duration and output tokens.
  3. Quote only the duration and output-token deltas. Cost in USD varies with cache order across configs, so ignore it.
  4. For the spawn cost, time 30 spawns of each hook body and take the median, then multiply by your weekly tool-call count.

Do this Monday

  • Run claude plugin validate on every mod you have installed and read the calls: and env reads: lines.
  • Disable any mod whose reach (process, fs write, model call) is larger than its job.
  • Disable next-steps if you mostly work in headless runs, the VS Code panel or the SDK, where its suggestions never draw.
  • Add a .catch handler that returns { deny: ... } to every guard mod you rely on.
  • Prove each guard fails closed: make it throw, run a command that creates a marker file, and check the file does not appear.
  • Do not rely on a masking mod to keep secrets out of ~/.claude/history.jsonl or the transcript. Check both on disk.
  • Replace per-call shell hooks that run node or python with an in-process mod, or with a compiled binary, if the spawn cost adds up over your weekly tool calls.
  • Learn the off switches: disable one mod in /plugin, --safe-mode for one session, "disableAllHooks": true in ~/.claude/settings.json for everywhere.

Go further

  • Build your own: a hands-on walkthrough of an ~80-line mod, with the pitfalls that matter (module-level state resets on hot reload, so keep data in $.state) s4.
  • Decide mod, hook, skill or settings from your own history: one practitioner suggests mining your session logs for recurring problems first s12.
  • Read the full event list and limits before writing a guard s3.
  • Org-wide management is out of scope here. The one solo-developer point: sec-default loads when the machine has managed settings or you are signed in with a Team or Enterprise plan, and it adds no other restrictions s6.
  • The origin of the design, including a worktree-isolation bug fixed in 2.1.288 and the runtime internals, sits in the open thread s5.
  • Sample mods from Anthropic (token-weather, blast-radius, replay-theater) are listed as shared without support s2.

Sources

FAQ

Are mods sandboxed?

No. Anthropic says mods run with the same access to your machine as Claude Code itself s1. Programs a mod starts also run outside the sandbox s2.

What happens if my guard mod crashes?

Without a .catch handler it is skipped and the command runs s7. Add a .catch that returns { deny: ... } so it fails closed.

Does a mod cost tokens?

Only if it calls the model. Of the ten we ran, next-steps and cache-keeper showed a measurable cost; the others showed no robust overhead in our run.

How do I turn mods off quickly?

Disable one in /plugin, start a session with --safe-mode, or set "disableAllHooks": true in ~/.claude/settings.json s2. None of these stop built-in mods.

Can I check what a mod does before installing it?

Yes. claude plugin validate lists the hooks, the API calls and the environment variables it reads s6.