# I Tested 10 Claude Code Mods. 3 Earned Their Place

Video: https://www.youtube.com/watch?v=k7BWw5gz8aQ
Article: https://aidive.dev/videos/claude-code-mods-measured-verdict/
Published: 2026-10-05

## Chapters

- [0:00](https://www.youtube.com/watch?v=k7BWw5gz8aQ&t=0s) Intro — ten mods, three survive
- [0:42](https://www.youtube.com/watch?v=k7BWw5gz8aQ&t=42s) The wave, and what we ran
- [2:22](https://www.youtube.com/watch?v=k7BWw5gz8aQ&t=142s) The decoration tier, measured
- [3:38](https://www.youtube.com/watch?v=k7BWw5gz8aQ&t=218s) The flagship bills you every turn
- [5:03](https://www.youtube.com/watch?v=k7BWw5gz8aQ&t=303s) Mod vs hook, the same job
- [6:30](https://www.youtube.com/watch?v=k7BWw5gz8aQ&t=390s) The guard that does nothing
- [8:09](https://www.youtube.com/watch?v=k7BWw5gz8aQ&t=489s) What you grant when you paste one
- [10:14](https://www.youtube.com/watch?v=k7BWw5gz8aQ&t=614s) Keep three, delete seven

## TL;DR

- Of 10 hyped Claude Code mods tested on a real week of work (85 sessions, ~6,000 tool calls), only 3 earn a keep: the collision guard, the model router, and the auto handoff.
- The flagship suggestion engine costs about 250 extra output tokens and ~2.9 s per qualifying answer, and its session fork fires even in headless runs where nothing is drawn.
- A guard mod that crashes fails open by default: the command runs anyway. One catch handler returning a deny makes it fail closed, 3 runs out of 3.
- Mods are not sandboxed: permission rules govern Claude's tool calls, not the mod's own file reads, and a child process walks through the network policy.
- The two-minute discipline: run `claude plugin validate` and read the audit before enabling any mod.

## Intro — ten mods, three survive

Claude Code mods are TypeScript functions inside plugins that can redraw Claude Code's interface or rewrite what it does, and within a day of their early-October launch, three separate video tours told developers to install ten of them. Two of those creators admit on camera that some of the mods save nothing, and none of them measured a single one. This test does: every one of the ten hyped mods ran on a real week of work, with a number on its overhead, its savings, and a keep-or-delete verdict. The result up front: only three of the ten earn a place on a working developer's machine, and one of them quietly spends tokens on every single answer.

## The wave, and what we ran

Anthropic's own definition fits in one breath: a mod is a function that hooks into an event, and it can run before it, after it, or instead of it. An event is a tool call, a submitted prompt, or a part of the interface being drawn. Mods are plain TypeScript, shipped inside a plugin you install like any other.

The launch tweet crossed four million views in about a day, with twenty thousand likes and more bookmarks than replies and reposts combined. Two days after launch, a community catalog had already scanned more than a thousand public mods across hundreds of repositories.

The bench for this test is one real week of work: 85 sessions across four projects, nearly 900 prompts, and just under 6,000 tool calls. Every mod went through the validator's audit, which lists what it hooks and what it can touch, and every mod ran the same task against a clean baseline, on the current release, on the same machine.

The headline split: six of the ten cost nothing you can measure at runtime, three cost real time or real tokens, and one breaks the single promise it makes.

## The decoration tier, measured

The quiet six come in within the noise of the connection. The goal meter, the repo heatmap, the flight recorder, the model router, the session bookmarks and the auto handoff all sit between a quarter second saved and a fifth of a second added, on a baseline task of about four seconds.

| Mod | Runtime delta | Notes |
| --- | --- | --- |
| Goal meter | within noise | decoration pane |
| Repo heatmap | within noise | lights up files as they are read |
| Flight recorder | within noise | live timeline of the turn |
| Model router | within noise | pays off on subagents, see verdict |
| Session bookmarks | within noise | broad capability footprint |
| Auto handoff | ~70 ms idle | writes one handoff at a context threshold |

They look great, and nothing broke: every run of every configuration finished the task correctly. Headless they cost nothing, because there is nothing to draw; in a terminal these panes redraw up to thirty times a second, so the honest cost of the decoration tier is attention rather than tokens.

The validator's audit is where it stops being funny. The bookmarks mod can call the model, start processes on your machine, and write files, and it reads your configuration paths from the environment. None of that is hidden and none of it is malicious, but it is a lot of reach for a bookmark. Four mods exit here: the goal meter, the heatmap and the flight recorder as decoration with zero measured benefit, and the bookmarks mod because it asks for more than it earns.

## The flagship bills you every turn

The suggestion engine is the mod every video demos first: your answer finishes, three prompt suggestions appear above the composer, you press a number and the draft fills itself. The mechanism is documented by its own author: when your turn completes, it forks the session to ask a model for those suggestions, and the fork shares the session's prompt cache, so it costs about one short reply. The tours never mention that line.

Measured on the bench, that is about 250 extra output tokens per qualifying answer and nearly three seconds of extra wall time, and a qualifying answer is almost every answer: anything longer than about eighty characters. The fork also has no surface gate. The suggestions only ever draw in a terminal, but the fork fires everywhere, including headless runs where nothing can be drawn at all.

| Mod | Measured cost | When it fires |
| --- | --- | --- |
| Suggestion engine | ~250 output tokens + ~2.9 s per qualifying answer | every answer over ~80 characters, including headless |
| Cache keeper | ~1.5 s per turn, plus warming-mode model calls | every turn, warming for hours |

The cache keeper has the same shape: about a second and a half per turn, with a warming mode that spends small model calls for hours to stop your prompt cache going cold. On a subscription plan the cache window is already an hour, so you are paying pings to solve a problem the plan mostly solved. Both are honest designs with documented costs, both are taxes on every turn that the install lists never price, and both come off the machine.

## Mod vs hook, the same job

Claude Code already had hooks: a shell script in your settings that fires on the same events. The documentation answers the choice in one table row — a mod is for interface and for rewriting events; a hook is for blocking, allowing, or logging with a script you already have.

The measurable difference is the process spawn. A settings hook starts a fresh process on every tool call. Timed on this machine, a shell hook doing nothing costs about 8 ms and a hook that starts Node about 43 ms, on every single call, before the script does anything. Across the bench week's 5,993 tool calls that is over four minutes of pure interpreter startup. A mod pays none of it: its handler runs inside the engine's own process, and the engine's log shows the hop settling in about a millisecond.

| Handler | Cost per call | A week of 5,993 calls |
| --- | --- | --- |
| Shell hook (no-op) | ~8 ms | ~48 s |
| Node hook (no-op) | ~43 ms | ~4.3 min |
| Mod (in-process) | ~1 ms | ~6 s |

The one real migration report in the wild says the same thing: twenty-seven shell hooks collapsed into five mods, and the spawn on every call disappeared with them. The rule that survives: interface or event rewriting, mod; block, allow, or log with a script you trust, hook — the spawn cost only matters at thousands of calls; knowledge you keep repeating, skill. A hook you've read beats a mod you haven't.

## The guard that does nothing

The simplest safety mod possible is a guard that watches every shell command — and this one was written to crash. Claude Code was asked to create a marker file; the guard threw; the command ran anyway and the file appeared. That is not a bug but the documented default: when a hook throws, times out, or returns the wrong shape, Claude Code skips it and moves on. A broken decoration should not brick a session, but a broken guard fails open, silently, with one line in a debug log nobody reads.

The fix is one catch handler that returns a deny. The same crashing guard with the catch refuses the command and names the failure. One line decides whether a guard fails open or fails closed, the documentation ships this exact pattern, and almost nobody installs it.

A community team re-ran the cases on the current release and judged by marker files instead of by what the model said. The catch pattern failed closed three runs out of three — and one path is still silently broken: a deny returned after the call was already forwarded does not stop the tool. The file landed three times out of three while the model was told the write failed.

The field report that named this problem ran a guard for days that was enabled, loaded, and doing nothing, because a stale flag had switched it off underneath: three green status chips over a counter stuck at zero. Silence that looks exactly like health.

The collision guard earns the first keep. It solves a real problem — two open chats editing the same file — and its failure mode is loud: it asks in a dialog and never silently allows. It costs about half a second on edits and adds nothing to the prompt. Install it, and give it the catch handler anyway.

## What you grant when you paste one

Anthropic says it in plain words on launch day: mods run with the same access to your machine as Claude Code itself. They aren't sandboxed; install them the way you'd install any code on your computer. Concretely, a mod can act on your machine as you: read your environment and settings, where API keys live; see every prompt and every tool call; rewrite them; approve a tool call before you are ever asked; and spend your plan's usage on model calls of its own.

Two traps catch even careful users. Permission rules govern Claude's tool calls, not the mod's own calls: deny Claude an env file, and a mod can still read that file directly with its own file access, or start a program that does. Network policy has the same edge: turn off web traffic and the mod's own fetch calls are refused, but a child process the mod starts reaches the network with full access. There is a built-in guard mod that loads ahead of everything, but only on managed machines and for Team or Enterprise seats; a solo seat on a personal subscription gets none of it.

None of this is theoretical. A user published a proof of concept days after launch: a mod whose button starts a program and writes into the home directory, installed from the catalog with no warning — and his point stands: the catalog looks like an app store, which suggests a vetting that isn't there. A separate hook bug broke subagent isolation for a day; the maintainer called it a big blunder and fixed it one release later.

The catalog's own scan of more than a thousand public mods: over four hundred start host processes, nearly four hundred read files, and over three hundred see every tool call. The catalog's caveat is the right frame — that is a footprint, not a verdict; a PR tracker has to run git. The discipline costs two minutes: run the validator before enabling anything, and know the exits — safe mode for one session, one setting to stop every installed hook for good.

## Keep three, delete seven

Of the ten, three earn their place: the collision guard, the model router, and the auto handoff.

| Mod | Verdict | The number behind it |
| --- | --- | --- |
| Collision guard | keep | ~0.5 s on edits, fails loud, nothing added to the prompt |
| Model router | keep | subagent billed on the cheap model, a third of the price |
| Auto handoff | keep | 70 ms of nothing, one handoff write at the context threshold |
| Suggestion engine | delete | ~250 output tokens + ~2.9 s on every qualifying answer |
| Cache keeper | delete | ~1.5 s per turn, warming pings against a 1-hour cache window |
| Recording mode | delete | masks the screen but not the disk |
| Goal meter | delete | decoration, zero measured benefit |
| Repo heatmap | delete | decoration, zero measured benefit |
| Flight recorder | delete | decoration, zero measured benefit |
| Session bookmarks | delete | reach far beyond its job |

The model router has a receipt: a session on the big model spawned one subagent, and the run's own usage readout showed the subagent billed on the cheap model, a third of the price for the same small job. On weeks heavy with subagents that is real money. The auto handoff costs nothing until the moment it pays: seventy milliseconds of idle overhead, and past a context threshold it writes the handoff for the cold start, once. One of the wave's own creators admits the manual handoff button doesn't really save time; as an automatic threshold write, it actually does.

The discipline that survives the test: read the validator's audit before enabling anything, give every guard its catch handler so it fails closed, and record demos in safe mode rather than trusting a masking mod. The limits are real — one week, one machine, one workload, three runs per point on the small model. Your three might differ, but now you know how to find them.

## FAQ

### What are Claude Code mods?

Mods are TypeScript functions inside Claude Code plugins that hook into an event — a tool call, a submitted prompt, part of the interface being drawn — and can run before it, after it, or instead of it. They can redraw the interface or rewrite what Claude Code does.

### Which Claude Code mods are actually worth installing?

On a measured week of real work, three of the ten hyped mods earned a keep: the collision guard (stops two chats editing the same file, fails loud), the model router (a subagent billed at a third of the price), and the auto handoff (70 ms of idle cost, one automatic handoff write).

### Do Claude Code mods cost tokens?

Most don't, but the suggestion engine forks the session after every qualifying answer, costing about 250 output tokens and nearly three seconds of wall time per turn, and the cache keeper's warming mode spends small model calls for hours.

### Are Claude Code mods safe to install?

Mods run with the same access to your machine as Claude Code itself and are not sandboxed. Permission rules govern Claude's tool calls, not the mod's own calls, so a mod can read files or start processes your rules deny to Claude. Run the validator's audit before enabling anything.

### What is the difference between a Claude Code mod and a hook?

Both fire on the same events. A hook is a shell script that pays a fresh process spawn on every tool call (about 8 ms for shell, 43 ms for Node), while a mod's handler runs inside the engine's process in about a millisecond. Use a mod for interface or event rewriting, a hook for blocking, allowing or logging with a script you trust.

### What happens if a Claude Code guard mod crashes?

By default it fails open: Claude Code skips the broken handler and the command runs anyway, with one line in a debug log. Adding a catch handler that returns a deny makes the guard fail closed, and the docs ship that exact pattern.

## Sources

- [Customize Claude Code with mods](https://claude.com/blog/claude-code-mods) — Anthropic
- [Mods overview — Claude Code docs](https://code.claude.com/docs/en/plugins/mods/overview) — Anthropic
- [React to events with a mod — fail-open and fail-closed handlers](https://code.claude.com/docs/en/plugins/mods/events) — Anthropic
- [Mods for admins — security caveats](https://code.claude.com/docs/en/plugins/mods/admin) — Anthropic
- [next-steps plugin README](https://github.com/anthropics/claude-plugins-community/tree/main/next-steps) — GitHub
- [awesome-claude-code-mods — community catalog and capability scan](https://github.com/karanb192/awesome-claude-code-mods) — GitHub
- [Hook fail-open behavior discussion and re-test](https://github.com/anthropics/claude-code/issues/91870) — GitHub
- [The Guard I Installed Was Enabled, Running, and Doing Nothing](https://www.practicalsystems.io/blog/claude-code-function-hooks-mods-layer) — Practical Systems
