AIDive

Video pack

Claude Code Mods: the fail-open experiment, measurements, checklist and sources

10 min read

TL;DR

  • A Claude Code Mod is three files: .claude-plugin/plugin.json, hooks/hooks.json holding {"modules":["./index.ts"]}, and a TypeScript module that exports register(on). Nothing else is needed to load it.
  • On build 2.1.272 the typed surface written by /plugin-types is 11,783 lines of claude-code.d.ts: 84 event or call names over 23 nouns. fs.readFile is gone, the noun is fs.read and fs.write.
  • A 42-line tool.call hook made the model read API_KEY=[REDACTED] instead of the real key for both the Read and the Bash tool, at 27.9 ms per healthy hop and about 0 tokens added to the session.
  • A hook that sleeps past the 10 s budget or throws is skipped and the command below it runs. The log says so, the effect is a bypass: a guard Mod fails open.
  • The generated path works: one sentence produced a 190-line Mod plus 53 lines of tests in about 4 minutes and $1.23, passing claude plugin validate with no warnings and claude plugin test 4 pass / 0 fail.
  • One thing did not reproduce: the Mod loaded under claude -p but stayed silent in a pty-driven interactive REPL across two attempts. Treat REPL loading as unverified until you test it on a real terminal.

What the measurements say

Everything below was run on Claude Code 2.1.272 with CLAUDE_CODE_ENABLE_FUNCTION_HOOKS set, on a hand-written Mod called redact-secrets and on three throwaway Mods built to break it. The feature itself is tracked in the Function Hooks proposal issue, which is still the closest thing to an official spec s3.

The tooling exists before the docs do. Build 2.1.272 ships claude plugin validate, test, eval and details; test works even though plugin --help does not list it in its command block s3. Running /plugin-types inside a session wrote 11,783 lines into .claude/types/claude-code.d.ts, plus a claude-code-mcp.d.ts of 3,438 lines covering 150 MCP tools from 7 servers s3. Counting the generated types gives 84 event or call names spread over 23 nouns, and the file renames one thing you may have read about in September posts: fs.readFile no longer exists, the surface is fs.read and fs.write s9. The same types say a tool.call hook returns { result, context? } or { deny }; text and ref come from core and are not part of a hook's own answer, so you rewrite result, not text s3.

claude plugin validate prints a footprint before the Mod ever runs: ./index.ts hooks: tool.call and ./index.ts calls: $.ui.toast, or calls: nothing on $ when the module touches no host capability s4. This static line is the only vetting a directory like awesome-claude-code-mods can automate today, which matters when you read the next paragraph.

The redaction works in -p mode. The 42-line hook intercepted tool.call, and the model received API_KEY=[REDACTED] where the file and the shell output held sk-test1234567890abcdef, for both the Read tool and the Bash tool s3. The debug log timed a healthy hop at 27.9 ms round trip, worker hop and next() included s3. claude plugin details priced the Mod at about 0 tokens added to every session: a Mod is code in the process, not prompt text, which is the main argument its promoters make against shell hooks and skills s7.

The guard fails open. A slow-guard Mod that sleeps 15 s was cut at the 10 s budget with hook failed: slow-guard: exceeded 10000ms budget (tool.call; skipped; what is below it ran in its place), and echo hi ran anyway s3. A throw-guard Mod that throws was skipped the same way, hook failed: throw-guard: boom (tool.call; skipped; what is below it ran in its place), reported 574.2 ms in s3. Loud in the log, bypassed in effect. Any Mod whose job is to block something must be read with that in mind: a bug in the guard is a hole, not a crash.

Generation is cheap. From one sentence, the model wrote a 190-line working Mod plus 53 lines of tests in about 4 minutes, 34 turns and $1.23 (19,053 output tokens, 497,282 cache-read, 8,357 thinking) s3. That Mod passed claude plugin validate with no warnings and claude plugin test with 4 pass / 0 fail in 0.31 s, and hid the keys live s3. On the hand-written Mod, claude plugin test ran with no API key and no model call: 1 pass in 0.25 s s3. A skill that teaches agents to write Mods already exists if you want to repeat this with a template s12.

What did not reproduce: $.ui.toast never rendered because the Mod did not load in the pty-driven REPL across 2 attempts; in -p the equivalent surfaced only as a debug line. The opposite of the "REPL-only" claim circulating on X was observed, and the root cause was not isolated s8. The 5 s wedged-worker heartbeat was not tested either; only the 10 s await budget and the throw path were measured.

Measurements

Case What the Mod does Outcome Time
redact-secrets, Read tool Rewrites result on tool.call Model sees API_KEY=[REDACTED] 27.9 ms per hop
redact-secrets, Bash tool Same hook, shell output Model sees API_KEY=[REDACTED] 27.9 ms per hop
slow-guard Sleeps 15 s inside tool.call Skipped, echo hi ran cut at 10000 ms
throw-guard Throws boom inside tool.call Skipped, command ran 574.2 ms
generated Mod 190 lines + 53 lines of tests from one sentence validate: no warnings; test: 4 pass / 0 fail ~4 min, 34 turns, $1.23
claude plugin test on redact-secrets No API key, no model call 1 pass 0.25 s
claude plugin details Session cost of the Mod ~0 tokens added n/a

Protocol: Claude Code 2.1.272, function hooks enabled by environment variable. Each Mod is a plugin directory with plugin.json, hooks/hooks.json and one index.ts. Runs went through claude -p with debug logging on; a planted file and a shell command both contained sk-test1234567890abcdef. Fail-open cases were driven by a Mod that sleeps 15 s and a Mod that throws, with echo hi as the command under guard. The interactive REPL was driven through a pty and did not load the Mod in two attempts.

Do this Monday

  • Run /plugin-types in a session and open .claude/types/claude-code.d.ts: search for fs.read and tool.call before trusting any snippet from a September post.
  • Write the three-file Mod skeleton (plugin.json, hooks/hooks.json with {"modules":["./index.ts"]}, index.ts exporting register(on)) and run claude plugin validate on it: read the hooks: and calls: footprint lines.
  • Port your most-used shell hook to a tool.call handler that rewrites result, then compare the hop time in the debug log with the shell version.
  • Add a claude plugin test file next to the Mod so the guard runs without an API key in CI.
  • Wrap every guard handler in a try/catch that returns { deny } on failure: on this build an exception or a 10 s stall skips you and lets the command through.
  • Test the Mod under claude -p and in your real interactive terminal separately, and write down which one loaded it.
  • Before installing a third-party Mod, run claude plugin validate on it and reject anything whose calls: line names host capabilities the Mod has no reason to use.

Go further

  • Read the proposal issue end to end, including the architecture PDF attached in the comments: it is the only written contract for register(on), budgets and the worker hop s3.
  • Compare against the Command Code Mods design, which runs TypeScript against a ModApi in the host process: the two systems share the shape and the early-access gating s2.
  • The claudefa.st preview was written while the feature was still a proposal: useful to see what changed between the Sep 3 text and the 2.1.272 binary s9.
  • Prathkum's note tweet is the clearest short statement of why in-process hooks beat shell hooks on tokens and latency s7.
  • cc-mod-waitwhat is a good first Mod to read: UI above the prompt, nothing written into the transcript s11.
  • cc-arcade shows how far $.ui reaches: games rendered above the prompt from a Mod s5.
  • The hooks reference still describes the shell model; keep it open to map each old event to its new noun.event name s1.
  • Skeptics on X argue that plugins already cover this and that the surface will break every release; the renamed fs noun is one data point for them s23.

Sources

  • Function Hooks proposal (issue #91870), GitHub, anthropics/claude-code. Why read it: the only spec-like text for Mods, with the architecture PDF and the shipping trail in the comments.
  • Hooks reference, Claude Code docs. Why read it: the shell-hook model you are migrating from, event by event.
  • Command Code Mods documentation, Command Code. Why read it: prior art with the same TypeScript-in-process shape, useful to spot what Anthropic copied or avoided.
  • awesome-claude-code-mods, GitHub, karanb192. Why read it: auto-scanned directory of public Mods, with the validate footprint as its only vetting.
  • cc-arcade, GitHub, sezaakgun. Why read it: the demo that made the feature visible, and a tour of $.ui.
  • Boris Cherny announcement tweet, X, Boris Cherny. Why read it: the launch statement from an Anthropic engineer, since no blog post exists.
  • Prathkum: Function Hooks explained, X, Prathkum. Why read it: the best short explanation of hooks vs Mods for someone who already writes shell hooks.
  • shipnotesai reaction thread, X, shipnotesai. Why read it: where the "REPL-only" claim circulated, which our run contradicted.
  • Claude Code Function Hooks: Preview Behind a Flag, claudefa.st. Why read it: pre-ship explainer, good for diffing the proposal against the binary.
  • cc-mod-waitwhat, GitHub, GGGODLIN. Why read it: a small readable Mod that writes to the UI and not the transcript.
  • claude-mods-skill, GitHub, BeLazy167. Why read it: a skill for generating Mods, if you want to repeat the one-sentence experiment.
  • AxialisSoftware reaction, X, AxialisSoftware. Why read it: the skeptical case, in one tweet.

FAQ

Does a Mod replace my shell hooks today?

Not yet for anything that must block. On 2.1.272 a hook that throws or stalls past 10 s is skipped and the command runs. Shell hooks keep working, so keep the blocking ones there until a declared catch or a fail-closed option lands.

Why does validate matter if the Mod runs fine?

Because its hooks: and calls: lines are the only static view of what a Mod touches on $. For your own Mod it confirms the footprint; for a third-party Mod it is the whole review you get before the code runs in your process.

How much does a Mod cost per session?

claude plugin details reported about 0 tokens added. The Mod is code running in the engine, not text in the prompt, which is the main advantage over a skill or a CLAUDE.md rule.

Why did the Mod load in -p but not in the REPL?

Unknown. Two pty-driven attempts stayed silent while claude -p loaded and applied the hook. The likely cause is the pty environment rather than the feature, so test on your own terminal before relying on either direction.