TL;DR
- Four tools sit on four different slices of a Claude Code bill: graphify on what gets read, rtk on shell output, Superpowers on history and model choice, caveman on what the agent writes. Rank them by the size of the slice they touch, not by the percentage on their README.
- Superpowers is the only one that moves the bill: a fresh subagent per task and the least powerful model that fits change what is read and who reads it. Its cost is a gate on every task, even the tiny ones.
- graphify is second: a local tree-sitter graph built with zero LLM credits, 91.8x fewer tokens per query than a naive read of our own corpus.
- rtk compresses the one slice a Bash hook can see. JetBrains measured 19.7% of what the model reads as compressible, a ceiling around 3% of the bill, and +7.6% per task in the end-to-end test.
- caveman advertises 65% and measured 8.5% of output tokens on agentic work. Claude Code ships the same idea as the Concise output style.
- The two headline numbers people quote, 11.6M tokens saved and +7.6% per task, measure different things: bytes of bash output versus the cost of a finished task.
What the measurements say
The map first. rtk's own savings doc draws the tree of what a session reads and marks bash output as the only part the proxy filters, then says it plainly: "A command showing 90% fewer output bytes does not make your session 90% cheaper" s3. JetBrains replayed 83 baseline sessions, 1.9 million characters of tool output, and split them three ways: shell output rtk can compress 373,339 (19.7%), shell output it has no rule for 879,326 (46.3%), file reading and search tools that skip rtk 646,613 (34.0%) s1. Read and Grep never pass through the Bash hook. The caveman README reaches the same conclusion from the other end: the reading is usually the bigger half of the bill s7.
graphify. The repo was created 2026-04-03 and stood at 113,946 stars and 11,078 forks on 2026-09-02, latest release v0.9.53, Python, Apache-2.0 s5. The README benchmark row reads "Graph build | LLM credits | 0": the code is parsed locally with tree-sitter across ~40 languages, communities are split with Leiden, nothing leaves the machine s5. The author's post on r/ClaudeAI reported 73k stars and 2.2M downloads in 2.5 months s10. Our own graphify benchmark on a 1,701,550-word project (about 2,268,733 tokens as a naive read) built 34,031 nodes and 56,865 edges, an average query cost of about 24,702 tokens, 91.8x fewer tokens per query; per question the spread ran from 679.1x for "what is the main entry point" to 34.0x for "what connects the data layer to the api" s5. The 91.8x is against a naive full read, which nobody does on purpose; the graph also goes stale as the code moves and the optional semantic pass is the one step that costs model calls.
rtk. Created 2026-01-22, 78,326 stars and 4,942 forks on 2026-09-02, release v0.47.0, Rust, Apache-2.0, "100+ supported commands, <10ms overhead" s4. The launch post claimed "cargo test: 155 lines → 3 lines (-98%)", "git status: 119 chars → 28 chars (-76%)" and "Total over 2 weeks: 10.2M tokens saved (89.2%)" s9. On our machine, rtk 0.42.3 reported 25599 commands and 11.6M tokens saved (41.6%), with find, read and grep at the top of the By Command table s4. JetBrains installed rtk v0.43.0 exactly as rtk init -g ships it, on Claude Code 2.1.201 and claude-sonnet-5, and ran 86 tasks twice. Only 1 in 3 shell commands is something rtk can rewrite (349 rewritable, 525 no rule, 182 skipped, of 1,056 commands). Squeeze all compressible output by 70% and the total saving tops out around 3% of the bill; the measured result was +7.6% more expensive per task, with no difference at high effort s1. The README now concedes the point: "RTK cuts up to 90% of the bash output your agent reads. That is what RTK measures, and it is not the same as cutting your bill by 90%", and the counts it reports are estimated as bytes / 4 s4.
Superpowers. Created 2025-10-09, 280,792 stars and 25,164 forks on 2026-09-02, release v6.3.0 with fourteen skills, MIT, one install command: /plugin install superpowers@claude-plugins-official s6. The brainstorming skill opens on a hard gate: no code, no scaffolding, no implementation skill until an explicit intent is approved; three paths (spike, bounded, architectural) and the rule "When in doubt between two paths, take the heavier one" s12. writing-plans assumes the engineer has zero context and cuts work into steps where "Each step is one action (2-5 minutes)" s13. subagent-driven-development gives each task a fresh subagent that receives only its task's context, never the session history, with a review after every task and a broad review of the branch at the end. Its model section says "Use the least powerful model that can handle each role to conserve cost", warns that an omitted model inherits the session's model, and states "Turn count beats token price". Five rounds maximum per task: rounds 1 to 3 resume the implementer, round 4 brings a fresh implementer on a more capable model, round 5 the orchestrator rules itself s14. No compression anywhere: the saving comes from reading less history and paying a smaller model for mechanical steps. The full walkthrough is in our earlier video s11.
caveman. Created 2026-04-04, 102,548 stars and 5,967 forks on 2026-09-02, release bin-v1.1.5, Go; the skill is MIT, the proxy runtime BSL-1.1 s7. Its own table, ten prompts through the Claude API, shows an average of 1214 output tokens without and 294 with, 65%, best case 87% and worst case 22% s7. The README's IMPORTANT block is honest: the skill only shortens output, input and reasoning tokens do not change, and the rules cost about 1 to 1.5k input tokens every turn, so whole-session savings land lower than the chart s7. JetBrains ran 82 paired tasks on Claude Code 2.1.200 with claude-sonnet-5 at effort low, about USD 106: "Advertised saving: 65%. Measured saving: 8.5%", 592k to 542k output tokens, no detectable quality degradation (sign test p = 0.82), recommendation "use it if you like it" s2. Claude Code's built-in Concise style does the same job: "Claude leads with the result, skips preamble and narration, and keeps responses short by default", requires v2.1.237 or later, is picked through /config and saved as outputStyle in .claude/settings.local.json. Styles apply to the main conversation only; a subagent runs its own system prompt s8.
Verdict table
| Tool | Slice of the bill | Advertised | Measured | Moves |
|---|---|---|---|---|
| Superpowers | history + model per task | no number | fresh context per task, least powerful model per role | the bill |
| graphify | what gets read | 0 LLM credits to build | 91.8x fewer tokens per query vs naive read (ours) | the bill, on reads |
| rtk | compressible shell output | 60-90% on commands | 19.7% compressible, ceiling around 3%, +7.6% per task (JetBrains) | the margins |
| caveman / Concise | what the agent writes | 65% | 8.5% of output tokens (JetBrains) | the margins |
Do this Monday
- Open your last long session and count where the input went: how much was Read and Grep, how much was shell output, how much was replayed history. Place each tool on its slice before installing anything.
- Install Superpowers with
/plugin install superpowers@claude-plugins-officialand run one real feature through brainstorming, writing-plans and subagent-driven-development. Watch the context gauge of the orchestrator stay flat. - Set an explicit model on every subagent you dispatch. An omitted model inherits the session's model and you pay the orchestrator's tier for mechanical work.
- Run
/graphify .on your largest repo, thengraphify benchmark, and compare the per-query cost with the naive corpus size it prints. Rebuild the graph when the code moves. - If you already run rtk, read
rtk gainas bytes of shell output, not money. Pair it with the JetBrains split: whateverfind,readandgrepsave, the Read and Grep tools never passed through the hook. - Switch to the Concise output style through
/configbefore adding caveman; it ships with Claude Code and costs no per-turn skill rules. - Keep one measurement of your own: task cost before and after each change, same task set, not a single command's byte count.
Go further
- JetBrains' full rtk methodology, with the 83-session replay and the 1,056-command split, is the reference for measuring any shell proxy s1.
- The caveman benchmark explains how 82 paired tasks and a sign test turn a 65% chart into 8.5% of output tokens s2.
- rtk's savings-explained page is the clearest tree of what a session actually reads; read it before trusting any "tokens saved" counter s3.
- The graphify README benchmarks section documents the zero-credit build and the semantic pass, the one step that costs model calls s5.
- The Superpowers model selection section, with "Turn count beats token price" and the five-round cap, is the part worth copying into your own agent definitions s14.
- The brainstorming hard gate and the three paths show how the ceremony scales with the task, and where it fires on fixes too small to deserve it s12.
- Output styles in the Claude Code docs: the Concise style, the v2.1.237 floor, and why subagents ignore it s8.
Sources
- Does rtk really save tokens in Claude Code?, JetBrains. Why read it: the only end-to-end test of rtk, with the 19.7% / 46.3% / 34.0% split of what an agent reads.
- Speak to AI agents like cavemen to save tokens, JetBrains. Why read it: 82 paired tasks that reduce the 65% claim to 8.5% without finding a quality loss.
- How RTK savings work, GitHub. Why read it: the maintainers' own tree of the bill and the sentence about 90% fewer bytes.
- rtk-ai/rtk on GitHub, GitHub. Why read it: the README now states what rtk measures and how the counts are estimated.
- Graphify-Labs/graphify on GitHub, GitHub. Why read it: benchmarks table, the zero-credit build, and the
graphify benchmarkcommand used here. - obra/superpowers on GitHub, GitHub. Why read it: the plugin that changes what gets read and which model reads it.
- JuliusBrussee/caveman on GitHub, GitHub. Why read it: the savings table and the IMPORTANT block that undercuts it.
- Output styles, Claude Code docs. Why read it: the built-in Concise style and its limits with subagents.
- rtk launch post on r/ClaudeAI, Reddit. Why read it: the original 10.2M claim, useful to compare with what JetBrains measured.
- Graphify hit 73k stars and 2.2M downloads (r/ClaudeAI), Reddit. Why read it: the author on adoption and what the graph is for.
- Superpowers: The Plugin That Disciplines Claude Code, AIDive. Why read it: our full walkthrough of the plugin, skill by skill.
- Superpowers brainstorming skill (SKILL.md), GitHub. Why read it: the hard gate and the three paths, verbatim.
- Superpowers writing-plans skill (SKILL.md), GitHub. Why read it: the 2 to 5 minute step granularity that keeps subagent contexts small.
- Superpowers subagent-driven-development skill (SKILL.md), GitHub. Why read it: model selection per role and the five-round cap.
FAQ
Why does rtk show 11.6M tokens saved if it barely moves the bill?
The counter measures bytes of shell output divided by 4, on the commands that went through the hook. Read and Grep tool calls, system prompt and replayed history never pass through it, and JetBrains found only 19.7% of what the model reads is compressible that way.
Does Superpowers compress anything?
No. It reads less by giving each task a fresh subagent with only that task's context, and pays less by assigning the least powerful model that can handle the role. The cost is a gate on every task.
Is caveman worth installing?
JetBrains found no quality loss and 8.5% fewer output tokens, so use it if you like the style. The Concise output style in Claude Code v2.1.237 or later gives the same effect without the 1 to 1.5k input tokens the skill's rules cost each turn.
When does graphify stop paying off?
When the graph is stale or the question needs the semantic pass, which does cost model calls. The 91.8x figure compares a query against reading the whole corpus, so smaller repos see a smaller gap.
AIDive