TL;DR
- Inside Claude Code, jev-gateway never forces a tool. One line,
steer: thinking || cached ? "hint" : "tool_choice", switches it to hint mode as soon as the request carries extended thinking or a cached conversation, and a real Claude Code request carries both on the first turn. - The hint is a two-sentence
<system-reminder>appended to the last user message. The model is free to ignore it, and when Claude Code has already appended its own system reminder as the last block, the hint is not attached at all. - The gateway's own benchmark (120 sessions) says routing pays on debugging and costs on feature work for Claude models: Opus 5 at +61% input tokens, +47% requests and +83% time on the feature task, Sonnet 5 at +16% input and +37% time.
- Codex is where Jev bites: the gateway forces the tool there, Jev steered 76 to 100% of Codex requests against 34 to 51% of Claude Code requests.
- Measured on our machine: a clean Claude Code 2.1.280 request already carries 24 tools and 47,411 prefix tokens; a normal setup with MCP servers carries 40 tools and 57,277 tokens, and the gateway re-sends that roster to Jev on every call.
- fast-jev-compaction, the most starred Jev tool, has open issues saying its hooks do not register on current Claude Code builds and that full transcripts leave for a third-party API. Not yet.
What the measurements say
Jev is a decision model, not a text generator. Its vendor prices input at $0.042 / MTok with output free, quotes an end-to-end response time of 70ms-500ms, and writes under its own "193.6x faster, 444.6x cheaper" headline that these figures "are on the higher end of real world gains" s3. The same post admits the reference answers average GPT-6 Astra and Fable 5.1, which biases the comparison toward OpenAI and Anthropic models s3.
jev-gateway plugs into Claude Code through a single environment variable: bin/clients.mjs sets ANTHROPIC_BASE_URL to the local gateway and leaves the Max login alone s1. In src/adapters/messages.ts the gateway asks Jev which tool fits the next step, then decides how to pass the answer on. When the request has thinking enabled or cache_control blocks, it hints. Otherwise it sets tool_choice s1. The hint reads: a tool-routing model suggests the named tool is the most relevant next step, ignore this if it does not fit what the user actually asked for. It is appended to the last user message as a <system-reminder> block s1.
The Anthropic API leaves no other choice. With manual extended thinking enabled, tool_choice: any and tool_choice: tool are not supported and return an error, and Claude Opus 5.5, Claude Fable 5.1 and Claude Mythos 5.1 return a 400 for forced tool use regardless of thinking s4. A nuance the gateway's own comment misses: the docs say Claude Opus 5 does support forced tool choice with thinking on s4. On caching, the hierarchy is tools then system then messages; a change to tool_choice invalidates only the messages cache, while editing a tool definition invalidates the entire cache, which is why the gateway appends a block instead of rewriting a tool description s5.
The benchmark almost nobody quotes is the gateway author's own. Six models, two tasks, five runs per mode, 120 agent sessions over 2026-09-18 and 19, GPT models in Codex 0.154, Claude models in Claude Code 2.1, every agent clean with no MCP servers, plugins or skills s2. On chess-bugfix every model used fewer tokens with routing and nothing got less correct. On chess-san, the feature task, routing made Opus 5 and Sonnet 5 clearly worse, and the authors name the cause: the gateway only hints with Claude models, so a hint that does not fit costs a detour instead of being ignored for free s2. Routing also cost correctness once: GPT-5.6 Luna solved chess-san five times out of five alone and three out of five with routing s2. The authors add that input tokens are mostly cached (80 to 96%), so an input saving is worth less money than the same saving in output tokens, and that Jev itself cost between half a cent and ten cents per five runs s2. One of the 120 runs, chess-bugfix.on.3 in the Luna series, is marked contaminated after the agent found another run's test script in /tmp s2.
The README footnote that motivated our own test: with --user-tools, one setup sent 285 tools and about 200,000 tokens with every Claude Code request, against 6 tools and 7,000 tokens clean s2. We measured the same seat. A clean Claude Code 2.1.280 request carries 24 tools, 87,547 characters of tool definitions and 47,411 billed prefix tokens (16,221 written, 31,190 read); the full setup carries 40 tools, 93,179 characters of definitions and 57,277 prefix tokens, all written. Both carry thinking: {type: "adaptive"} and 3 cache_control blocks, no tool_choice, which is exactly the condition that locks hint mode in messages.ts s1. A single "ok" reply costs $0.07 API-equivalent on the clean Sonnet 5 run and $1.15 on the full Fable 5.1 run, read from Claude Code's own total_cost_usd and usage fields, the same seat the gateway's launcher occupies s1.
On fast-jev-compaction, the plugin behind the "instant compaction" thread (score 495, 117 comments) s9, the open issues matter more than the star count: #21 reports Hooks (0) on install because session.compact and turn.complete are not recognized hook events on Claude Code 2.1.272, #88 says hooks cannot replace compaction and full transcripts are sent to a third-party API, #65 documents 9 consecutive fabricated "work done" reports after one compaction, #89 says compaction is undone on --resume s7. The top complaint on the thread, at 84 points, is the vendor's ToS and data controls s9. The skill-suggestion cookbook is the one measured win the vendor publishes for an agent roster: wrong skill loaded drops from 16.8% to 7.3%, and a skill loaded when nothing fits drops from 9.8% to 4.0% s13.
Measurements
The gateway's benchmark, percentages against the same model with routing off s2.
chess-bugfix: find and fix five injected bugs
| Model | Solved, on / off | Output tokens | Input tokens | LLM requests | Seconds | Jev steered |
|---|---|---|---|---|---|---|
| GPT-6 Astra | 5/5 · 5/5 | 1,226 (-57%) | 96k (-7%) | 5 (0%) | 41 (-39%) | 100% |
| GPT-5.6 Sol | 5/5 · 5/5 | 3,211 (-57%) | 202k (-40%) | 9 (-36%) | 78 (-36%) | 93% |
| GPT-5.6 Luna | 1/4 · 0/5 | 10,519 (-12%) | 506k (-10%) | 19.5 (-15%) | 200 (+10%) | 86% |
| Fable 5.1 | 5/5 · 5/5 | 8,675 (-13%) | 276k (-19%) | 14 (-22%) | 148 (+6%) | 45% |
| Opus 5 | 5/5 · 5/5 | 16,693 (-7%) | 406k (-22%) | 18 (-14%) | 218 (+2%) | 38% |
| Sonnet 5 | 5/5 · 5/5 | 16,623 (-41%) | 616k (-48%) | 26 (-26%) | 243 (-25%) | 34% |
chess-san: add algebraic notation to a working engine
| Model | Solved, on / off | Output tokens | Input tokens | LLM requests | Seconds | Jev steered |
|---|---|---|---|---|---|---|
| GPT-6 Astra | 5/5 · 5/5 | 3,663 (0%) | 143k (+2%) | 7 (0%) | 88 (+8%) | 95% |
| GPT-5.6 Sol | 5/5 · 5/5 | 5,096 (-9%) | 147k (-39%) | 7 (-36%) | 78 (-16%) | 86% |
| GPT-5.6 Luna | 3/5 · 5/5 | 6,809 (-14%) | 315k (-51%) | 14 (-42%) | 121 (-14%) | 76% |
| Fable 5.1 | 5/5 · 5/5 | 13,497 (-24%) | 331k (-27%) | 13 (-19%) | 167 (-26%) | 51% |
| Opus 5 | 5/5 · 5/5 | 20,152 (+22%) | 676k (+61%) | 25 (+47%) | 390 (+83%) | 44% |
| Sonnet 5 | 5/5 · 5/5 | 23,487 (+9%) | 991k (+16%) | 32 (+3%) | 327 (+37%) | 42% |
Our own request capture, the seat jev-gateway occupies s1.
| Clean | Full | |
|---|---|---|
| Model Claude Code chose | claude-sonnet-5 | claude-fable-5-1 (user setting, 1M) |
thinking in the request |
{type: "adaptive"} |
{type: "adaptive"} |
cache_control blocks |
3 | 3 |
tool_choice |
absent (auto) | absent (auto) |
| Tools in the request | 24 | 40 (28 built-in + 12 MCP) |
| Tool definitions, characters | 87,547 | 93,179 |
| System prompt, characters | 27,754 | 12,436 |
| Whole request, characters | 134,882 | 155,718 |
| Prefix tokens billed (cache write + read) | 47,411 (16,221 written, 31,190 read) | 57,277 (all written) |
| Output tokens | 4 | 4 |
| API-equivalent cost of one "ok" | $0.07 | $1.15 |
Protocol: a 60-line logging proxy on 127.0.0.1:8790 forwards every request to https://api.anthropic.com byte for byte and logs what it carries, the exact seat bin/clients.mjs gives jev-gateway. Claude Code 2.1.280 ran headless, claude -p "Reply with the single word ok. Do not use any tool." --output-format json --max-turns 1, from a private Expo repo of 1,021 tracked files, on a claude.ai subscription. Clean: CLAUDE_CONFIG_DIR on an empty directory, --strict-mcp-config, --setting-sources project. Full: the machine's normal user settings, project .mcp.json, user MCP servers and installed plugins. One request per configuration, first turn only; no Jev key, so the regression figures are the gateway's bench replayed, not reproduced.
Do this Monday
- Before adding any router, measure your own seat: start a logging proxy, point
ANTHROPIC_BASE_URLat it, runclaude -p "Reply with the single word ok." --output-format json --max-turns 1, and readcache_creation_input_tokenspluscache_read_input_tokensin the output. - Count the tools in that request. If MCP servers you rarely use push the roster up, remove them from
.mcp.jsonor scope them per project; that cut lands on every request, router or not. - If you still want Jev in Claude Code, open
src/adapters/messages.tsin your clone of jev-gateway and confirm thesteerline: with thinking or caching on, you are buying a hint, not a route. - Run the gateway's bench on your own repo with
--user-toolsrather than trusting the chess tables; keep routing only if a bug-hunt task shows fewer requests with no solved/unsolved change. - Do not install fast-jev-compaction until issues #21, #88 and #89 are closed; check that
/hookslists more than zero hooks after install. - Read the vendor ToS before pasting a key: every routed request ships your tool roster and last message, and the compaction plugin ships full transcripts.
- If you also run Codex, test Jev there first: forced
tool_choiceis what the bench shows paying off.
Go further
- The forced tool use table per model, including which models return a 400 and which thinking modes block
anyandtools4. - The cache invalidation table:
toolsthensystemthenmessages, and thetool_choicerow that explains the gateway's design s5. - Issue #24 on jev-gateway: the session-invariant tool roster is re-sent to Jev on every request, the cost center the dashboard hides s14.
- The bench README's data-integrity section: 119 of 120 runs kept to themselves, one run marked
contaminatedinruns.jsonls2. - An independent read of Jev as a classifier or filter on public and private data, outside the coding-agent framing s12.
- Why the vendor's evals measure against two models rather than ground truth, and what that does to the headline multipliers s11.
- A third-party review of jev-gateway that walks the "Jev picks, the LLM writes" split and its localhost exposure, fixed the same day s8.
- The HN launch thread, where the pricing and subsidy question is argued in the open s10.
Sources
- jev-gateway, GitHub, vinilana. Why read it:
src/adapters/messages.tsholds the one line that decides hint versus forced tool, andbin/clients.mjsshows the launcher sets onlyANTHROPIC_BASE_URL. - jev-gateway-bench, GitHub, vinilana. Why read it: the full 120-session table and the authors' own reading, including the one contaminated run.
- Introducing System One Models & Jev, TypeSafe AI. Why read it: pricing, latency and the footnote that qualifies the 444.6x claim, from the vendor.
- Forcing tool use, Anthropic docs. Why read it: the per-model table of what
tool_choicevalues error out. - Prompt caching, Anthropic docs. Why read it: the invalidation hierarchy that forces the hint design.
- fast-jev-compaction, GitHub, tamaratran. Why read it: open the issues tab before the README.
- jev-gateway Review: Jev Picks, the LLM Writes, mrjev.com. Why read it: an outside walkthrough of the gateway's architecture.
- Instant Claude Code compaction is my favorite use of Jev so far, r/ClaudeCode. Why read it: the practitioner thread, with the ToS objection at the top.
- HN: Introducing System One Models and Jev, Hacker News. Why read it: the launch debate on pricing and sustainability.
- The Evals: Measured Against Two Models, Not Against Truth, novcog. Why read it: a critique of the evaluation method behind the vendor multipliers.
- Testing Jev on public and private data: classifier or filter, Aman Kumar. Why read it: an independent measurement outside coding agents.
- Skill suggestion cookbook, TypeSafe docs. Why read it: the only published roster-selection numbers, 16.8% to 7.3%.
- jev-gateway issue #24, GitHub. Why read it: the roster re-send cost nobody counts.
- Jev + Claude Code: compaction and a Sonnet 5 source check, jevmodel.ai. Why read it: a second look at the compaction claim with a Sonnet 5 check.
FAQ
Does jev-gateway ever force a tool inside Claude Code?
Only when the request has neither extended thinking nor cache_control blocks. Our captured first-turn requests had both, on a clean and on a full setup, so in practice the gateway hints.
Why does the gateway not just rewrite the tool descriptions to steer harder?
Modifying tool definitions invalidates the entire prompt cache, tools, system and messages. Appending a block to the last user message only touches the messages level, which is the cheapest place to put a hint.
Is Jev useless for coding, then?
No. The bench shows it paying on Codex, where the tool is forced and 76 to 100% of requests are steered, and on debugging tasks for every model. It is the "cheapest Claude Code" framing that the gateway's own numbers do not support.
Should I try fast-jev-compaction?
Wait until the hook registration issues (#21, #88) and the --resume issue (#89) are closed, and decide whether full transcripts leaving for a third-party API is acceptable in your repos.
AIDive