AIDive

Video pack

DeepSeek Harness and V4 Pro: architecture notes, pricing grid, verdict table and sources

10 min read

TL;DR

  • DeepSeek Harness (dsh) is an MIT-licensed coding agent where models, tools, sandboxes, storage, loops and the UI are all plugins you swap in config. The repository already counts more than 21,000 stars and more than 12,000 commits.
  • The strongest idea is the append-only session log: you replay a derailed session event by event instead of scrolling a transcript.
  • DeepSeek V4 Pro 0813 claims about 80.6% on SWE-bench Verified against 80.8% for Claude Opus 4.6. Every score is self-reported; nobody independent has reproduced one yet.
  • Launch pricing is $0.435 per million input tokens and $0.87 output. From August 16 the API moves to peak and off-peak billing and output climbs as high as $3.96 at peak, still roughly four times cheaper than Opus 5.
  • dsh is a 0.1 developer preview whose README promises breaking changes. Tool builders should install it this week, cost cutters should test V4 Pro on a side project at the August 16 rates, daily Claude Code users should stay put.

What the sources say

Architecture

The whole harness rests on a kernel called Cordis, and the official page describes every capability as something you can select, swap or extend in configuration without touching the harness code s8. That covers more than Claude Code's skills and MCP servers: in Claude Code you extend the agent at its edges, in dsh you can recompose the sandbox, the session storage and the loop itself s1. A GitHub topic, dsh-plugin, already lists community plugins, including model plugins that let the harness run on models other than DeepSeek's s10.

Installation is one command: npx @deepseek-ai/dsh web starts a local web UI on port 3080, no account needed, and the same code base serves the terminal mode. Cloning and building the full repository takes four pnpm commands s2.

dsh ships four execution modes. Standard is the complete coding agent (file edits, shell, search, planning). Code makes the model write TypeScript that orchestrates multi-step operations itself instead of chaining tool calls one by one, which cuts API round trips on long tasks. Minimal strips the agent down to bash plus a file editor, mainly to benchmark the bare model. Creator is for building your own presets with live runtime inspection s8.

Everything the model sees goes into an append-only session log: prompts, reasoning, tool calls, context injections. That log is the harness's source of truth, so any session can be resumed, forked, searched or replayed from the event stream s1.

The model

V4 Pro 0813 is a mixture-of-experts model with a 1M-token context window, up to 384,000 output tokens, tool calls and JSON output s4. The DeepSeek API advertises Anthropic API compatibility, so a tool written for Claude can target V4 Pro by changing a base URL and a key s3.

DeepSeek's own numbers put the Max variant at about 80.6% on SWE-bench Verified against 80.8% for Claude Opus 4.6, with claimed wins on Terminal Bench 2.1 and several agent benchmarks against Opus 4.8. On the Artificial Analysis index V4 Pro sits at 53 while Opus 5 is at 63 and Fable 5 at 62. On speed it outputs 83 tokens per second where Opus 5 does 52 s5.

None of those scores has been reproduced by a third party. The benchmarks first circulated in DeepSeek's official WeChat group, then in a Reddit post the moderators removed, then as an ASCII table on Hacker News. There is no official announcement page and open weights are not confirmed, although the April V4 Pro and the July V4 Flash both ended up on Hugging Face. The one hands-on observation so far: three reasoning levels produced three radically different results on the same drawing prompt, something not seen on any other model s5.

Pricing

At launch V4 Pro costs $0.435 per million input tokens and $0.87 per million output tokens, and the 1M context is billed at the standard rate with no long-context surcharge s3. Claude Opus 5 is at $5 input and $25 output, Fable 5 at $10 and $50, Sonnet 5 at $2 and $10 s6. That makes V4 Pro about 11 times cheaper than Opus 5 on input and 28 times on output, and still 4 to 11 times cheaper than Sonnet 5 s3.

A typical agent session of 50,000 input tokens and 15,000 output tokens costs about $0.62 on Opus 5 and $0.035 on V4 Pro at launch rates. The hidden line for agents is the cache: a harness resends the same context every turn, so most input tokens are re-reads. A cache hit costs $0.50 per million on Opus 5 s6 and $0.0036 on DeepSeek, a ratio of 138 to one exactly where an agent spends the most s3.

From August 16, three days after launch, the API switches to peak and off-peak billing. Off-peak, V4 Pro rises to $0.66 input and $1.98 output. Peak, between 1h and 4h and between 6h and 10h UTC, it is $1.32 and $3.96 s3. Peak output is four and a half times the launch-day rate, a 355% increase. Even at the worst rate V4 Pro stays almost four times cheaper than Opus 5 s6. Peak hours match the Chinese working day, so cron jobs and European night runs can target off-peak, while live afternoon coding will hit some peak slots.

Reception

The Hacker News post passed 990 points in one day s7. Bloomberg framed the launch as a direct challenge to Claude Code, which is Anthropic's harness bound to Claude models s9.

Verdict

Piece Keep, try or skip Why
dsh plugin architecture Try Sandbox, storage and loop are all swappable; the most interesting harness design of the moment
dsh replayable session log Try Event-by-event replay beats reading a transcript when an agent derails
dsh as daily driver Skip for now 0.1 developer preview, README promises breaking changes, no production track record
V4 Pro on side projects Try 1M context, Anthropic API compatibility, $0.0036 cache hits
V4 Pro in production Wait Scores self-reported only, no announcement page, weights unconfirmed
Launch pricing in your spreadsheet Skip Expires August 16; model with $0.66/$1.98 off-peak and $1.32/$3.96 peak
Claude Code for daily work Keep Verified model scores, mature ecosystem, flat subscription

Do this Monday

  • Run npx @deepseek-ai/dsh web, open 127.0.0.1:3080 and spend thirty minutes in Standard mode on a throwaway repo.
  • Switch the same task to Code mode and count the API round trips against Standard mode.
  • Force a failure, then replay the session log to the event where it went wrong and compare that with reading a Claude Code transcript.
  • Point one existing Anthropic-compatible script at the DeepSeek base URL with a V4 Pro key and check tool calls and JSON output still work.
  • Rebuild your cost estimate with the August 16 rates: $0.66/$1.98 off-peak, $1.32/$3.96 peak, and mark which of your runs fall in 1h to 4h or 6h to 10h UTC.
  • Measure your cache-hit share on one real agent session before trusting the 138 to one cache ratio.
  • Watch the dsh-plugin GitHub topic for a model plugin that runs your current model inside dsh.
  • Keep Claude Code as the default and set a calendar reminder to re-evaluate once a third party publishes a SWE-bench Verified reproduction.

Go further

  • Read the quickstart and build dsh from source with pnpm to see how the web and terminal modes share one code base s2.
  • Study the Cordis kernel and the "Everything is a Plugin" section before writing a plugin, since the preset system is where the recomposition happens s8.
  • Follow the dsh-plugin topic to see which community plugins appear first: models, sandboxes or storage tell you where the ecosystem is heading s10.
  • Read the WeChat to Reddit to Hacker News chain of custody for the benchmark numbers before quoting the 80.6% figure anywhere s5.
  • Check the OpenRouter listing for the 384,000 output-token ceiling and the provider mix if you want V4 Pro without a direct DeepSeek account s4.
  • Compare the Claude cache pricing page with DeepSeek's cache hit line; the per-turn cache is where agent bills diverge most s6.
  • Scan the Hacker News thread for early plugin authors and the first reports of breaking changes between preview builds s7.

Sources

FAQ

Can I run dsh with a Claude model?

The harness accepts models other than DeepSeek's through model plugins, and the dsh-plugin topic already lists some. The two halves of the launch test separately: dsh with your current model, or your current harness with V4 Pro through the Anthropic-compatible API.

Is the 28x price gap real?

At launch rates, yes: $0.87 against $25 per million output tokens. From August 16 the gap shrinks to roughly 4x at peak hours, so budget on the later grid.

Why not trust the 80.6% SWE-bench score?

It is DeepSeek's own number, surfaced through a WeChat group and an ASCII table, with no announcement page and no independent reproduction so far. It may be accurate; nobody outside DeepSeek has checked.

What does the append-only log change in practice?

When an agent derails at tool call 42, you replay the session to that event and see exactly what the model had in context, instead of reconstructing it from a flat transcript.