Page last updated on
DeepSeek-V4-Pro-0813
By DeepSeek
The official DeepSeek-V4-Pro build. A 1.6T MoE with 49B active parameters and 1M context, retrained for agents and coding, with 87.9 on Terminal Bench 2.1.
Specifications
- Total parameters
- 1.6T (49B active)
- Active parameters
- 49B
- Architecture
- MoE
- Architecture class
- DeepseekV4ForCausalLM
- Attention
- Compressed Sparse Attention (CSA) + Heavily Compressed Attention (HCA)
- Context window
- 1M
- Vocab size
- 129,280
- Modality
- Text
- Precision
- FP4 + FP8 Mixed
- License
- MIT
- Released
- August 2026
- Recommended hardware
- 8× H2008× B200
- Best for
- Complex agent workflows, coding, security research, and long-context analysis
Good to know
-
DeepSeek-V4-Pro-0813 is the official release of DeepSeek-V4-Pro, superseding the April 2026 preview. It went live on the API on August 12, 2026, and DeepSeek published the release notes on August 13.
-
DeepSeek open-sourced an agent harness alongside this model, DeepSeek Harness, which is MIT-licensed. It is the framework DeepSeek used to produce the code-agent benchmark scores on this page. Note that version 0.1 is a developer preview and the maintainers warn of compatibility-breaking changes.
-
This build natively supports the OpenAI Responses API format with a one-click Codex setup, matching what V4-Flash-0731 shipped in July.
-
The API price increase takes effect at 16:00 UTC on August 16, 2026, and it is large. See the pricing section for the peak and off-peak tables.
Architecture
DeepSeek-V4-Pro-0813 keeps the V4 architecture unchanged and rebuilds the post-training, the same playbook that produced DeepSeek-V4-Flash-0731 two weeks earlier.
Two things are specific to this checkpoint:
- DSpark speculative decoding ships attached. The released weights bundle the module. Both vLLM and SGLang can enable it.
- Selectable reasoning effort. The
reasoning_effortparameter takes low, high, or max. DeepSeek suggests temperature 1.0 with top_p 0.95 for agentic work, and allowing up to 384K output tokens at the higher settings.
For more models like this one, browse the full open source LLM directory or read the open source LLM ecosystem statistics.
Benchmarks
DeepSeek published eleven agent and coding benchmarks for the 0813 build, and the jump over the April preview is the story. One detail matters for reading these numbers: the code-agent rows were not run on a bare API. DeepSeek evaluated them inside DeepSeek Harness in minimal mode, a two-tool setup limited to bash and a text editor, at max reasoning effort.
DS-V4-Pro-0813 vs frontier models
Higher is better on every benchmark shown. Scores marked n/a were not reported. The Fable-5 column is the w/ fallback configuration. The V4-Flash Preview column from the model card is omitted here for readability.
Each model keeps the same color across every benchmark, stacked top to bottom in this order. Hover or focus a row to read all six scores.
- Opus-4.8
- Fable-5
- Kimi K3
- GLM-5.2
- V4-Pro Preview
- V4-Flash-0731
- DS-V4-Pro-0813
View all scores as a table
| Benchmark | Opus-4.8 | Fable-5 | Kimi K3 | GLM-5.2 | V4-Pro Preview | V4-Flash-0731 | DS-V4-Pro-0813 |
|---|---|---|---|---|---|---|---|
| Reasoning | |||||||
| HLE (wo tools) | 49.8 | 53.3 | 43.5 | 40.5 | 37.7 | 37.8 | 42.7 |
| HLE (w tools) | 57.9 | 63.0 | 56.0 | 54.7 | 48.2 | 51.5 | 60.0 |
| Agentic & Coding | |||||||
| Terminal Bench 2.1 (Acc) | 85.0 | 88.0 | 88.3 | 81.0 | 72.1 | 82.7 | 87.9 |
| NL2Repo (Pass@1) | 69.7 | n/a | n/a | 48.9 | 38.5 | 54.2 | 61.5 |
| Cybergym (Pass@1) | 78.3 | 83.1 | 80.0 | n/a | 52.7 | 76.7 | 83.3 |
| DeepSWE (Resolved) | 58.0 | 70.0 | 67.5 | 46.2 | 12.8 | 54.4 | 62.7 |
| Toolathlon-Verified (Pass@1) | 76.2 | 77.9 | 76.5 | 59.9 | 55.9 | 70.3 | 74.1 |
| Agents' Last Exam (Pass@1) | 25.7 | n/a | 27.6 | 23.8 | 16.5 | 25.2 | 25.7 |
| AutomationBench Public (Pass@1) | 27.2 | 29.1 | 30.8 | 12.9 | 12.8 | 25.1 | 31.8 |
| DSBench-FullStack (Pass@1) | 71.6 | 77.2 | 73.7 | 61.8 | 41.8 | 68.7 | 71.1 |
| DSBench-Hard (Pass@1) | 71.7 | 68.3 | 63.0 | 54.5 | 31.1 | 59.6 | 67.2 |
Source: the DS-V4-Pro-0813 model card. Best score in each row is marked.
Third-party evaluations
Independent numbers arrived within a day of the API rollout, and they broadly confirm the direction DeepSeek reported. The model lands in the top tier of open weights, it is strong on code, and the price is what puts it on the frontier.
API pricing
The price increase DeepSeek warned about previously landed alongside this release. From 16:00 UTC on August 16, 2026, the flat rate is gone and a peak and off-peak structure replaces it. Peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC. Every other hour is off-peak, at half the peak rate. The table below shows both tiers.
| Model | Input · cache hit | Input · cache miss | Output |
|---|---|---|---|
| DeepSeek-V4-Pro-0813 | $0.044 | $1.32 | $3.96 |
| DeepSeek-V4-Pro-0813, off-peak | $0.022 | $0.66 | $1.98 |
| DeepSeek-V4-Flash-0731 | $0.014 | $0.44 | $1.32 |
| DeepSeek-V4-Flash-0731, off-peak | $0.007 | $0.22 | $0.66 |
Per 1M tokens, as of August 16, 2026. Official pricing.
The jump is steep. V4-Pro output went from $0.87 per million tokens to $3.96 at peak, so even the off-peak rate of $1.98 is more than double the old flat price. Cache-hit input moved from $0.003625 to $0.044 at peak, roughly a twelve-fold increase. V4-Flash took the same treatment, from $0.28 output to $1.32 at peak and $0.66 off-peak.
Even after the increase, the model is inexpensive against the closed frontier. Off-peak output at $1.98 per million tokens is about 92% below Claude Opus 4.8 at $25, and roughly 93% below GPT-5.6 Sol at $30. Arena measured a blended $0.76 per million tokens at the old rates and placed the model on the cost-performance frontier. What changed is the internal comparison: V4-Flash-0731 off-peak output at $0.66 now costs a third of V4-Pro, so the choice between the two models is a real budget decision rather than a rounding error.
Two practical consequences. First, scheduling matters now. A batch job moved out of the two peak windows costs half as much for the same tokens. Second, prompt caching matters more than it did: cache-hit input at $0.022 off-peak is 30 times cheaper than a cache miss, so reusing system prompts, tool schemas, and repository context is where the savings are.
Sources: DeepSeek pricing · DeepSeek V4-Pro release notes · Claude pricing · OpenAI pricing
What people are saying
“We ran DeepSeek v4 Pro 0813 on our cybersecurity benchmark, it outperformed EVERY (!) other model at finding vulnerabilities - At pass@3, it rediscovered 87.5% of the benchmark CVEs. Far above Opus 5 and Qwen 3.8 at 81.3% - The tradeoff is precision.”
Philippe Dourassov
@pilvar222
X Aug 13, 2026
“DeepSeek V4 Pro 0813 delivers the same high-end performance as Pro Preview, at 29% lower cost. we benchmarked Pro 0813, Pro Preview, and Flash 0731 across 100 deep-research questions. 0813 was strongest on: Law: 100% Academic: 83% Medicine: 67%”
GMI Cloud
@gmi_cloud
X Aug 13, 2026
“DeepSeek V4 Pro 0813 is live on OpenRouter. @deepseek_ai reports large agent gains over V4 Pro Preview: DeepSWE 62.7 (+49.9), CyberGym 83.3 (+30.6), NL2Repo 61.5 (+23.0), and Terminal Bench 2.1 87.9 (+15.8) More providers coming online soon”
OpenRouter
@OpenRouter
X Aug 12, 2026
“Deepseek V4 Pro GA scoring only 1 point higher than Flash on Artificial Analysis while being like 5 times larger is crazy work”
Lentils
@Lentils80
X Aug 13, 2026
“We ran DeepSeek v4 Pro 0813 on our cybersecurity benchmark, it outperformed EVERY (!) other model at finding vulnerabilities - At pass@3, it rediscovered 87.5% of the benchmark CVEs. Far above Opus 5 and Qwen 3.8 at 81.3% - The tradeoff is precision.”
Philippe Dourassov
@pilvar222
X Aug 13, 2026
“DeepSeek V4 Pro 0813 delivers the same high-end performance as Pro Preview, at 29% lower cost. we benchmarked Pro 0813, Pro Preview, and Flash 0731 across 100 deep-research questions. 0813 was strongest on: Law: 100% Academic: 83% Medicine: 67%”
GMI Cloud
@gmi_cloud
X Aug 13, 2026
“DeepSeek V4 Pro 0813 is live on OpenRouter. @deepseek_ai reports large agent gains over V4 Pro Preview: DeepSWE 62.7 (+49.9), CyberGym 83.3 (+30.6), NL2Repo 61.5 (+23.0), and Terminal Bench 2.1 87.9 (+15.8) More providers coming online soon”
OpenRouter
@OpenRouter
X Aug 12, 2026
“Deepseek V4 Pro GA scoring only 1 point higher than Flash on Artificial Analysis while being like 5 times larger is crazy work”
Lentils
@Lentils80
X Aug 13, 2026
Run it yourself
vllm serve deepseek-ai/DeepSeek-V4-Pro-0813 sglang serve --model-path deepseek-ai/DeepSeek-V4-Pro-0813 Frequently asked questions
Is DeepSeek-V4-Pro-0813 open source?
Yes. The weights ship under the MIT license on Hugging Face, which allows unrestricted commercial use, modification, and redistribution. The API went live on August 12, 2026 and the weights followed shortly after.
What is DeepSeek Harness?
DeepSeek Harness, or dsh, is an open-source agent harness from DeepSeek, released under the MIT license alongside DeepSeek-V4-Pro-0813. It turns a language model into a coding agent, and it is built so that every capability is a plugin, including models, tools, skills, sessions, sandboxes, and storage. It ships four modes: standard for full coding work, code for letting the model orchestrate multiple rounds of tool calls, minimal for benchmarking with just bash and a text editor, and creator for building custom presets. Every run is written to an append-only session log that you can resume, fork, search, and replay. Version 0.1 is a developer preview, and the maintainers warn of compatibility-breaking changes.
How much does DeepSeek-V4-Pro-0813 cost?
From 16:00 UTC on August 16, 2026 the API uses peak and off-peak rates. Peak is $0.044 per million cache-hit input tokens, $1.32 per million cache-miss input tokens, and $3.96 per million output tokens. Off-peak is half of each: $0.022, $0.66, and $1.98. Peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC, and every other hour is off-peak. The previous flat rate was $0.87 per million output tokens.
Is DeepSeek-V4-Pro-0813 better than DeepSeek-V4-Flash-0731?
On the eleven benchmarks DeepSeek published, yes, on all of them. Terminal Bench 2.1 is 87.9 against 82.7, DeepSWE 62.7 against 54.4, and Cybergym 83.3 against 76.7. The gaps are meaningful but not huge, and V4-Pro now costs three times more per output token, so Flash remains the better default for high-volume work.
How does DeepSeek-V4-Pro-0813 compare to Claude Opus 4.8 and Fable 5?
It trades wins. DeepSeek reports 87.9 on Terminal Bench 2.1 against 85.0 for Opus 4.8 and 88.0 for Fable 5, and 83.3 on Cybergym against 78.3 and 83.1. Opus 4.8 keeps a clear lead on NL2Repo, 69.7 against 61.5, and on DSBench-Hard, 71.7 against 67.2. Fable 5 leads on DeepSWE, 70.0 against 62.7. All of these are vendor-reported numbers from the DeepSeek model card.
Keep exploring
vllm.ai
DeepSeek V4 in vLLM: Efficient Long-context Attention
lmsys.org
DeepSeek-V4 on Day 0: From Fast Inference to Verified RL with SGLang and Miles
developer.nvidia.com
Build with DeepSeek V4 Using NVIDIA Blackwell and GPU-Accelerated Endpoints
inferencex.semianalysis.com
MI355X DeepSeek-V4-Pro on SGLang: 110.5x Throughput per GPU in 26 Days
From the blog
Is DeepSeek Better Than ChatGPT? 25+ DeepSeek Stats (2026)
From the blog