DeepSeek-V4-Pro
By DeepSeek · April 2026
Page last updated on
Flagship MoE model with 1.6T total parameters and 1M context. Built for advanced reasoning, coding, and long-horizon agent workflows.
Specifications
- Total parameters
- 1.6T (49B active)
- Active parameters
- 49B
- Architecture
- MoE
- Architecture class
- DeepseekV4ForCausalLM
- Attention
- Compressed Sparse Attention (CSA) + Heavily Compressed Attention (HCA)
- Context window
- 1M
- Vocab size
- 129,280
- Modality
- Text
- Precision
- FP4 + FP8 Mixed
- License
- MIT
- Released
- April 2026
- Recommended hardware
- 8× H2008× B200
- Best for
- Advanced reasoning, coding, and long-horizon agent workflows
Good to know
-
This page covers the DeepSeek-V4 preview, released on April 24, 2026. It is an early open-weight build. The official V4 release is expected around late July 2026. Benchmarks, pricing, and behavior may still change in the final version.
Architecture
The defining choice of DeepSeek-V4-Pro is the attention stack. It interleaves Compressed Sparse Attention (CSA) with Heavily Compressed Attention (HCA).
- CSA compresses blocks of the KV cache and attends to only a few relevant entries.
- HCA squeezes much larger spans into single entries. That way the model keeps a full 1M-token history within reach, without paying dense-attention memory for it.
The payoff is efficiency at long context. At 1M tokens, DeepSeek reports that V4-Pro uses 27% of the per-token inference FLOPs and 10% of the KV-cache memory of DeepSeek-V3.2. The model also adds manifold-constrained hyper-connections (mHC) for cleaner residual flow. It trains with the Muon optimizer on more than 32T tokens. The released checkpoint uses mixed-precision quantization: MoE expert weights in FP4, and most of the rest in FP8.
For more models like this one, browse the full open source LLM directory or read the open source LLM ecosystem statistics.
Benchmarks
DeepSeek published a full benchmark set for V4-Pro. It is strongest at coding and tool use. It leads on LiveCodeBench (93.5%) and Codeforces (3206). It sits within a point of the best score on SWE-bench Verified (80.6%). On broad knowledge tests like GPQA Diamond and HLE, it gives a little ground to the closed frontier models.
DS-V4-Pro Max vs frontier models
Scores published by DeepSeek in the V4-Pro model card. DS-V4-Pro Max runs in Think Max mode. Higher is better on every benchmark shown. Scores marked n/a were not reported.
Each model keeps the same color across every benchmark, stacked top to bottom in this order. Hover or focus a row to read all six scores.
Separate axis. Rating is not a percentage, so this row is plotted as points on a fitted scale starting at 2,900.
Separate axis. Elo is not a percentage, so this row is plotted as points on a fitted scale starting at 1,000.
- Opus-4.6 Max
- GPT-5.4 xHigh
- Gemini-3.1-Pro High
- K2.6 Thinking
- GLM-5.1 Thinking
- DS-V4-Pro Max
View all scores as a table
| Benchmark | Opus-4.6 Max | GPT-5.4 xHigh | Gemini-3.1-Pro High | K2.6 Thinking | GLM-5.1 Thinking | DS-V4-Pro Max |
|---|---|---|---|---|---|---|
| Knowledge & Reasoning | ||||||
| MMLU-Pro (EM) | 89.1 | 87.5 | 91.0 | 87.1 | 86.0 | 87.5 |
| SimpleQA-Verified (Pass@1) | 46.2 | 45.3 | 75.6 | 36.9 | 38.1 | 57.9 |
| Chinese-SimpleQA (Pass@1) | 76.4 | 76.8 | 85.9 | 75.9 | 75.0 | 84.4 |
| GPQA Diamond (Pass@1) | 91.3 | 93.0 | 94.3 | 90.5 | 86.2 | 90.1 |
| HLE (Pass@1) | 40.0 | 39.8 | 44.4 | 36.4 | 34.7 | 37.7 |
| LiveCodeBench (Pass@1) | 88.8 | n/a | 91.7 | 89.6 | n/a | 93.5 |
| HMMT 2026 Feb (Pass@1) | 96.2 | 97.7 | 94.7 | 92.7 | 89.4 | 95.2 |
| IMOAnswerBench (Pass@1) | 75.3 | 91.4 | 81.0 | 86.0 | 83.8 | 89.8 |
| Apex (Pass@1) | 34.5 | 54.1 | 60.9 | 24.0 | 11.5 | 38.3 |
| Apex Shortlist (Pass@1) | 85.9 | 78.1 | 89.1 | 75.5 | 72.4 | 90.2 |
| Codeforces (Rating) | n/a | 3,168 | 3,052 | n/a | n/a | 3,206 |
| Long Context | ||||||
| MRCR 1M (MMR) | 92.9 | n/a | 76.3 | n/a | n/a | 83.5 |
| CorpusQA 1M (ACC) | 71.7 | n/a | 53.8 | n/a | n/a | 62.0 |
| Agentic | ||||||
| Terminal Bench 2.0 (Acc) | 65.4 | 75.1 | 68.5 | 66.7 | 63.5 | 67.9 |
| SWE Verified (Resolved) | 80.8 | n/a | 80.6 | 80.2 | n/a | 80.6 |
| SWE Pro (Resolved) | 57.3 | 57.7 | 54.2 | 58.6 | 58.4 | 55.4 |
| SWE Multilingual (Resolved) | 77.5 | n/a | n/a | 76.7 | 73.3 | 76.2 |
| BrowseComp (Pass@1) | 83.7 | 82.7 | 85.9 | 83.2 | 79.3 | 83.4 |
| HLE w/ tools (Pass@1) | 53.1 | 52.0 | 51.6 | 54.0 | 50.4 | 48.2 |
| MCPAtlas Public (Pass@1) | 73.8 | 67.2 | 69.2 | 66.6 | 71.8 | 73.6 |
| Toolathlon (Pass@1) | 47.2 | 54.6 | 48.8 | 50.0 | 40.7 | 51.8 |
| GDPval-AA (Elo) | 1,619 | 1,674 | 1,314 | 1,482 | 1,535 | 1,554 |
Source: the DS-V4-Pro model card. Best score in each row is marked.
Third-party evaluations
The first third-party evaluations landed right after the open-source preview. They line up with what DeepSeek reported. V4 sits in the top tier of open models, and it is strongest at coding.
API pricing
Price is where V4 pulls ahead. V4-Pro serves the same 1M-token context and up to 384K output tokens as the frontier models. Output costs $0.87 per million tokens. That is roughly 94% below GPT-5.4 at $15, and 97% below Claude Opus 4.8 at $25.
| Model | Input · cache hit | Input · cache miss | Output |
|---|---|---|---|
| DeepSeek-V4-Pro | $0.003625 | $0.435 | $0.87 |
| DeepSeek-V4-Flash | $0.0028 | $0.14 | $0.28 |
Per 1M tokens, as of July 21, 2026. Official pricing.
V4-Flash goes lower still, at $0.28 per million output tokens. That is about 99% cheaper than Opus 4.8, and among the least expensive models at any scale. Prompt caching drops V4-Pro input to $0.003625 per million tokens on a cache hit. Workloads that reuse a system prompt, tool schema, or repository context pay a small fraction of the cache-miss price.
Sources: DeepSeek pricing · Claude pricing · OpenAI pricing
What people are saying
“🚀 DeepSeek-V4 Preview is officially live & open-sourced! Welcome to the era of cost-effective 1M context length. DeepSeek-V4-Pro: 1.6T total / 49B active params. Performance rivaling the world's top closed-source models.”
DeepSeek
@deepseek_ai
X Apr 24, 2026
“IMO DeepSeek v4 demonstrated utter confidence and competence by not benchmaxxing... just showed up, demonstrated SOTA long context efficiency techniques (CSA, HCA, mHC, flash at 8% cost of pro...), dropped the best open base models in the world, peaced out. BYO posttraining... bravo.”
swyx
@swyx · Latent Space
X Apr 2026
“A few more notes on DeepSeek-V4: it seems to be a ~GPT-5.2/Opus 4.5+ tier model... at 1.6T params they now have a model that's in the same weight class as GPT-5.4... the technical paper is a big deal... the arch is very efficient for long context.”
Lisan al Gaib
@scaling01
X Apr 24, 2026
“DeepSeek v4 is out, and it's groundbreaking... it's the first open model to have solved long context... beats Opus 4.6 on LiveCodeBench... tops GPT-5.4 on Codeforces... the open landscape is competitive across the board, exciting times!”
Merve Noyan
@mervenoyann · Hugging Face
X Apr 2026
“🚀 DeepSeek-V4 Preview is officially live & open-sourced! Welcome to the era of cost-effective 1M context length. DeepSeek-V4-Pro: 1.6T total / 49B active params. Performance rivaling the world's top closed-source models.”
DeepSeek
@deepseek_ai
X Apr 24, 2026
“IMO DeepSeek v4 demonstrated utter confidence and competence by not benchmaxxing... just showed up, demonstrated SOTA long context efficiency techniques (CSA, HCA, mHC, flash at 8% cost of pro...), dropped the best open base models in the world, peaced out. BYO posttraining... bravo.”
swyx
@swyx · Latent Space
X Apr 2026
“A few more notes on DeepSeek-V4: it seems to be a ~GPT-5.2/Opus 4.5+ tier model... at 1.6T params they now have a model that's in the same weight class as GPT-5.4... the technical paper is a big deal... the arch is very efficient for long context.”
Lisan al Gaib
@scaling01
X Apr 24, 2026
“DeepSeek v4 is out, and it's groundbreaking... it's the first open model to have solved long context... beats Opus 4.6 on LiveCodeBench... tops GPT-5.4 on Codeforces... the open landscape is competitive across the board, exciting times!”
Merve Noyan
@mervenoyann · Hugging Face
X Apr 2026
Run it yourself
vllm serve deepseek-ai/DeepSeek-V4-Pro python3 -m sglang.launch_server --model-path deepseek-ai/DeepSeek-V4-Pro Frequently asked questions
Is DeepSeek-V4-Pro open source?
Yes. DeepSeek-V4-Pro ships under the MIT license, with open weights available on Hugging Face. You can use it commercially, fine-tune it, and redistribute it without restrictions.
What hardware do I need to run DeepSeek-V4-Pro?
The recommended setups are 8× H200 or 8× B200 GPU nodes. Mixed FP4 + FP8 precision keeps memory requirements manageable despite the 1.6T total parameter count.
What is the difference between DeepSeek-V4-Pro and DeepSeek-V4-Flash?
V4-Pro is the flagship, with 1.6T total parameters and 49B active per token. V4-Flash is a lighter sibling at 284B total parameters and 13B active, aimed at cost-sensitive workloads. Both support a 1M token context window and ship under the MIT license.
How does DeepSeek-V4-Pro perform on coding benchmarks?
Launch coverage reports 80.6% on SWE-bench Verified, the strongest open-weights result at release and roughly on par with leading closed models.
Keep exploring
vllm.ai
DeepSeek V4 in vLLM: Efficient Long-context Attention
lmsys.org
DeepSeek-V4 on Day 0: From Fast Inference to Verified RL with SGLang and Miles
newsletter.semianalysis.com
DeepSeekV4 1.6T Day 0 to Day 43 Performance Over Time
From the blog