DeepSeek-V4-Pro is the flagship of the V4 generation: a 1.6-trillion-parameter MoE model with only 49 billion active parameters per token and a standard context window of one million tokens. New attention mechanisms with token compression (Compressed Sparse Attention and Heavily Compressed Attention, combined with DeepSeek Sparse Attention) make this massive context manageable. In benchmarks, V4-Pro performs on par with leading proprietary models—while being available under an open MIT license.
DeepSeek V4 Pro (Preview)
DeepSeek AI
April 2026
MIT (Open Source)
MoE Flagship, Dual Mode (Think/Non-Think)
1.6 billion total (49 billion active per token)
Compressed Sparse Attention + Heavily Compressed Attention + DSA, text-only
1,000,000 tokens
V4-Pro is a Frontier model with corresponding infrastructure requirements. We’ll work with you to evaluate its benefits, hardware, and operational concept.
DeepSeek-V4-Pro combines several attention mechanisms with token-wise compression: Compressed Sparse Attention and Heavily Compressed Attention work together with DeepSeek Sparse Attention to keep the 1-million-token context efficient.
The model is purely text-based (not multimodal) and, according to reports, also runs on Huawei Ascend hardware—a testament to the robustness of the infrastructure.
For organizations that want to run Frontier Quality and very long contexts themselves.
1-million-token context as the default
Frontier benchmarks under an open MIT license
Efficient CSA/HCA Attention
Broad agent and API compatibility
Enormous hardware and storage requirements
Not multimodal (text-only)
Marked as “Preview” – Please note the level of maturity
We’ll work with you to plan the cluster configuration for DeepSeek-V4-Pro and handle the deployment, quantization, and operation—entirely in Germany, if you wish.