DeepSeek Model Release Timeline: From DeepSeek-Coder to V4.1-Flash
A complete chronological guide to DeepSeek’s open-weights model releases, architectural milestones, and key developments.
DeepSeek-Coder
DeepSeek’s initial public release, specialized for code completion, generation, and repository-level tasks trained on 2 trillion tokens.
DeepSeek-LLM
The company’s first general-purpose dense language models, released in 7B and 67B parameter variants.
DeepSeek-MoE
Introduced DeepSeek’s initial Mixture-of-Experts (MoE) architecture, establishing high-efficiency routing for fine-grained expert specialization.
DeepSeek-Math
A dedicated model series fine-tuned specifically for mathematical reasoning using Group Relative Policy Optimization (GRPO).
DeepSeek-Prover
Designed specifically for formal theorem proving in the Lean 4 interactive theorem prover environment.
DeepSeek-V2 & DeepSeek-Coder-V2
Introduced Multi-Head Latent Attention (MLA) to drastically reduce KV cache memory footprints alongside enhanced coding and conversational capabilities.
DeepSeek-V2.5
Merged general conversation capabilities with code specialization into a single unified flagship model.
DeepSeek-R1-Lite-Preview
An early preview release showcasing chain-of-thought reasoning performance prior to full open-weights availability.
DeepSeek-VL2
Second-generation vision-language MoE models supporting dynamic resolution image understanding and document analysis.
DeepSeek-V3
Major release of the 671B parameter Mixture-of-Experts model featuring 37B active parameters per token, Multi-Token Prediction (MTP), and FP8 execution.
DeepSeek-R1 & Distilled Models
Full open-weights release of DeepSeek-R1, R1-Zero, and six distilled models leveraging Llama and Qwen bases for lightweight local deployment.
DeepSeek-Prover-V2
Scaled formal theorem-proving capabilities to 671B parameter sizes, combining RL search with formal verification engines.
DeepSeek-R1-0528
Updated checkpoint for DeepSeek-R1, enhancing math/coding precision, multi-turn reasoning consistency, function calling, and structured JSON output support.
DeepSeek-V3.1
Introduced a hybrid execution architecture capable of switching dynamically between thinking (reasoning) and non-thinking inference paths within a single endpoint.
DeepSeek-V3.2
Reasoning-first iteration introducing Sparse Attention optimizations and the reasoning-focused Speciale variant designed for high-compute agentic loops.
DeepSeek-V4 (Pro & Flash Previews)
Architectural leap introducing a 1-million-token context window, available in V4-Pro (1.6T parameter / 49B active) and V4-Flash (284B parameter / 18B active) configurations.
DeepSeek-V4.1-Flash
Integrated native visual understanding directly into the Flash model series via a causal encoder-decoder structure without requiring external vision heads.
