Language Models2024advanced12 min read
DeepSeek-V3 Technical Report
التقرير التقني لـ DeepSeek-V3
DeepSeek-AI — arXiv
The problem
By late 2024, the leading large language models (GPT-4o, Claude 3.5 Sonnet) had set high quality bars but at enormous costs. Open-source models lagged behind their closed-source counterparts, and scaling up dense architectures hit diminishing returns in compute efficiency. The field needed an architecture that could match frontier quality while drastically reducing training cost and memory.
The contribution
DeepSeek-V3 introduces a 671B-parameter Mixture-of-Experts model that activates only 37B parameters per . It combines Multi-Head Latent Attention (MLA) for compressed with an auxiliary-loss-free load balancing strategy for MoE. A multi-token prediction objective strengthens performance and enables . FP8 mixed precision training and DualPipe parallelism reduce training cost to $5.576M on 2048 H800 GPUs — over ten times cheaper than comparable models — while matching or exceeding GPT-4o on standard benchmarks.
The impact
DeepSeek-V3 proved that open-source models can match closed-source frontier quality at a fraction of the cost. Its auxiliary-loss-free MoE balancing, MLA attention, and FP8 training became reference architectures for the field. Most importantly, it served as the base model for DeepSeek-R1, which pioneered reasoning via and became one of the most influential open-source models of 2025.
Imagine a hospital with 671 specialist doctors. When a patient arrives, a triage system picks the 37 most relevant specialists — an orthopedist, a neurologist, a radiologist — while the other 634 rest. Each specialist keeps compressed case notes rather than full transcripts, saving shelf space for thousands of simultaneous patients. The hospital outperforms competitors with far larger staffs because it routes smartly, stores efficiently, and never wastes a specialist's time.
DeepSeek-V3 is this hospital: 671 billion parameters, 37 billion active per patient (token), compressed memory, and intelligent routing — delivering frontier performance at a fraction of the cost.
Multi-Head Latent Attention: compressing the memory bottleneck
In standard , every token stores a separate key vector and value vector for each attention head. During inference, all these key-value pairs from past tokens must remain in GPU memory — this is the KV Cache, and it grows linearly with sequence length and number of heads. For a model with 128 heads generating a 128K context, the KV cache alone can consume hundreds of gigabytes.
(GQA) addressed this by sharing key-value heads across groups, cutting memory by a constant factor. But MLA goes further with a fundamentally different approach: low-rank joint compression. Instead of storing full key and value vectors for each head, MLA compresses the token's representation into a single small latent vector of dimension , where .
Think of it like this: instead of each specialist writing separate case notes for each department, all specialists share a single compressed summary. When a specialist needs the original details, they reconstruct them from the summary on the fly. The shelf space (KV cache) drops dramatically, while the quality of consultation (attention) stays the same.
There is one subtlety. encodes position information into keys and queries. But RoPE is position-dependent: it cannot be applied before compression because the compressed latent is position-agnostic. DeepSeek-V3 solves this by adding a separate small decoupled rotary key that carries only position information. The final key becomes a concatenation — content from the latent, position from the decoupled RoPE key. This keeps position encoding correct while preserving the compression benefit.
DeepSeekMoE: 671B parameters, 37B active
A dense Transformer activates every parameter for every token — wasteful for simple tokens like "the" that do not need the same compute as a complex math derivation. A model replaces each dense feed-forward layer with a collection of smaller expert networks and a that routes each token to only a few experts.
DeepSeek-V3 has 256 routed experts plus 1 shared expert per MoE layer. For each token, the gating function selects the top-8 experts, so only of total parameters are activated — roughly 5.5%. The shared expert processes every token and captures common knowledge, while the routed experts specialize in different domains or patterns.
Picture a conference room. One permanent moderator (the shared expert) always participates. For each topic, 8 of 256 specialists (the routed experts) are called in based on relevance. The rest stay available but do not consume meeting time.
The load balancing problem. In traditional MoE, some experts become "popular" and get overloaded while others sit idle — wasting parameters and creating compute bottlenecks. Prior work solved this with an auxiliary loss that penalizes imbalanced routing. But this loss competes with the primary language modeling objective: the model must choose between routing to the best expert and distributing load evenly. Even small auxiliary loss coefficients degrade model quality.
DeepSeek-V3 pioneers an auxiliary-loss-free approach. Instead of adding a loss term, the gating function uses a dynamic bias for each expert. If expert is receiving too many tokens, decreases, making that expert less likely to be selected. If it is underutilized, increases. The bias updates are detached from the — they steer routing without contaminating the training signal.
Multi-Token Prediction: looking ahead while learning
Standard language models predict one token at a time: given all previous tokens, predict the next. Multi-Token Prediction (MTP) extends this: at each position, the model predicts not just the next token but the next tokens simultaneously. In DeepSeek-V3, (one extra look-ahead token).
Why does this help? Predicting multiple future tokens forces the model to build richer representations of the current context. A model that must predict token from position needs to understand deeper structure than one that only predicts token . It is like a chess player who plans two moves ahead versus one — the two-move planner develops stronger positional understanding even for the immediate next move.
Architecturally, each additional prediction depth has its own MTP module containing a shared layer, a projection matrix, a single Transformer block, and an output head. The MTP module at depth takes the main model's representation at position and combines it with the embedding of the ground-truth token at position to predict token .
FP8 Training: halving precision, doubling speed
Most large models train in (16-bit brain floating-point). DeepSeek-V3 pushes precision lower: all linear-layer matrix multiplications use FP8 (8-bit floating-point), cutting compute and memory roughly in half. This is the first validated use of FP8 training at the 600B+ parameter scale.
The key challenge is that 8 bits provide much less numerical range and precision than 16. Naively using FP8 causes training instability — gradients overflow, activations quantize poorly, and the model diverges. DeepSeek-V3 solves this with three techniques:
First, : instead of one scaling factor per tensor, they use one per 128-element tile for activations and one per 128-element block for weights. This captures local variation that a single scale factor would miss.
Second, high-precision accumulation: FP8 multiplications accumulate into FP32 registers every 128 elements, preventing rounding errors from compounding.
Third, online : scaling factors are computed from the current batch statistics rather than from the previous batch, ensuring they always reflect the actual data distribution.
DualPipe: hiding communication behind computation
Training a 671B-parameter model across 2048 GPUs requires sophisticated parallelism. DeepSeek-V3 uses (splitting layers across GPUs), (distributing MoE experts across nodes), and (replicating across training batches) — but not , which is memory-expensive.
The DualPipe algorithm is the key innovation for pipeline parallelism. Standard pipeline parallelism has "bubbles" — idle GPU time when a pipeline stage waits for data from the previous stage. DualPipe minimizes bubbles by overlapping the of one micro-batch with the of a previous one, and crucially, overlapping All-to-All communication (needed for MoE expert dispatch) with computation.
Think of it like a factory assembly line where workers receive raw materials for the next product while finishing the current one. No one is ever idle, and material deliveries happen while hands are already busy.
The result: near-zero communication overhead. As long as the computation-to-communication ratio stays constant, DeepSeek-V3 can scale to larger models with more fine-grained experts across more nodes without paying a communication tax.
Pre-training: 14.8 trillion tokens, zero rollbacks
DeepSeek-V3 was pre-trained on 14.8 trillion high-quality and diverse tokens, primarily in English and Chinese, over approximately 55 days on a cluster of 2048 NVIDIA H800 GPUs. The training was remarkably stable — the team reported no irrecoverable loss spikes and no rollbacks throughout the entire run. For context, most frontier models encounter multiple loss spikes requiring restoration during training.
The training used a two-stage schedule: a warmup followed by cosine annealing from a peak learning rate of $2.2 \times 10^-4 to \2.2 \times 10^-5$. After , a two-stage context extension pushed the from 4K to 32K and then to 128K tokens, using only 119K additional GPU hours.
Total training cost: 2.788 million H800 GPU hours, or approximately $5.576 million at $2 per GPU hour. For comparison, GPT-4's training is estimated at $50–100 million. DeepSeek-V3 achieved comparable quality for roughly one-tenth the cost.
Post-training: aligning the model with human preferences
After pre-training, DeepSeek-V3 undergoes two post-training stages: (SFT) and Reinforcement Learning (RL).
For SFT, the team curated 1.5 million instruction-following examples spanning reasoning, coding, mathematics, creative writing, and general knowledge. A key innovation is the use of reasoning data distilled from DeepSeek-R1 — a separate reasoning-focused model that generates detailed chain-of-thought solutions. The team used DeepSeek-R1 to produce verified reasoning traces for math and code problems, then filtered for correctness. This transferred R1's reasoning patterns (including verification and reflection) into V3 without making V3's outputs unnecessarily long.
For RL, DeepSeek-V3 uses (GRPO) instead of standard . GRPO eliminates the need for a separate critic () network by estimating baselines from group scores within each batch. It uses two types of reward models: rule-based rewards for problems with verifiable answers (math, code) and model-based rewards for open-ended tasks (writing, general QA). This combination maintains a balance between accuracy and output style.
Results: matching GPT-4o at 1/10th the cost
DeepSeek-V3 achieves 88.5 on MMLU, 75.9 on MMLU-Pro, and 59.1 on GPQA — matching or exceeding GPT-4o and Claude 3.5 Sonnet on knowledge benchmarks. On math, it scores 90.2 on MATH-500, outperforming even o1-preview on this specific benchmark. On coding, it leads all models on LiveCodeBench and Codeforces, and achieves competitive scores on SWE-Bench Verified for software engineering tasks. In Chinese knowledge benchmarks, DeepSeek-V3 outperforms all models including GPT-4o, reflecting its strong bilingual training data.
Legacy: the foundation for DeepSeek-R1
2017
Transformer architecture (Vaswani et al.)
Introduced Multi-Head Attention and the encoder-decoder Transformer, establishing the foundation for all modern large language models.
2022
Mixtral and early MoE LLMs
Showed that Mixture-of-Experts could produce competitive LLMs with far fewer active parameters than dense models.
2024
DeepSeek-V2 — MLA + DeepSeekMoE
Validated MLA for KV cache compression and the fine-grained MoE architecture. Proved the combination could match larger dense models.
2024
DeepSeek-V3 (this paper)
Scaled to 671B parameters with auxiliary-loss-free MoE, multi-token prediction, and FP8 training. Matched GPT-4o at 1/10th the cost.
2025
DeepSeek-R1 — reasoning via RL
Built on DeepSeek-V3 as its base model, introducing reasoning through reinforcement learning without supervised chain-of-thought data. Became one of the most influential open-source models of 2025.
DeepSeek-V3 reshaped the economics of frontier AI. It demonstrated that careful architectural design — , compressed attention, hardware-aware parallelism — can deliver top-tier quality at a cost accessible to academic labs and smaller companies, not just tech giants. Its open-source release democratized access to frontier capability. And by serving as the foundation for DeepSeek-R1, it enabled the breakthrough that showed reasoning could emerge from reinforcement learning alone.
CitationDeepSeek-AI. DeepSeek-V3 Technical Report. arXiv, 2024.
Terms in this paper
- Mixture of Expertsمزيج الخبراء
- Multi-Head Attentionالانتباه المتعدد المسارات
- Grouped Query Attentionانتباه الاستعلام المُجمَّع
- KV Cacheذاكرة المفاتيح والقيم
- Gating Mechanismآلية البوابات
- Sparse Activationالتنشيط المتفرّق
- Knowledge Distillationتقطير المعرفة
- Reinforcement Learningالتعلم المعزز
- Supervised Fine-Tuningالضبط الدقيق الخاضع للإشراف
- Pipeline Parallelismالتوازي المتسلسل للطبقات
- Tensor Parallelismتوازي مصفوفات الموتّرات
- Data Parallelismتوازي البيانات
- Speculative Decodingفك الترميز التخميني
- Rotary Position Embedding (RoPE)التضمين الموضعي الدوار
- RMSNormتسوية الجذر التربيعي للمتوسط