Language Models2024advanced12 min read

DeepSeek-V3 Technical Report

التقرير التقني لـ DeepSeek-V3

DeepSeek-AI — arXiv

The problem

By late 2024, the leading large language models (GPT-4o, Claude 3.5 Sonnet) had set high quality bars but at enormous costs. Open-source models lagged behind their closed-source counterparts, and scaling up dense architectures hit diminishing returns in compute efficiency. The field needed an architecture that could match frontier quality while drastically reducing training cost and memory.

The contribution

DeepSeek-V3 introduces a 671B-parameter Mixture-of-Experts model that activates only 37B parameters per . It combines Multi-Head Latent Attention (MLA) for compressed with an auxiliary-loss-free load balancing strategy for MoE. A multi-token prediction objective strengthens performance and enables . FP8 mixed precision training and DualPipe parallelism reduce training cost to $5.576M on 2048 H800 GPUs — over ten times cheaper than comparable models — while matching or exceeding GPT-4o on standard benchmarks.

The impact

DeepSeek-V3 proved that open-source models can match closed-source frontier quality at a fraction of the cost. Its auxiliary-loss-free MoE balancing, MLA attention, and FP8 training became reference architectures for the field. Most importantly, it served as the base model for DeepSeek-R1, which pioneered reasoning via and became one of the most influential open-source models of 2025.

Imagine a hospital with 671 specialist doctors. When a patient arrives, a triage system picks the 37 most relevant specialists — an orthopedist, a neurologist, a radiologist — while the other 634 rest. Each specialist keeps compressed case notes rather than full transcripts, saving shelf space for thousands of simultaneous patients. The hospital outperforms competitors with far larger staffs because it routes smartly, stores efficiently, and never wastes a specialist's time.

DeepSeek-V3 is this hospital: 671 billion parameters, 37 billion active per patient (token), compressed memory, and intelligent routing — delivering frontier performance at a fraction of the cost.

Multi-Head Latent Attention: compressing the memory bottleneck

In standard , every token stores a separate key vector and value vector for each attention head. During inference, all these key-value pairs from past tokens must remain in GPU memory — this is the KV Cache, and it grows linearly with sequence length and number of heads. For a model with 128 heads generating a 128K context, the KV cache alone can consume hundreds of gigabytes.

(GQA) addressed this by sharing key-value heads across groups, cutting memory by a constant factor. But MLA goes further with a fundamentally different approach: low-rank joint compression. Instead of storing full key and value vectors for each head, MLA compresses the token's representation into a single small latent vector ctKV\mathbf{c}_t^{KV} of dimension dcd_c, where dc≪dh⋅nhd_c \ll d_h \cdot n_h.

Think of it like this: instead of each specialist writing separate case notes for each department, all specialists share a single compressed summary. When a specialist needs the original details, they reconstruct them from the summary on the fly. The shelf space (KV cache) drops dramatically, while the quality of consultation (attention) stays the same.

ctKV=WDKV ht,ktC=WUK ctKV,vtC=WUV ctKV\mathbf{c}_t^{KV} = W^{DKV}\,\mathbf{h}_t, \quad \mathbf{k}_t^C = W^{UK}\,\mathbf{c}_t^{KV}, \quad \mathbf{v}_t^C = W^{UV}\,\mathbf{c}_t^{KV}
MLA low-rank KV compression — compress once, reconstruct on the fly — WDKVW^{DKV} projects the hidden state ht\mathbf{h}_t down to a compact latent vector ctKV\mathbf{c}_t^{KV}. During attention, WUKW^{UK} and WUVW^{UV} reconstruct the full keys and values from this latent. Only ctKV\mathbf{c}_t^{KV} is stored in cache — a fraction of what standard MHA or even GQA would require.

There is one subtlety. encodes position information into keys and queries. But RoPE is position-dependent: it cannot be applied before compression because the compressed latent is position-agnostic. DeepSeek-V3 solves this by adding a separate small decoupled rotary key ktR\mathbf{k}_t^R that carries only position information. The final key becomes a concatenation kt,i=[kt,iC;ktR]\mathbf{k}_{t,i} = [\mathbf{k}_{t,i}^C ; \mathbf{k}_t^R] — content from the latent, position from the decoupled RoPE key. This keeps position encoding correct while preserving the compression benefit.

Open in Lab
Compare KV cache memory for MHA, GQA, and MLA. Slide the sequence length to see how MLA's compressed latent keeps memory flat while others grow.
The demo wakes as you arrive…

DeepSeekMoE: 671B parameters, 37B active

A dense Transformer activates every parameter for every token — wasteful for simple tokens like "the" that do not need the same compute as a complex math derivation. A model replaces each dense feed-forward layer with a collection of smaller expert networks and a that routes each token to only a few experts.

DeepSeek-V3 has 256 routed experts plus 1 shared expert per MoE layer. For each token, the gating function selects the top-8 experts, so only 37B37B of 671B671B total parameters are activated — roughly 5.5%. The shared expert processes every token and captures common knowledge, while the routed experts specialize in different domains or patterns.

Picture a conference room. One permanent moderator (the shared expert) always participates. For each topic, 8 of 256 specialists (the routed experts) are called in based on relevance. The rest stay available but do not consume meeting time.

Open in Lab
Watch tokens get routed to different experts. Each token activates only 8 of 256 specialists. The shared expert processes every token.
The demo wakes as you arrive…

The load balancing problem. In traditional MoE, some experts become "popular" and get overloaded while others sit idle — wasting parameters and creating compute bottlenecks. Prior work solved this with an auxiliary loss that penalizes imbalanced routing. But this loss competes with the primary language modeling objective: the model must choose between routing to the best expert and distributing load evenly. Even small auxiliary loss coefficients degrade model quality.

DeepSeek-V3 pioneers an auxiliary-loss-free approach. Instead of adding a loss term, the gating function uses a dynamic bias bib_i for each expert. If expert ii is receiving too many tokens, bib_i decreases, making that expert less likely to be selected. If it is underutilized, bib_i increases. The bias updates are detached from the — they steer routing without contaminating the training signal.

gi,t=si,t⋅1(si,t+bi∈TopK)∑j=1Nrsj,t⋅1(sj,t+bj∈TopK),si,t=Sigmoid(ei⊤ht)g_{i,t} = \frac{s_{i,t} \cdot \mathbb{1}(s_{i,t} + b_i \in \text{TopK})} {\sum_{j=1}^{N_r} s_{j,t} \cdot \mathbb{1}(s_{j,t} + b_j \in \text{TopK})}, \quad s_{i,t} = \text{Sigmoid}(\mathbf{e}_i^\top \mathbf{h}_t)
Auxiliary-loss-free gating — bias steers routing without touching gradients — The affinity score si,ts_{i,t} uses Sigmoid (not Softmax), so experts are scored independently. The bias bib_i is added only for the TopK selection decision, not for the final gate value. This means the bias influences *which* experts are chosen but not *how much* weight they receive. The gradient flows cleanly through the Sigmoid scores.

Multi-Token Prediction: looking ahead while learning

Standard language models predict one token at a time: given all previous tokens, predict the next. Multi-Token Prediction (MTP) extends this: at each position, the model predicts not just the next token but the next DD tokens simultaneously. In DeepSeek-V3, D=2D = 2 (one extra look-ahead token).

Why does this help? Predicting multiple future tokens forces the model to build richer representations of the current context. A model that must predict token t+2t+2 from position tt needs to understand deeper structure than one that only predicts token t+1t+1. It is like a chess player who plans two moves ahead versus one — the two-move planner develops stronger positional understanding even for the immediate next move.

Architecturally, each additional prediction depth kk has its own MTP module containing a shared layer, a projection matrix, a single Transformer block, and an output head. The MTP module at depth kk takes the main model's representation at position tt and combines it with the embedding of the ground-truth token at position t+k−1t+k-1 to predict token t+kt+k.

Open in Lab
Standard next-token prediction vs Multi-Token Prediction. See how MTP forces the model to look further ahead, producing richer internal representations.
The demo wakes as you arrive…
LMTP=1D∑k=1D(λkT∑t=1TLCE ⁣(p^t(k),  xt+k))\mathcal{L}_{\text{MTP}} = \frac{1}{D} \sum_{k=1}^{D} \left( \frac{\lambda_k}{T} \sum_{t=1}^{T} \mathcal{L}_{\text{CE}}\!\big(\hat{p}_{t}^{(k)},\; x_{t+k}\big) \right)
MTP training loss — average cross-entropy across prediction depths — The final training objective adds the MTP loss to the standard next-token loss with a weighting factor λk\lambda_k. DeepSeek-V3 uses λ=0.3\lambda = 0.3 for the extra prediction depth. During inference, the MTP modules are discarded — they only improve training. Alternatively, they can be kept for speculative decoding: the extra predictions serve as draft tokens that are then verified by the main model, accelerating generation by 1.8× with no quality loss.

FP8 Training: halving precision, doubling speed

Most large models train in (16-bit brain floating-point). DeepSeek-V3 pushes precision lower: all linear-layer matrix multiplications use FP8 (8-bit floating-point), cutting compute and memory roughly in half. This is the first validated use of FP8 training at the 600B+ parameter scale.

The key challenge is that 8 bits provide much less numerical range and precision than 16. Naively using FP8 causes training instability — gradients overflow, activations quantize poorly, and the model diverges. DeepSeek-V3 solves this with three techniques:

First, : instead of one scaling factor per tensor, they use one per 128-element tile for activations and one per 128-element block for weights. This captures local variation that a single scale factor would miss.

Second, high-precision accumulation: FP8 multiplications accumulate into FP32 registers every 128 elements, preventing rounding errors from compounding.

Third, online : scaling factors are computed from the current batch statistics rather than from the previous batch, ensuring they always reflect the actual data distribution.

Open in Lab
Explore how FP8 represents numbers vs BF16. See how fine-grained quantization preserves the signal that naive FP8 would lose.
The demo wakes as you arrive…

DualPipe: hiding communication behind computation

Training a 671B-parameter model across 2048 GPUs requires sophisticated parallelism. DeepSeek-V3 uses (splitting layers across GPUs), (distributing MoE experts across nodes), and (replicating across training batches) — but not , which is memory-expensive.

The DualPipe algorithm is the key innovation for pipeline parallelism. Standard pipeline parallelism has "bubbles" — idle GPU time when a pipeline stage waits for data from the previous stage. DualPipe minimizes bubbles by overlapping the of one micro-batch with the of a previous one, and crucially, overlapping All-to-All communication (needed for MoE expert dispatch) with computation.

Think of it like a factory assembly line where workers receive raw materials for the next product while finishing the current one. No one is ever idle, and material deliveries happen while hands are already busy.

The result: near-zero communication overhead. As long as the computation-to-communication ratio stays constant, DeepSeek-V3 can scale to larger models with more fine-grained experts across more nodes without paying a communication tax.

Pre-training: 14.8 trillion tokens, zero rollbacks

DeepSeek-V3 was pre-trained on 14.8 trillion high-quality and diverse tokens, primarily in English and Chinese, over approximately 55 days on a cluster of 2048 NVIDIA H800 GPUs. The training was remarkably stable — the team reported no irrecoverable loss spikes and no rollbacks throughout the entire run. For context, most frontier models encounter multiple loss spikes requiring restoration during training.

The training used a two-stage schedule: a warmup followed by cosine annealing from a peak learning rate of $2.2 \times 10^-4 to \2.2 \times 10^-5$. After , a two-stage context extension pushed the from 4K to 32K and then to 128K tokens, using only 119K additional GPU hours.

Total training cost: 2.788 million H800 GPU hours, or approximately $5.576 million at $2 per GPU hour. For comparison, GPT-4's training is estimated at $50–100 million. DeepSeek-V3 achieved comparable quality for roughly one-tenth the cost.

Post-training: aligning the model with human preferences

After pre-training, DeepSeek-V3 undergoes two post-training stages: (SFT) and Reinforcement Learning (RL).

For SFT, the team curated 1.5 million instruction-following examples spanning reasoning, coding, mathematics, creative writing, and general knowledge. A key innovation is the use of reasoning data distilled from DeepSeek-R1 — a separate reasoning-focused model that generates detailed chain-of-thought solutions. The team used DeepSeek-R1 to produce verified reasoning traces for math and code problems, then filtered for correctness. This transferred R1's reasoning patterns (including verification and reflection) into V3 without making V3's outputs unnecessarily long.

For RL, DeepSeek-V3 uses (GRPO) instead of standard . GRPO eliminates the need for a separate critic () network by estimating baselines from group scores within each batch. It uses two types of reward models: rule-based rewards for problems with verifiable answers (math, code) and model-based rewards for open-ended tasks (writing, general QA). This combination maintains a balance between accuracy and output style.

Results: matching GPT-4o at 1/10th the cost

DeepSeek-V3 achieves 88.5 on MMLU, 75.9 on MMLU-Pro, and 59.1 on GPQA — matching or exceeding GPT-4o and Claude 3.5 Sonnet on knowledge benchmarks. On math, it scores 90.2 on MATH-500, outperforming even o1-preview on this specific benchmark. On coding, it leads all models on LiveCodeBench and Codeforces, and achieves competitive scores on SWE-Bench Verified for software engineering tasks. In Chinese knowledge benchmarks, DeepSeek-V3 outperforms all models including GPT-4o, reflecting its strong bilingual training data.

Open in Lab
Benchmark comparison across knowledge, math, code, and Chinese tasks. DeepSeek-V3 matches frontier closed-source models while being fully open-source.
The demo wakes as you arrive…

Legacy: the foundation for DeepSeek-R1

  1. 2017

    Transformer architecture (Vaswani et al.)

    Introduced Multi-Head Attention and the encoder-decoder Transformer, establishing the foundation for all modern large language models.

  2. 2022

    Mixtral and early MoE LLMs

    Showed that Mixture-of-Experts could produce competitive LLMs with far fewer active parameters than dense models.

  3. 2024

    DeepSeek-V2 — MLA + DeepSeekMoE

    Validated MLA for KV cache compression and the fine-grained MoE architecture. Proved the combination could match larger dense models.

  4. 2024

    DeepSeek-V3 (this paper)

    Scaled to 671B parameters with auxiliary-loss-free MoE, multi-token prediction, and FP8 training. Matched GPT-4o at 1/10th the cost.

  5. 2025

    DeepSeek-R1 — reasoning via RL

    Built on DeepSeek-V3 as its base model, introducing reasoning through reinforcement learning without supervised chain-of-thought data. Became one of the most influential open-source models of 2025.

DeepSeek-V3 reshaped the economics of frontier AI. It demonstrated that careful architectural design — , compressed attention, hardware-aware parallelism — can deliver top-tier quality at a cost accessible to academic labs and smaller companies, not just tech giants. Its open-source release democratized access to frontier capability. And by serving as the foundation for DeepSeek-R1, it enabled the breakthrough that showed reasoning could emerge from reinforcement learning alone.

CitationDeepSeek-AI. DeepSeek-V3 Technical Report. arXiv, 2024.

Terms in this paper