DeepSeek V4's full paper reveals FP4 training that doubles QK speed without losing recall.

This week, DeepSeek released the full version of their V4 paper, expanding on the technical insights previewed in April.

3 min readMachine Learning

DeepSeek's V4 technical paper arrives at a pivotal moment for AI development, offering concrete evidence that efficiency gains need not come at the expense of capability. The introduction of FP4 quantization aware training represents a significant shift in how we think about model deployment, particularly when compared to earlier approaches like the 4-bit optimizations we explored with Chaperone-Thinking-LQ-1.0. This technique enables direct inference on FP4 weights while maintaining 99.7% recall on critical selection operations, effectively doubling computational throughput without meaningful quality degradation. For practitioners managing multi-agent systems where individual tasks spawn numerous model calls, these improvements compound dramatically, transforming what was once computationally prohibitive into routine execution.

The efficiency table tells a compelling story that extends beyond raw numbers. V4-Pro achieves just 27% of baseline FLOPs while reducing KV cache requirements to 10% of the V3.2 baseline, and V4-Flash pushes this further to 10% and 7% respectively. These aren't marginal improvements but fundamental restructurings of resource allocation that make large context windows genuinely practical for everyday applications. When we consider that our previous work with 4-bit quantization already demonstrated substantial memory savings, the FP4 approach suggests we're approaching a tipping point where high-performance models become accessible on modest hardware configurations.

Beyond computational efficiency, DeepSeek addresses the elephant in the room for trillion-parameter mixture-of-experts models: training instability. The documentation of anticipatory routing and SwiGLU clamping provides valuable insights for the broader research community grappling with similar divergence problems. Rather than treating these as proprietary secrets, the team shares concrete mechanisms—deliberately desynchronizing router updates and imposing hard limits on activation ranges—that others can adapt and refine. This transparency accelerates collective progress while establishing new baselines for robust MoE training. The generative reward model approach further streamlines development by unifying generation and evaluation within a single framework, reducing the traditional dependency on extensive human labeling while maintaining grounded reasoning capabilities.

Human evaluation results validate that these technical improvements translate to real-world utility. The 62.7% win rate against Gemini 3.1 Pro in Chinese writing tasks, combined with strong performance across white-collar professional tasks, suggests we're witnessing genuine capability advancement rather than benchmark-specific optimization. Perhaps most telling is the coding agent evaluation where 91% of users expressed confidence in V4-Pro as their default coding model—numbers that align with my own experience integrating these models into daily workflows. The question moving forward isn't whether FP4 quantization scales, but how quickly these efficiency gains propagate across the ecosystem and what new applications emerge when computational constraints shift from limiting factors to enablers.

From Machine Learning

DeepSeek dropped the full V4 paper this week. preview from april was 58 pages, this version adds a lot of technical depth.

FP4 quantization aware training. theyre running FP4 QAT directly in late stage training. MoE expert weights quantized to FP4 (the main gpu memory consumer). QK path in the CSA indexer uses FP4 activations. 2x speedup on QK selector with 99.7% recall preserved. inference runs directly on the FP4 weights.

Read the original at Machine Learning