3 articles
Deepseek
All posts tagged with #Deepseek
The Post-Training Era: Three AI Upgrades That Prove Bigger Isn't Better
Grok 4.6, Gemini 3.7 Flash, and DeepSeek V4-Pro all shipped this week without new base models. The performance gains came entirely from better post-training — and that tells you where model development is actually heading.
A 13-Billion-Parameter Model Just Beat Its Own Big Brother. Here's Why That Matters.
DeepSeek's V4-Flash-0731 activates only 13B parameters per token but outscored its own flagship V4-Pro on independent benchmarks. Here's the training recipe that made it happen, and what it means for your inference costs.
Distillation in 2026: Why the Best Teachers Are No Longer Bigger Models
Distillation isn't compression anymore — it's synthesis. Frontier labs now merge specialized RL checkpoints of the same base model via on-policy distillation, creating students that exceed their teachers without ever seeing a larger model.