AI competition is shifting toward what happens after pretraining.
Pretraining remains the most compute-intensive stage of model development. Chinese labs faced tighter compute constraints, so they pushed more capability gains into post-training through reinforcement learning, distillation and better training data.
DeepSeek-R1 showed that models could learn to reason by combining GRPO with rewards that were simple to verify. Because GRPO avoids using a separate model to score every answer, the training process is simpler and cheaper.
ByteDance, Qwen and MiniMax then adapted the approach to reduce repetitive answers, improve training stability and make better use of feedback. GRPO became the starting point for a broader post-training toolkit.
The same post-training push is now extending beyond short reasoning problems. Kimi K3 breaks long agentic episodes into resumable chunks and rewards models for staying within compute budgets. Where outcomes can be checked directly, verifiable environments can replace costly human preference labels.
Distillation can then transfer reasoning capabilities from a large model to smaller models capable of running on phones, vehicles and robots. On-policy distillation goes further by having the student generate its own attempts and learn from token-level feedback rather than copying fixed answers.
Together, these advances are changing the compute and data mix. More post-training compute goes toward generating rollouts, executing environments and verifying outcomes. Generic training text becomes easier to replace, while novel real-world data and high-quality verifiable environments become more valuable.
免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。