1 nota · última: 28 Oct 2025
The post argues that on-policy distillation combines on-policy sampling with dense teacher scoring, improving reasoning, personalization, and continual learning while using less compute than RL.