#Reinforcement-Learning
2 notas · última:
What the hell happened with AGI timelines in 2025?
The episode explains why AGI timelines shortened in early 2025 and then moved back out later in the year, pointing to limits in reasoning generalization, inference scaling, reinforcement learning, and autonomy.
On-Policy Distillation
The post argues that on-policy distillation combines on-policy sampling with dense teacher scoring, improving reasoning, personalization, and continual learning while using less compute than RL.
