This paper develops a method to improve reasoning models by reducing the discrepancy between a stronger teacher and an on-policy student, which helps to prevent the student from learning the teacher's own flaws. Practitioners might care because it can lead to more accurate models that better represent the capability gap between teachers and students.
Firehose
Filtered to tagged “on-policy distillation” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives
Browse by tag
This paper proposes a new method for self-retiring on-policy distillation in agentic reinforcement learning, which allows agents to learn more effectively by switching between teacher-student training and reinforcement learning alone. Practitioners might care about this approach because it can improve the performance of agents in complex tasks.
This paper investigates why some AI models can generate excessively long responses and proposes a solution to mitigate this issue by aligning the models' termination tokens. Practitioners might care about this problem because it can lead to wasted generation budgets and inefficient AI applications.