This paper develops a method to improve reasoning models by reducing the discrepancy between a stronger teacher and an on-policy student, which helps to prevent the student from learning the teacher's own flaws. Practitioners might care because it can lead to more accurate models that better represent the capability gap between teachers and students.
Firehose
Filtered to Papers, tagged “calibration” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives