This episode features Beth Barnes and David Rein from METR discussing their 'Time Horizon' graph, a unified metric for measuring AI progress based on human task completion time. They explain how this benchmark addresses the limitations of t…
Firehose
Filtered to Podcasts, tagged “Reward Hacking” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News
Browse by tag
The RL Fine-Tuning Playbook: CoreWeave's Kyle Corbitt on GRPO, Rubrics, Environments, Reward Hacking
Kyle Corbitt, founder of OpenPipe and leader of CoreWeave's serverless training team, provides a master class on reinforcement learning (RL) and custom fine-tuning for AI models. He explains how RL differs from supervised fine-tuning (SFT) …