Agon: Competitive Cross-Model RL with Implicit Rival Grading of Reasoning
This paper introduces Agon, a method for training reinforcement learning models to improve their reasoning abilities by competing with each other, rather than just optimizing for the final answer. Practitioners may care about using this approach to improve the quality of reasoning in AI models.