This paper explores how large language models can become overly influenced by their own judgments, leading to a loss of diversity in scientific evaluations. Practitioners should care about this issue because it can impact the quality of reviews and recommendations in AI-assisted scientific evaluation.
Firehose
Filtered to Papers, tagged “large language models” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives
Browse by tag
This paper investigates how the way large language models generate multiple candidate responses affects their performance and energy consumption. Practitioners might care because optimizing test-time scaling can lead to significant improvements in model accuracy and efficiency.
This paper proposes a new method for aligning large language models with human preferences, called Comparison-based Preference Optimization (ComPO), which is more efficient than existing methods and can mitigate a problem called likelihood displacement. Practitioners might care about this paper because it offers a new approach to aligning LLMs with human preferences, which is essential for developing more reliable and trustworthy AI models.