This paper introduces AREX, a self-improving agent that recursively refines its answers by verifying intermediate results and using the partially verified state to guide subsequent refinement, aiming to improve the efficiency of deep research. Practitioners might care about AREX's approach to reducing the cost of discovery and verification in complex research tasks.
Firehose
Filtered to tagged “benchmarking” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News
Browse by tag
This paper introduces a new framework for training educational language models, specifically designed to evaluate their ability to understand curriculum knowledge and its visual presentation. Practitioners might care about this work because it aims to improve language models' performance in educational settings.
This paper investigates whether large language models, like Google's Gemma-4-E4B-it, represent scientific concepts and governing physics, and whether this representation affects their answers. Practitioners caring about the accuracy and reliability of language models in scientific domains might find this research valuable.
This paper introduces a new framework for generating synchronized multi-view RGB-T data for aerial object detection, which can help practitioners improve their object detection models by training on more realistic data and by reducing the need for expensive real-world data collection.
This paper creates a benchmark for testing the ability of AI agents to understand and analyze complex financial documents, and uses it to evaluate the performance of different agents in this task. Practitioners in finance and AI research can care about this work because it aims to improve the accuracy and reliability of financial document analysis.