Papers

Filtered to benchmarking · clear filter

Browse by term

continual learning 33reinforcement learning 21large language models 9vision-language models 8language models 6benchmarking 5autoregressive models 4diffusion transformers 4generative models 4multimodal learning 4robotics 4video generation 4benchmarks 3computer vision 3policy optimization 3self-distillation 3vision-language-action models 3world modeling 3agent-based systems 2autonomous agents 2calibration 2coding agents 2diffusion models 2foundation models 2image synthesis 2in-context learning 2knowledge graphs 2LLMs 2multimodal evaluation 2multimodal large language models 2

Matching papers

AREX: Towards a Recursively Self-Improving Agent for Deep Research

115 upvotes · 23 JUL 2026 · Shuqi Lu, Chaofan Li, Kun Luo et al.

This paper introduces AREX, a self-improving agent that recursively refines its answers by verifying intermediate results and using the partially verified state to guide subsequent refinement, aiming to improve the efficiency of deep research. Practitioners might care about AREX's approach to reducing the cost of discovery and verification in complex research tasks.

K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs

39 upvotes · 23 JUL 2026 · Hao Liang, Qihan Lin, Zhaoyang Han et al.

This paper introduces a new framework for training educational language models, specifically designed to evaluate their ability to understand curriculum knowledge and its visual presentation. Practitioners might care about this work because it aims to improve language models' performance in educational settings.

Reading and Steering Representations of Materials-Science Mechanisms in an Open-Weight Language Model

5 upvotes · 22 JUL 2026 · Markus J. Buehler

This paper investigates whether large language models, like Google's Gemma-4-E4B-it, represent scientific concepts and governing physics, and whether this representation affects their answers. Practitioners caring about the accuracy and reliability of language models in scientific domains might find this research valuable.

G-MAD: A Game-Based Data Generation Framework for Multi-View RGB-T Aerial Object Detection

4 upvotes · 22 JUL 2026 · Yechan Kim, JongHyun Park, Dongho Yoon et al.

This paper introduces a new framework for generating synchronized multi-view RGB-T data for aerial object detection, which can help practitioners improve their object detection models by training on more realistic data and by reducing the need for expensive real-world data collection.

FinanceComplexQA: Benchmarking Agentic Reasoning on Industrial-grade Financial Documents

4 upvotes · 21 JUL 2026 · Xianfu Cheng, Shiwei Zhang, Jiyu Zhao et al.

This paper creates a benchmark for testing the ability of AI agents to understand and analyze complex financial documents, and uses it to evaluate the performance of different agents in this task. Practitioners in finance and AI research can care about this work because it aims to improve the accuracy and reliability of financial document analysis.