Papers

Filtered to world models · clear filter

Browse by term

continual learning 86reinforcement learning 48benchmarking 13large language models 12benchmarks 11vision-language models 10language models 9robotics 7world models 7natural language processing 6recursive self-improvement 6generative models 5on-policy distillation 5video generation 5attention mechanisms 4coding agents 4LLMs 4multi-agent systems 4multimodal learning 4multimodal models 4self-distillation 4self-supervised learning 4transformers 4vision-language-action models 4world modeling 4agent-based systems 3agentic models 3agentic search 3agents 3autonomous systems 3

Matching papers

Scaling Automatic Research Agents via World Models

452 upvotes · 29 AUG 2026 · Xiyuan Yang, Sheikh Sarwar, Jingru Cheng et al.

This paper proposes a method to scale automatic research agents by replacing environment execution with a world model, which can reduce training costs and improve performance. Practitioners might care about this approach because it can accelerate training times and lead to better results for complex AI tasks.

Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models

185 upvotes · 26 AUG 2026 · Pengfei Zhou, Hexin Wang, Zhengfeiyang Zhang et al.

This paper proposes a way to improve the efficiency of training world models by using game development as a source of reward signals and trajectory data, allowing for more effective post-training of large language models using reinforcement learning. Practitioners might care about this approach because it could lead to more scalable and effective world models for applications like dialogue systems and visual question answering.

Can MiniMax-H3 Reason About the Physical World? An Evaluation of Omni-Modal Generative Model

115 upvotes · 16 SEP 2026 · Haoyu Zhao, Zihao Zhao, Tianyu Deng et al.

This paper evaluates whether a multimodal generative model can reason about the physical world by testing its ability to integrate information from different modalities, such as text, images, video, and audio. Practitioners might care about this research because it can help develop more advanced models that can better understand and generate complex scenarios.

HarnessEval-W: Agentifying the Evaluation of Visual Worlds

108 upvotes · 17 AUG 2026 · Weiliang Chen, Haowen Sun, Jun Gao et al.

This paper develops a new method for evaluating world models, called HarnessEval-W, which provides more detailed and justifiable results than existing benchmarks. Practitioners might care about HarnessEval-W because it can help them build more trustworthy world models that better align with human preferences.