Papers

Filtered to scene understanding · clear filter

Browse by term

continual learning 33reinforcement learning 21large language models 9vision-language models 8language models 6benchmarking 5autoregressive models 4diffusion transformers 4generative models 4multimodal learning 4robotics 4video generation 4benchmarks 3computer vision 3policy optimization 3self-distillation 3vision-language-action models 3world modeling 3agent-based systems 2autonomous agents 2calibration 2coding agents 2diffusion models 2foundation models 2image synthesis 2in-context learning 2knowledge graphs 2LLMs 2multimodal evaluation 2multimodal large language models 2

Matching papers

Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers

9 upvotes · 23 JUL 2026 · Sicheng Mo, Yuheng Li, Ziyang Leng et al.

This paper introduces a new method for generating videos in multi-agent environments, where each agent has its own view of the world. It's useful for applications like video games or simulations where multiple agents need to interact with each other and the environment.

SeededGrasp: Language-Guided Grasping in Complex Scenes with Multiple Embodiments

3 upvotes · 22 JUL 2026 · Yang Xu, Gurpreet Singh Mukker, Raymond Wang et al.

This paper develops a new framework for robotic grasping in complex scenes that uses language to specify task requirements, allowing robots to grasp objects with more precision and flexibility. Practitioners working on robotic grasping and manipulation might care about this approach because it could improve the performance and efficiency of robots in real-world applications.