Papers

Filtered to video generation · clear filter

Browse by term

continual learning 33reinforcement learning 21large language models 9vision-language models 8language models 6benchmarking 5autoregressive models 4diffusion transformers 4generative models 4multimodal learning 4robotics 4video generation 4benchmarks 3computer vision 3policy optimization 3self-distillation 3vision-language-action models 3world modeling 3agent-based systems 2autonomous agents 2calibration 2coding agents 2diffusion models 2foundation models 2image synthesis 2in-context learning 2knowledge graphs 2LLMs 2multimodal evaluation 2multimodal large language models 2

Matching papers

HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement

52 upvotes · 20 JUL 2026 · Yiyang Cai, Nan Chen, Rongchang Xie et al.

This paper develops a video personalization method that focuses on human-object interactions, aiming to improve the accuracy of video generation by better understanding human-object relationships and incorporating intra-subject references. Practitioners may care about this research as it could lead to more realistic and engaging video content.

Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence

40 upvotes · 17 JUL 2026 · Runmao Yao, Kairui Hu, Yukang Cao et al.

This paper introduces a benchmark to evaluate video generation models' ability to reason about physical laws, which is crucial for creating reliable world simulators. Practitioners caring about developing more realistic and physically intelligent AI models will find this research valuable.

Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers

9 upvotes · 23 JUL 2026 · Sicheng Mo, Yuheng Li, Ziyang Leng et al.

This paper introduces a new method for generating videos in multi-agent environments, where each agent has its own view of the world. It's useful for applications like video games or simulations where multiple agents need to interact with each other and the environment.

FVAttn: Adaptive Sparse Attention with Runtime Load Balancing for Video Generation

8 upvotes · 17 JUL 2026 · Hao Liu, Chenghuan Huang, Ye Huang et al.

This paper develops a more efficient way to generate high-quality videos by balancing the workload across multiple GPUs during training, which can improve the performance of video generation models like those used in FVAttn. Practitioners in video generation and deep learning might care about this research because it can lead to faster and more efficient video generation models.