Papers

Filtered to multimodal learning · clear filter

Browse by term

continual learning 33reinforcement learning 21large language models 9vision-language models 8language models 6benchmarking 5autoregressive models 4diffusion transformers 4generative models 4multimodal learning 4robotics 4video generation 4benchmarks 3computer vision 3policy optimization 3self-distillation 3vision-language-action models 3world modeling 3agent-based systems 2autonomous agents 2calibration 2coding agents 2diffusion models 2foundation models 2image synthesis 2in-context learning 2knowledge graphs 2LLMs 2multimodal evaluation 2multimodal large language models 2

Matching papers

K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs

39 upvotes · 23 JUL 2026 · Hao Liang, Qihan Lin, Zhaoyang Han et al.

This paper introduces a new framework for training educational language models, specifically designed to evaluate their ability to understand curriculum knowledge and its visual presentation. Practitioners might care about this work because it aims to improve language models' performance in educational settings.

FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicry

31 upvotes · 20 JUL 2026 · Dingyun Zhang, Lixue Gong, Wei Liu

This paper creates a new AI model that can edit and generate videos without needing masks, and can also learn to mimic image editing capabilities. Practitioners might care about this because it could lead to more diverse and realistic video editing data, and enable AI models to understand and generate human-like video editing instructions.

ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models

2 upvotes · 22 JUL 2026 · Karan Goyal, Afreen Hossain, Debojyoti Das et al.

This paper introduces a new dataset and tool to study contextual entrainment in vision-language models, which is the tendency for models to respond to irrelevant or false context in their inputs. Practitioners in AI and ML might care about this because it can affect the accuracy and reliability of vision-language models in real-world applications.

ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning

1 upvotes · 20 JUL 2026 · Ting Huang, Zhenyu Zhang, Wenyuan Huang et al.

This paper proposes a new framework for video spatial reasoning that focuses on learning geometric consistency to improve the accuracy and stability of models in tasks like navigation and question answering. Practitioners working on multimodal models and spatial reasoning tasks may benefit from this approach.