Papers

Filtered to vision-language models · clear filter

Browse by term

continual learning 33reinforcement learning 21large language models 9vision-language models 8language models 6benchmarking 5autoregressive models 4diffusion transformers 4generative models 4multimodal learning 4robotics 4video generation 4benchmarks 3computer vision 3policy optimization 3self-distillation 3vision-language-action models 3world modeling 3agent-based systems 2autonomous agents 2calibration 2coding agents 2diffusion models 2foundation models 2image synthesis 2in-context learning 2knowledge graphs 2LLMs 2multimodal evaluation 2multimodal large language models 2

Matching papers

Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment

13 upvotes · 15 JUL 2026 · Dwip Dalal, Shivansh Patel, Chahit Jain et al.

This paper proposes a new method for fine-tuning vision-language models on robot demonstrations to improve their performance on real-world tasks, by preventing the overwrite of pre-trained representations and aligning language and action predictions. Practitioners may care about this work because it aims to improve the generalizability and robustness of vision-language-action policies in real-world applications.

HPD-Parsing: Hierarchical Parallel Document Parsing

8 upvotes · 21 JUL 2026 · Shu Wei, Jingjing Wu, Lingshu Zhang et al.

This paper introduces HPD-Parsing, a new approach to document parsing that uses hierarchical parallel decoding to improve efficiency and throughput. Practitioners in natural language processing and computer vision might care because it could lead to faster and more accurate document parsing models.

Robostral Navigate

8 upvotes · 22 JUL 2026 · Arjun Majumdar, Avinash Sooriyarachchi, Benjamin Tibi et al.

This paper introduces Robostral Navigate, a vision-language model that enables robots to navigate using only a single monocular RGB camera, making it more scalable and cost-effective for deployment across various robotic platforms. Practitioners might care about this because it can simplify navigation tasks for robots in real-world environments.

SVR-R1: Bootstrapping Multi-modal Reasoning with Self-verification in Reinforcement Learning

3 upvotes · 13 JUL 2026 · Mingyuan Wu, Jingcheng Yang, Shengyi Qian et al.

This paper introduces a new reinforcement learning framework called SVR-R1 that helps models improve their reasoning abilities by giving them the chance to correct their own mistakes. Practitioners might care about this because it could lead to better performance in tasks that require complex reasoning.

SeededGrasp: Language-Guided Grasping in Complex Scenes with Multiple Embodiments

3 upvotes · 22 JUL 2026 · Yang Xu, Gurpreet Singh Mukker, Raymond Wang et al.

This paper develops a new framework for robotic grasping in complex scenes that uses language to specify task requirements, allowing robots to grasp objects with more precision and flexibility. Practitioners working on robotic grasping and manipulation might care about this approach because it could improve the performance and efficiency of robots in real-world applications.

UI2App: Benchmarking Visual Interaction Inference in Executable Web Application Generation

2 upvotes · 7 JUL 2026 · Grace Man Chen, Litao Guo, Yifan Wu et al.

This paper introduces UI2App, a benchmark to evaluate the ability of large language models to infer interaction behavior from screenshots of web applications, which is crucial for real-world development workflows. Practitioners might care about this research because it highlights the limitations of current visual-driven approaches and the need for better interaction inference capabilities in web application generation.

ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models

2 upvotes · 22 JUL 2026 · Karan Goyal, Afreen Hossain, Debojyoti Das et al.

This paper introduces a new dataset and tool to study contextual entrainment in vision-language models, which is the tendency for models to respond to irrelevant or false context in their inputs. Practitioners in AI and ML might care about this because it can affect the accuracy and reliability of vision-language models in real-world applications.