13 upvotes · 15 JUL 2026 · Dwip Dalal, Shivansh Patel, Chahit Jain et al.
This paper proposes a new method for fine-tuning vision-language models on robot demonstrations to improve their performance on real-world tasks, by preventing the overwrite of pre-trained representations and aligning language and action predictions. Practitioners may care about this work because it aims to improve the generalizability and robustness of vision-language-action policies in real-world applications.
8 upvotes · 21 JUL 2026 · Shu Wei, Jingjing Wu, Lingshu Zhang et al.
This paper introduces HPD-Parsing, a new approach to document parsing that uses hierarchical parallel decoding to improve efficiency and throughput. Practitioners in natural language processing and computer vision might care because it could lead to faster and more accurate document parsing models.
8 upvotes · 22 JUL 2026 · Arjun Majumdar, Avinash Sooriyarachchi, Benjamin Tibi et al.
This paper introduces Robostral Navigate, a vision-language model that enables robots to navigate using only a single monocular RGB camera, making it more scalable and cost-effective for deployment across various robotic platforms. Practitioners might care about this because it can simplify navigation tasks for robots in real-world environments.
3 upvotes · 13 JUL 2026 · Mingyuan Wu, Jingcheng Yang, Shengyi Qian et al.
This paper introduces a new reinforcement learning framework called SVR-R1 that helps models improve their reasoning abilities by giving them the chance to correct their own mistakes. Practitioners might care about this because it could lead to better performance in tasks that require complex reasoning.
3 upvotes · 22 JUL 2026 · Yang Xu, Gurpreet Singh Mukker, Raymond Wang et al.
This paper develops a new framework for robotic grasping in complex scenes that uses language to specify task requirements, allowing robots to grasp objects with more precision and flexibility. Practitioners working on robotic grasping and manipulation might care about this approach because it could improve the performance and efficiency of robots in real-world applications.
2 upvotes · 7 JUL 2026 · Grace Man Chen, Litao Guo, Yifan Wu et al.
This paper introduces UI2App, a benchmark to evaluate the ability of large language models to infer interaction behavior from screenshots of web applications, which is crucial for real-world development workflows. Practitioners might care about this research because it highlights the limitations of current visual-driven approaches and the need for better interaction inference capabilities in web application generation.
2 upvotes · 22 JUL 2026 · Karan Goyal, Afreen Hossain, Debojyoti Das et al.
This paper introduces a new dataset and tool to study contextual entrainment in vision-language models, which is the tendency for models to respond to irrelevant or false context in their inputs. Practitioners in AI and ML might care about this because it can affect the accuracy and reliability of vision-language models in real-world applications.
0 upvotes · 18 JUL 2026 · Haoru Tan, Wang Wang, Sitong Wu et al.
This paper proposes a new method for dataset distillation that focuses on aligning the final outcome of training, rather than imitating the training process. Practitioners might care about this method because it can improve the accuracy of model training on limited data.