Papers

Filtered to visual-linguistic integration · clear filter

Browse by term

continual learning 33reinforcement learning 21large language models 9vision-language models 8language models 6benchmarking 5autoregressive models 4diffusion transformers 4generative models 4multimodal learning 4robotics 4video generation 4benchmarks 3computer vision 3policy optimization 3self-distillation 3vision-language-action models 3world modeling 3agent-based systems 2autonomous agents 2calibration 2coding agents 2diffusion models 2foundation models 2image synthesis 2in-context learning 2knowledge graphs 2LLMs 2multimodal evaluation 2multimodal large language models 2

Matching papers

See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models

4 upvotes · 13 JUL 2026 · Byungkun Lee, Dongyoon Hwang, Dongjin Kim et al.

This paper proposes a way to improve vision-language-action models so they can better understand the world from a robot's perspective, which is important for robots to make accurate decisions. By using robot-centric pointmaps, these models can generalize better across different camera setups and viewpoints.