Papers

Filtered to vision-language-action models · clear filter

Browse by term

continual learning 86reinforcement learning 48benchmarking 13large language models 12benchmarks 11vision-language models 10language models 9robotics 7world models 7natural language processing 6recursive self-improvement 6generative models 5on-policy distillation 5video generation 5attention mechanisms 4coding agents 4LLMs 4multi-agent systems 4multimodal learning 4multimodal models 4self-distillation 4self-supervised learning 4transformers 4vision-language-action models 4world modeling 4agent-based systems 3agentic models 3agentic search 3agents 3autonomous systems 3

Matching papers

TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM

118 upvotes · 29 JUL 2026 · Hengyi Xie, Chenfei Yao, Xianjin Wu et al.

This paper introduces TurboVLA, a new vision-language-action model that reduces computation and memory overhead by directly exchanging information between visual observations and language instructions, allowing for faster and more efficient robotic manipulation. Practitioners might care about this approach for building more efficient and effective VLA models.

GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System Architecture

97 upvotes · 16 AUG 2026 · GigaBrain Team, Angen Ye, Axiang Sun et al.

This paper presents GigaBrain-0.7, a new embodied foundation model that achieves strong generalization across diverse robot embodiments and tasks, by improving the architecture and scaling it to large amounts of data. Practitioners might care about this research if they're working on developing generalist robots that can adapt to new tasks and environments.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

59 upvotes · 16 JUL 2026 · Xiaomi Robotics Team, Jun Guo, Piaopiao Jin et al.

This paper introduces a vision-language-action model that can perform mobile manipulation tasks in unseen environments with minimal training data, and how it can be scaled up to achieve better performance. Practitioners might care about this model for building robots that can adapt to new tasks with minimal fine-tuning.

Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models

55 upvotes · 27 AUG 2026 · Senqiao Yang, Chengyao Wang, Yuxin Chen et al.

This paper proposes a new approach to training Vision-Language-Action models by using a pre-trained backbone that captures generalizable visual-action knowledge from a large, diverse dataset of robot trajectories. This allows the model to perform well on new, unseen tasks without requiring a large amount of task-specific data.