59 upvotes · 16 JUL 2026 · Xiaomi Robotics Team, Jun Guo, Piaopiao Jin et al.
This paper introduces a vision-language-action model that can perform mobile manipulation tasks in unseen environments with minimal training data, and how it can be scaled up to achieve better performance. Practitioners might care about this model for building robots that can adapt to new tasks with minimal fine-tuning.
4 upvotes · 17 JUL 2026 · Haoran Sun, Wentao Zhang, Junyang Hua et al.
This paper develops a service-oriented framework, JoyNexus, to efficiently train and deploy Vision-Language-Action models across multiple tenants, improving resource utilization and reducing costs. Practitioners may care about JoyNexus for its potential to streamline the training process and make VLA models more accessible.
4 upvotes · 13 JUL 2026 · Byungkun Lee, Dongyoon Hwang, Dongjin Kim et al.
This paper proposes a way to improve vision-language-action models so they can better understand the world from a robot's perspective, which is important for robots to make accurate decisions. By using robot-centric pointmaps, these models can generalize better across different camera setups and viewpoints.