NaVILA: Legged Robot Vision-Language-Action Model for Navigation

An-Chieh Cheng; Yandong Ji; Zhaojing Yang; Zaitian Gongye; Xueyan Zou; Jan Kautz; Erdem Biyik; Hongxu Yin; Sifei Liu; Xiaolong Wang

2025 RSS RSS 2025

NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Abstract

This paper proposes to solve the problem of Vision and-Language Navigation with legged robots, which not only provides a flexible way for humans to command but also allows the robot to navigate through more challenging and cluttered scenes. However, it is non-trivial to translate human language instructions all the way to low-level leg joint actions. We propose NaVILA, a 2-level framework that unifies a Vision-Language Action model (VLA) with locomotion skills. Instead of directly predicting low-level actions from VLA, NaVILA first generates mid-level actions with spatial information in the form of language, (e.g., “moving forward 75cm”), which serves as an input for a visual locomotion RL policy for execution. NaVILA substantially improves previous approaches on existing benchmarks. The same advantages are demonstrated in our newly developed benchmarks with IsaacLab, featuring more realistic scenes, low-level controls, and real-world robot experiments.

🌉 Interdisciplinary Bridge — Artificial Intelligence and Reinforcement Learning

🐝 Cross-Pollinator — Artificial Intelligence, Computer Science, Computer Vision, Data Science & Analytics, Deep Learning, Interdisciplinary, Knowledge & Reasoning, Machine Learning, Natural Language Processing, Reinforcement Learning, Robotics

Authors

An-Chieh Cheng , Yandong Ji , Zhaojing Yang , Zaitian Gongye , Xueyan Zou , Jan Kautz , Erdem Biyik , Hongxu Yin , Sifei Liu , Xiaolong Wang

Topics

Artificial Intelligence > Core AI > Autonomous Vehicles Reinforcement Learning > Applications > Robotics

Keywords

vision-language navigation locomotion skill legged robot navigation

Download PDF

Related papers

Enhancing Autonomous Driving Systems with On-Board Deployed Large Language Models 2025

Debiasing 6-DOF IMU via Hierarchical Learning of Continuous Bias Dynamics 2025

SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Models 2025

RoboVerse: A Unified Platform, Benchmark and Dataset for Scalable and Generalizable Robot Learning 2025

Learning Humanoid Standing-up Control across Diverse Postures 2025