Neural Volumetric Memory for Visual Locomotion Control
Ruihan Yang, Ge Yang, Xiaolong Wang
Abstract
Legged robots have the potential to expand the reach of autonomy beyond paved roads. In this work, we consider the difficult problem of locomotion on challenging terrains using a single forward-facing depth camera. Due to the partial observability of the problem, the robot has to rely on past observations to infer the terrain currently beneath it. To solve this problem, we follow the paradigm in computer vision that explicitly models the 3D geometry of the scene and propose Neural Volumetric Memory (NVM), a geometric memory architecture that explicitly accounts for the SE(3) equivariance of the 3D world. NVM aggregates feature volumes from multiple camera views by first bringing them back to the ego-centric frame of the robot. We test the learned visual-locomotion policy on a physical robot and show that our approach, which explicitly introduces geometric priors during training, offers superior performance than more naïve methods. We also include ablation studies and show that the representations stored in the neural volumetric memory capture sufficient geometric information to reconstruct the scene. Our project page with videos is https://rchalyang.github.io/NVM
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7b501e52-d722-4958-a4b4-9016b14cd4c7Cited by top-tier papers6
- Hybrid Internal Model: Learning Agile Legged Locomotion with Simulated Robot ResponseJunfeng Long, Zirui Wang, Quanyi Li, Liu Cao et al.ICLR 2024 · 66 citations
- Deep SE(3)-Equivariant Geometric Reasoning for Precise Placement TasksBen Eisner, Yi Yang, Todor Davchev, Mel Vecerík et al.ICLR 2024 · 24 citations
- VoxDet: Voxel Learning for Novel Instance DetectionBowen Li, Jiashun Wang, Yaoyu Hu, Chen Wang et al.NeurIPS 2023 · 14 citations
- Active Vision Reinforcement Learning under Limited Visual ObservabilityJinghuan Shang, Michael S. RyooNeurIPS 2023 · 1 citation
- Let Humanoids Hike! Integrative Skill Development on Complex TrailsKwan-Yee Lin, Stella X. YuCVPR 2025
Builds on8
- AMP: adversarial motion priors for stylized physics-based character controlXue Bin Peng, Ze Ma, Pieter Abbeel, Sergey Levine et al.SIGGRAPH 2021 · 392 citations
- Learning Vision-Guided Quadrupedal Locomotion End-to-End with Cross-Modal TransformersRuihan Yang, Minghao Zhang, Nicklas Hansen, Huazhe Xu et al.ICLR 2022 · 146 citations
- HoloGAN: Unsupervised Learning of 3D Representations From Natural ImagesThu Nguyen-Phuoc, Chuan Li, Lucas Theis, Christian Richardt et al.ICCV 2019 · 98 citations
- PixelSynth: Generating a 3D-Consistent Experience from a Single ImageChris Rockwell, David F. Fouhey, Justin JohnsonICCV 2021 · 98 citations
- Worldsheet: Wrapping the World in a 3D Sheet for View Synthesis from a Single ImageRonghang Hu, Nikhila Ravi, Alexander C. Berg, Deepak PathakICCV 2021 · 97 citations
Related papers
- Building 3D Representations and Generating Motions From a Single Image via Video-GenerationWeiming Zhi, Ziyong Ma, Tianyi Zhang, Matthew Johnson-RobersonNeurIPS 2025 · 1 citation
- SE(3) Equivariant Convolution and Transformer in Ray SpaceYinshuang Xu, Jiahui Lei, Kostas DaniilidisNeurIPS 2023 · 6 citations
- AMB3R: Accurate Feed-forward Metric-scale 3D Reconstruction with BackendHengyi Wang, Lourdes AgapitoCVPR 2026 · 17 citations
- Volumetric Environment Representation for Vision-Language NavigationRui Liu, Wenguan Wang, Yi YangCVPR 2024 · 25 citations
- Dream-to-Recon: Monocular 3D Reconstruction with Diffusion-Depth Distillation from Single ImagesPhilipp Wulff, Felix Wimbauer, Dominik Muhle, Daniel CremersICCV 2025 · 1 citation
