Neural Volumetric Memory for Visual Locomotion Control
Ruihan Yang, Ge Yang, Xiaolong Wang
摘要
Legged robots have the potential to expand the reach of autonomy beyond paved roads. In this work, we consider the difficult problem of locomotion on challenging terrains using a single forward-facing depth camera. Due to the partial observability of the problem, the robot has to rely on past observations to infer the terrain currently beneath it. To solve this problem, we follow the paradigm in computer vision that explicitly models the 3D geometry of the scene and propose Neural Volumetric Memory (NVM), a geometric memory architecture that explicitly accounts for the SE(3) equivariance of the 3D world. NVM aggregates feature volumes from multiple camera views by first bringing them back to the ego-centric frame of the robot. We test the learned visual-locomotion policy on a physical robot and show that our approach, which explicitly introduces geometric priors during training, offers superior performance than more naïve methods. We also include ablation studies and show that the representations stored in the neural volumetric memory capture sufficient geometric information to reconstruct the scene. Our project page with videos is https://rchalyang.github.io/NVM
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Hybrid Internal Model: Learning Agile Legged Locomotion with Simulated Robot ResponseJunfeng Long, Zirui Wang, Quanyi Li, Liu Cao 等ICLR 2024 · 被引用 66 次
- Deep SE(3)-Equivariant Geometric Reasoning for Precise Placement TasksBen Eisner, Yi Yang, Todor Davchev, Mel Vecerík 等ICLR 2024 · 被引用 24 次
- VoxDet: Voxel Learning for Novel Instance DetectionBowen Li, Jiashun Wang, Yaoyu Hu, Chen Wang 等NeurIPS 2023 · 被引用 14 次
- Active Vision Reinforcement Learning under Limited Visual ObservabilityJinghuan Shang, Michael S. RyooNeurIPS 2023 · 被引用 1 次
- Let Humanoids Hike! Integrative Skill Development on Complex TrailsKwan-Yee Lin, Stella X. YuCVPR 2025
它引用的顶会 Paper8
- AMP: adversarial motion priors for stylized physics-based character controlXue Bin Peng, Ze Ma, Pieter Abbeel, Sergey Levine 等SIGGRAPH 2021 · 被引用 392 次
- Learning Vision-Guided Quadrupedal Locomotion End-to-End with Cross-Modal TransformersRuihan Yang, Minghao Zhang, Nicklas Hansen, Huazhe Xu 等ICLR 2022 · 被引用 146 次
- HoloGAN: Unsupervised Learning of 3D Representations From Natural ImagesThu Nguyen-Phuoc, Chuan Li, Lucas Theis, Christian Richardt 等ICCV 2019 · 被引用 98 次
- PixelSynth: Generating a 3D-Consistent Experience from a Single ImageChris Rockwell, David F. Fouhey, Justin JohnsonICCV 2021 · 被引用 98 次
- Worldsheet: Wrapping the World in a 3D Sheet for View Synthesis from a Single ImageRonghang Hu, Nikhila Ravi, Alexander C. Berg, Deepak PathakICCV 2021 · 被引用 97 次
相关 Paper
- Building 3D Representations and Generating Motions From a Single Image via Video-GenerationWeiming Zhi, Ziyong Ma, Tianyi Zhang, Matthew Johnson-RobersonNeurIPS 2025 · 被引用 1 次
- SE(3) Equivariant Convolution and Transformer in Ray SpaceYinshuang Xu, Jiahui Lei, Kostas DaniilidisNeurIPS 2023 · 被引用 6 次
- AMB3R: Accurate Feed-forward Metric-scale 3D Reconstruction with BackendHengyi Wang, Lourdes AgapitoCVPR 2026 · 被引用 17 次
- Volumetric Environment Representation for Vision-Language NavigationRui Liu, Wenguan Wang, Yi YangCVPR 2024 · 被引用 25 次
- Dream-to-Recon: Monocular 3D Reconstruction with Diffusion-Depth Distillation from Single ImagesPhilipp Wulff, Felix Wimbauer, Dominik Muhle, Daniel CremersICCV 2025 · 被引用 1 次
