Hierarchical Adaptive Value Estimation for Multi-modal Visual Reinforcement Learning
Yangru Huang, Peixi Peng, Yifan Zhao, Haoran Xu, Mengyue Geng, Yonghong Tian
Abstract
Integrating RGB frames with alternative modality inputs is gaining increasing traction in many vision-based reinforcement learning (RL) applications. Existing multi-modal vision-based RL methods usually follow a Global Value Estimation (GVE) pipeline, which uses a fused modality feature to obtain a unified global environmental description. However, such a feature-level fusion paradigm with a single critic may fall short in policy learning as it tends to overlook the distinct values of each modality. To remedy this, this paper proposes a Local modalitycustomized Value Estimation (LVE) paradigm, which dynamically estimates the contribution and adjusts the importance weight of each modality from a valuelevel perspective. Furthermore, a task-contextual re-fusion process is developed to achieve a task-level re-balance of estimations from both feature and value levels. To this end, a Hierarchical Adaptive Value Estimation (HAVE) framework is formed, which adaptively coordinates the contributions of individual modalities as well as their collective efficacy. Agents trained by HAVE are able to exploit the unique characteristics of various modalities while capturing their intricate interactions, achieving substantially improved performance. We specifically highlight the potency of our approach within the challenging landscape of autonomous driving, utilizing the CARLA benchmark with neuromorphic event and depth data to demonstrate HAVE's capability and the effectiveness of its distinct components. The code of our paper can be found at https://github.com/Yara-HYR/HAVE . RGB Critic 𝑞 𝑀 1 (d) Local modality-customized Value Estimation (LVE) 𝑞 𝑡 𝑙 𝑞 𝑀 2 𝑤 𝑀 1 𝑤 𝑀 2 𝑟 𝑡 𝑞 𝑡 𝑔 Critic (c) Global Value Estimation (GVE)
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 16bb3f54-4837-4697-9cea-c9af26c9a440Cited by top-tier papers2
- Seek Commonality but Preserve Differences: Dissected Dynamics Modeling for Multi-modal Visual RLYangru Huang, Peixi Peng, Yifan Zhao, Guangyao Chen et al.NeurIPS 2024 · 3 citations
- VLMs-Guided Representation Distillation for Efficient Vision-Based Reinforcement LearningHaoran Xu, Peixi Peng, Guang Tan, Yiqian Chang et al.CVPR 2025
Builds on16
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 1,261 citations
- Target-aware Dual Adversarial Learning and a Multi-scenario Multi-Modality Benchmark to Fuse Infrared and Visible for Object DetectionJinyuan Liu, Xin Fan, Zhanbo Huang, Guanyao Wu et al.CVPR 2022 · 929 citations
- Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from PixelsDenis Yarats, Ilya Kostrikov, Rob FergusICLR 2021 · 911 citations
- Reinforcement Learning with Augmented DataMichael Laskin, Kimin Lee, Adam Stooke, Lerrel Pinto et al.NeurIPS 2020 · 833 citations
Related papers
- DMR: Decomposed Multi-Modality Representations for Frames and Events Fusion in Visual Reinforcement LearningHaoran Xu, Peixi Peng, Guang Tan, Yuan Li et al.CVPR 2024 · 5 citations
- EE-RL: Vision Language Guided Reinforcement Learning with Explorer and Expert model for End-to-End Autonomous DrivingXiaolong Li, Lan Yang, Ruyang Li, Shan Fang et al.CVPR 2026
- VMLoc: Variational Fusion For Learning-Based Multimodal Camera LocalizationKaichen Zhou, Changhao Chen, Bing Wang, Muhamad Risqi Utama Saputra et al.AAAI 2021 · 25 citations
- Learning Mixture of Domain-Specific Experts via Disentangled Factors for Autonomous DrivingInhan Kim, Joonyeong Lee, Daijin KimAAAI 2022 · 9 citations
- Parameterized Decision-Making with Multi-Modality Perception for Autonomous DrivingYuyang Xia, Shuncheng Liu, Quanlin Yu, Liwei Deng et al.ICDE 2024 · 26 citations
