DMR: Decomposed Multi-Modality Representations for Frames and Events Fusion in Visual Reinforcement Learning
Haoran Xu, Peixi Peng, Guang Tan, Yuan Li, Xinhai Xu, Yonghong Tian
Abstract
We explore visual reinforcement learning (RL) using two complementary visual modalities: frame-based RGB cam-era and event-based Dynamic Vision Sensor (DVS). Ex-isting multi-modality visual RL methods often encounter challenges in effectively extracting task-relevant information from multiple modalities while suppressing the in-creased noise, only using indirect reward signals instead of pixel-level supervision. To tackle this, we propose a Decomposed Multi-Modality Representation (DMR) framework for visual RL. It explicitly decomposes the inputs into three distinct components: combined task-relevant features (co-features), RGB-specific noise, and DVS-specific noise. The co-features represent the full information from both modalities that is relevant to the RL task; the two noise components, each constrained by a data reconstruction loss to avoid information leak, are contrasted with the co-features to maximize their difference. Extensive experiments demonstrate that, by explicitly separating the different types of information, our approach achieves substan-tially improved policy performance compared to state-of-the-art approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3655b80b-3d40-4f66-a5e2-cee7ccadc422Cited by top-tier papers7
- RGB-Event based Pedestrian Attribute Recognition: A Benchmark Dataset and An Asymmetric RWKV Fusion FrameworkXiao Wang, Haiyang Wang, Shiao Wang, Qiang Chen et al.CVPR 2026 · 8 citations
- Multimodal Decomposed Distillation with Instance Alignment and Uncertainty Compensation for Thermal Object DetectionYanfeng Liu, Lefei ZhangACM MM 2025 · 2 citations
- SEBA: Sample-Efficient Black-Box Attacks on Visual Reinforcement LearningTairan Huang, Yulin Jin, Junxu Liu, Qingqing Ye et al.CVPR 2026
- Resolving the Stability-Plasticity Dilemma in Reinforcement Learning via Complementary Continual CriticsBo Sun, Peixi Peng, Guang Tan, Haoran Xu et al.CVPR 2026
- Semore: VLM-guided Enhanced Semantic Motion Representations for Visual Reinforcement LearningWentao Wang, Chunyang Liu, Kehua Sheng, Bo Zhang et al.AAAI 2026
Builds on24
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 1,261 citations
- Improving Sample Efficiency in Model-Free Reinforcement Learning from ImagesDenis Yarats, Amy Zhang, Ilya Kostrikov, Brandon Amos et al.AAAI 2021 · 506 citations
- Trajectory-guided Control Prediction for End-to-end Autonomous Driving: A Simple yet Strong BaselinePenghao Wu, Xiaosong Jia, Li Chen, Junchi Yan et al.NeurIPS 2022 · 444 citations
- Contrastive Learning as Goal-Conditioned Reinforcement LearningBenjamin Eysenbach, Tianjun Zhang, Sergey Levine, Ruslan SalakhutdinovNeurIPS 2022 · 331 citations
- Scalable Methods for Computing State Similarity in Deterministic Markov Decision ProcessesPablo Samuel CastroAAAI 2020 · 171 citations
Related papers
- Seek Commonality but Preserve Differences: Dissected Dynamics Modeling for Multi-modal Visual RLYangru Huang, Peixi Peng, Yifan Zhao, Guangyao Chen et al.NeurIPS 2024 · 3 citations
- Leveraging Modality-Specific Representations for Audio-Visual Speech Recognition via Reinforcement LearningChen Chen, Yuchen Hu, Qiang Zhang, Heqing Zou et al.AAAI 2023 · 35 citations
- Representation Learning for Event-based Visuomotor PoliciesSai Vemprala, Sami Mian, Ashish KapoorNeurIPS 2021 · 42 citations
- Leveraging Conditional Dependence for Efficient World Model DenoisingShaowei Zhang, Jiahan Cao, Dian Cheng, Xunlan Zhou et al.NeurIPS 2025
- Focus On What Matters: Separated Models For Visual-Based RL GeneralizationDi Zhang, Bowen Lv, Hai Zhang, Feifan Yang et al.NeurIPS 2024 · 14 citations
