Simoun: Synergizing Interactive Motion-appearance Understanding for Vision-based Reinforcement Learning
Yangru Huang, Peixi Peng, Yifan Zhao, Yunpeng Zhai, Haoran Xu, Yonghong Tian
Abstract
Efficient motion and appearance modeling are critical for vision-based Reinforcement Learning (RL). However, existing methods struggle to reconcile motion and appearance information within the state representations learned from a single observation encoder. To address the problem, we present Synergizing Interactive Motion-appearance Understanding (Simoun), a unified framework for vision-based RL Given consecutive observation frames, Simoun deliberately and interactively learns both motion and appearance features through a dual-path network architecture. The learning process collaborates with a structural interactive module, which explores the latent motion-appearance structures from the two network paths to leverage their complementarity. To promote sample efficiency, we further design a consistency-guided curiosity module to encourage the exploration of under-learned observations. During training, the curiosity module provides intrinsic rewards according to the consistency of environmental temporal dynamics, which are deduced from both motion and appearance network paths. Experiments conducted on Deep-Mind control suite and CARLA automatic driving benchmarks demonstrate the effectiveness of Simoun, where it performs favorably against state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0130f59a-da0c-4625-9818-321c03d129c1Cited by top-tier papers4
- DMR: Decomposed Multi-Modality Representations for Frames and Events Fusion in Visual Reinforcement LearningHaoran Xu, Peixi Peng, Guang Tan, Yuan Li et al.CVPR 2024 · 5 citations
- DyMoDreamer: World Modeling with Dynamic ModulationBoxuan Zhang, Runqing Wang, Wei Xiao, Weipu Zhang et al.NeurIPS 2025 · 2 citations
- Semore: VLM-guided Enhanced Semantic Motion Representations for Visual Reinforcement LearningWentao Wang, Chunyang Liu, Kehua Sheng, Bo Zhang et al.AAAI 2026
- VLMs-Guided Representation Distillation for Efficient Vision-Based Reinforcement LearningHaoran Xu, Peixi Peng, Guang Tan, Yiqian Chang et al.CVPR 2025
Related papers
- Learning Generalizable Representations for Reinforcement Learning via Adaptive Meta-learner of Behavioral SimilaritiesJianda Chen, Sinno Jialin PanICLR 2022 · 6 citations
- Active Vision Reinforcement Learning under Limited Visual ObservabilityJinghuan Shang, Michael S. RyooNeurIPS 2023 · 1 citation
- Resolving the Stability-Plasticity Dilemma in Reinforcement Learning via Complementary Continual CriticsBo Sun, Peixi Peng, Guang Tan, Haoran Xu et al.CVPR 2026
- Stabilizing Visual Reinforcement Learning via Asymmetric Interactive CooperationYunpeng Zhai, Peixi Peng, Yifan Zhao, Yangru Huang et al.ICCV 2023 · 4 citations
- Curious Representation Learning for Embodied IntelligenceYilun Du, Chuang Gan, Phillip IsolaICCV 2021 · 50 citations
