Dual-Agent Reinforcement Learning for Adaptive and Cost-Aware Visual-Inertial Odometry
Feiyang Pan, Shenghe Zheng, Chunyan Yin, Guangbin Dou
Abstract
Visual-Inertial Odometry (VIO) is a critical component for robust ego-motion estimation, enabling foundational capabilities such as autonomous navigation in robotics and real-time 6-DoF tracking for augmented reality. Existing methods face a well-known trade-off: filter-based approaches are efficient but prone to drift, while optimization-based methods, though accurate, rely on computationally prohibitive Visual-Inertial Bundle Adjustment (VIBA) that is difficult to run on resource-constrained platforms. Rather than removing VIBA altogether, we aim to reduce how often and how heavily it must be invoked. To this end, we cast two key design choices in modern VIO, when to run the visual frontend and how strongly to trust its output, as sequential decision problems, and solve them with lightweight reinforcement learning (RL) agents. Our framework introduces a lightweight, dual-pronged RL policy that serves as our core contribution: (1) a Select Agent intelligently gates the entire VO pipeline based only on high-frequency IMU data; and (2) a composite Fusion Agent that first estimates a robust velocity state via a supervised network, before an RL policy adaptively fuses the full (p, v, q) state. Experiments on the EuRoC MAV and TUM-VI datasets show that, in our unified evaluation, the proposed method achieves a more favorable accuracy-efficiency-memory trade-off than prior GPU-based VO/VIO systems: it attains the best average ATE while running up to 1.77 times faster and using less GPU memory. Compared to classical optimization-based VIO systems, our approach maintains competitive trajectory accuracy while substantially reducing computational load.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4f2c36b2-cd20-4701-a6a1-aebdc6644b56Builds on6
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 1,248 citations
- Learning To Explore Using Active Neural SLAMDevendra Singh Chaplot, Dhiraj Gandhi, Saurabh Gupta, Abhinav Gupta et al.ICLR 2020 · 603 citations
- Deep Patch Visual OdometryZachary Teed, Lahav Lipson, Jia DengNeurIPS 2023 · 323 citations
- Adaptive VIO: Deep Visual-Inertial Odometry with Online Continual LearningYouqi Pan, Wugen Zhou, Yingdian Cao, Hongbin ZhaCVPR 2024 · 17 citations
- Robust Tightly-Coupled Visual-Inertial Odometry with Pre-built Maps in High Latency SituationsHujun Bao, Weijian Xie, Quanhao Qian, Danpeng Chen et al.IEEE VR 2022 · 16 citations
Related papers
- A Rotation-Translation-Decoupled Solution for Robust and Efficient Visual-Inertial InitializationYijia He, Bo Xu, Zhanpeng Ouyang, Hongdong LiCVPR 2023
- AVA-VLA: Improving Vision-Language-Action models with Active Visual AttentionLei Xiao, Jifeng Li, Juntao Gao, Feiyang Ye et al.CVPR 2026 · 26 citations
- UAV: A Unified and Adaptive Scheduling Framework for UAV Autopilot System with Reinforcement LearningZeying Li, shuai zhao, Chaowen Wu, Boyang Li et al.ICML 2026
- Information-Driven Direct RGB-D OdometryAlejandro Fontán, Javier Civera, Rudolph TriebelCVPR 2020
- 100-Phones: A Large VI-SLAM Dataset for Augmented Reality Towards Mass Deployment on Mobile PhonesGuofeng Zhang, Jin Yuan, Haomin Liu, Zhen Peng et al.IEEE VR 2024 · 5 citations
