From Motion to Localization: Cross-view Optimization with Stationary Event and RGB Cameras for Enhanced Pose Estimation
Yukun Zhao, Xinyuan Song, Huajian Huang, Tristan Braud
Abstract
Applications such as Augmented Reality (AR) require accurate device positioning to minimize alignment errors. While visual positioning techniques offer high accuracy, their performance can degrade due to environmental changes like lighting variations and object movements. This paper introduces a new approach to visual positioning, relying on a stationary joint event/RGB sensing platform to track scene dynamics in real-time. This platform is at the core of a localization pipeline to predict the pose of user devices. First, a cross-modal object tracker matches dynamic objects between RGB and event images captured by the platform. These objects contribute to building a dynamic map, combined with the initial static 3D Structure from Motion (SfM) model to form a global feature map. Finally, a cross-view pose optimizer estimates pose uncertainties between modalities to refine and improve localization accuracy. To validate our approach, we collect a large-scale dataset over three scenes to account for typical AR scenarios where dynamics can affect the quality of visual positioning. We contribute this dataset to the community for future research on scene dynamics. Our approach shows significant improvement over existing methods, reducing translation and rotation errors by 12.9% and 13.4%, respectively, for weekly data over 4 weeks, and by 38.5% and 16.2% for monthly data over 4 months, compared to HLoc (SP+SG). It also reduces performance degradation by up to 50% after only 4 weeks.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 074cacc3-d03b-4cc6-afb4-bfc256fcd116Related papers
- EgoPoseVR: Spatiotemporal Multi-Modal Reasoning for Egocentric Full-Body Pose in Virtual RealityHaojie Cheng, Shaun Jing Heng Ong, Shaoyu Cai, Aiden Tat Yang Koh et al.IEEE VR 2026 · 1 citation
- CroCoDL: Cross-device Collaborative Dataset for LocalizationHermann Blum, Alessandro Mercurio, Joshua O'Reilly, Tim Engelbracht et al.CVPR 2025
- RGB2LIDAR: Towards Solving Large-Scale Cross-Modal Visual LocalizationNiluthpol Chowdhury Mithun, Karan Sikka, Han-Pang Chiu, Supun Samarasekera et al.ACM MM 2020 · 22 citations
- 360Loc: A Dataset and Benchmark for Omnidirectional Visual Localization with Cross-Device QueriesHuajian Huang, Changkun Liu, Yipeng Zhu, Hui Cheng et al.CVPR 2024
- CrowdDriven: A New Challenging Dataset for Outdoor Visual LocalizationAra Jafarzadeh, Manuel López-Antequera, Pau Gargallo, Yubin Kuang et al.ICCV 2021 · 18 citations
