OFVL-MS: Once for Visual Localization across Multiple Indoor Scenes
Tao Xie, Kun Dai, Siyi Lu, Ke Wang, Zhiqiang Jiang, Jinghan Gao, Dedong Liu, Jie Xu, Lijun Zhao, Ruifeng Li
Abstract
In this work, we seek to predict camera poses across scenes with a multi-task learning manner, where we view the localization of each scene as a new task. We propose OFVL-MS, a unified framework that dispenses with the traditional practice of training a model for each individual scene and relieves gradient conflict induced by optimizing multiple scenes collectively, enabling efficient storage yet precise visual localization for all scenes. Technically, in the forward pass of OFVL-MS, we design a layer-adaptive sharing policy with a learnable score for each layer to automatically determine whether the layer is shared or not. Such sharing policy empowers us to acquire task-shared parameters for a reduction of storage cost and task-specific parameters for learning scene-related features to alleviate gradient conflict. In the backward pass of OFVL-MS, we introduce a gradient normalization algorithm that homogenizes the gradient magnitude of the task-shared parameters so that all tasks converge at the same pace. Furthermore, a sparse penalty loss is applied on the learnable scores to facilitate parameter sharing for all tasks without performance degradation. We conduct comprehensive experiments on multiple benchmarks and our new released indoor dataset LIVL, showing that OFVL-MS families significantly outperform the state-of-the-arts with fewer parameters. We also verify that OFVL-MS can generalize to a new scene with much few parameters while gaining superior localization performance. The dataset and evaluation code is available at https://github.com/mooncake199809/UFVL-Net .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- From Sparse to Dense: Camera Relocalization with Scene-Specific Detector from Feature Gaussian SplattingZhiwei Huang, Hailin Yu, Yichun Shentu, Jin Yuan et al.CVPR 2025
- Mixture of Submodules for Domain Adaptive Person SearchMinsu Kim, Seungryong Kim, Kwanghoon SohnCVPR 2025
Builds on19
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Conflict-Averse Gradient Descent for Multi-task learningBo Liu, Xingchao Liu, Xiaojie Jin, Peter Stone et al.NeurIPS 2021 · 686 citations
- Efficiently Identifying Task Groupings for Multi-Task LearningChris Fifty, Ehsan Amid, Zhe Zhao, Tianhe Yu et al.NeurIPS 2021 · 352 citations
- Just Pick a Sign: Optimizing Deep Multitask Models with Gradient Sign DropoutZhao Chen, Jiquan Ngiam, Yanping Huang, Thang Luong et al.NeurIPS 2020 · 313 citations
- Towards Impartial Multi-task LearningLiyang Liu, Yi Li, Zhanghui Kuang, Jing-Hao Xue et al.ICLR 2021 · 228 citations
Related papers
- Simultaneous Scene-independent Camera Localization and Category-level Object Pose Estimation via Multi-level Feature FusionJunyi Wang, Yue QiIEEE VR 2023 · 6 citations
- MDL-NAS: A Joint Multi-domain Learning Framework for Vision TransformerShiguang Wang, Tao Xie, Jian Cheng, Xingcheng Zhang et al.CVPR 2023
- Towards Consistent Multi-Task Learning: Unlocking the Potential of Task-Specific ParametersXiaohan Qin, Xiaoxing Wang, Junchi YanCVPR 2025
- LSM: Learning Subspace Minimization for Low-Level VisionChengzhou Tang, Lu Yuan, Ping TanCVPR 2020
- FULLER: Unified Multi-modality Multi-task 3D Perception via Multi-level Gradient CalibrationZhijian Huang, Sihao Lin, Guiyu Liu, Mukun Luo et al.ICCV 2023 · 18 citations
