OFVL-MS: Once for Visual Localization across Multiple Indoor Scenes
Tao Xie, Kun Dai, Siyi Lu, Ke Wang, Zhiqiang Jiang, Jinghan Gao, Dedong Liu, Jie Xu, Lijun Zhao, Ruifeng Li
摘要
In this work, we seek to predict camera poses across scenes with a multi-task learning manner, where we view the localization of each scene as a new task. We propose OFVL-MS, a unified framework that dispenses with the traditional practice of training a model for each individual scene and relieves gradient conflict induced by optimizing multiple scenes collectively, enabling efficient storage yet precise visual localization for all scenes. Technically, in the forward pass of OFVL-MS, we design a layer-adaptive sharing policy with a learnable score for each layer to automatically determine whether the layer is shared or not. Such sharing policy empowers us to acquire task-shared parameters for a reduction of storage cost and task-specific parameters for learning scene-related features to alleviate gradient conflict. In the backward pass of OFVL-MS, we introduce a gradient normalization algorithm that homogenizes the gradient magnitude of the task-shared parameters so that all tasks converge at the same pace. Furthermore, a sparse penalty loss is applied on the learnable scores to facilitate parameter sharing for all tasks without performance degradation. We conduct comprehensive experiments on multiple benchmarks and our new released indoor dataset LIVL, showing that OFVL-MS families significantly outperform the state-of-the-arts with fewer parameters. We also verify that OFVL-MS can generalize to a new scene with much few parameters while gaining superior localization performance. The dataset and evaluation code is available at https://github.com/mooncake199809/UFVL-Net .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- From Sparse to Dense: Camera Relocalization with Scene-Specific Detector from Feature Gaussian SplattingZhiwei Huang, Hailin Yu, Yichun Shentu, Jin Yuan 等CVPR 2025
- Mixture of Submodules for Domain Adaptive Person SearchMinsu Kim, Seungryong Kim, Kwanghoon SohnCVPR 2025
它引用的顶会 Paper19
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Conflict-Averse Gradient Descent for Multi-task learningBo Liu, Xingchao Liu, Xiaojie Jin, Peter Stone 等NeurIPS 2021 · 被引用 686 次
- Efficiently Identifying Task Groupings for Multi-Task LearningChris Fifty, Ehsan Amid, Zhe Zhao, Tianhe Yu 等NeurIPS 2021 · 被引用 352 次
- Just Pick a Sign: Optimizing Deep Multitask Models with Gradient Sign DropoutZhao Chen, Jiquan Ngiam, Yanping Huang, Thang Luong 等NeurIPS 2020 · 被引用 313 次
- Towards Impartial Multi-task LearningLiyang Liu, Yi Li, Zhanghui Kuang, Jing-Hao Xue 等ICLR 2021 · 被引用 228 次
相关 Paper
- Simultaneous Scene-independent Camera Localization and Category-level Object Pose Estimation via Multi-level Feature FusionJunyi Wang, Yue QiIEEE VR 2023 · 被引用 6 次
- MDL-NAS: A Joint Multi-domain Learning Framework for Vision TransformerShiguang Wang, Tao Xie, Jian Cheng, Xingcheng Zhang 等CVPR 2023
- Towards Consistent Multi-Task Learning: Unlocking the Potential of Task-Specific ParametersXiaohan Qin, Xiaoxing Wang, Junchi YanCVPR 2025
- LSM: Learning Subspace Minimization for Low-Level VisionChengzhou Tang, Lu Yuan, Ping TanCVPR 2020
- FULLER: Unified Multi-modality Multi-task 3D Perception via Multi-level Gradient CalibrationZhijian Huang, Sihao Lin, Guiyu Liu, Mukun Luo 等ICCV 2023 · 被引用 18 次
