Simultaneous Scene-independent Camera Localization and Category-level Object Pose Estimation via Multi-level Feature Fusion
Junyi Wang, Yue Qi
Abstract
In AR/MR applications, camera localization and object pose estimation both play crucial roles. The universality of learning techniques, often referred to as scene-independent localization and category-level pose estimation, presents challenges for both tasks. The two missions maintain close relationships due to the spatial geometry constraint, but differing task requirements result in distinct feature extraction. In this paper, we focus on simultaneous scene-independent camera localization and category-level object pose estimation with a unified learning framework. The system consists of a localization branch called SLO-LocNet, a pose estimation branch called SLO-ObjNet, a feature fusion module for feature sharing between two tasks, and two decoders for creating coordinate maps. In SLO-LocNet, localization features are produced for anticipating the relative pose between two adjusted frames using inputs of color and depth images. Furthermore, we establish an image fusion module in order to promote feature sharing in depth and color branches. With SLO-ObjNet, we take the detected depth image and its corresponding point cloud as inputs, and produce object pose features for pose estimation. A geometry fusion module is created to combine depth and point cloud information simultaneously. Between the two tasks, the image fusion module is also exploited to accomplish feature sharing. In terms of the loss function, we present a mixed optimization function that is composed of the relative camera pose, geometry constraint, absolute and relative object pose terms. To verify how well our algorithm could perform, we conduct experiments on both localization and pose estimation datasets, covering 7 Scenes, ScanNet, REAL275 and YCB-Video. All experiments demonstrate superior performance to other existing methods. We specifically train the network on ScanNet and test it on 7 Scenes to demonstrate the universality performance. Additionally, the positive effects of fusion modules and loss function are also demonstrated.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 6faa6310-925f-4e51-a97a-82a1df4a49d9Cited by top-tier papers1
Ask how each one uses itRelated papers
- SGLoc: Scene Geometry Encoding for Outdoor LiDAR LocalizationWen Li, Shangshu Yu, Cheng Wang, Guosheng Hu et al.CVPR 2023
- Feature-Level Collaboration: Joint Unsupervised Learning of Optical Flow, Stereo Depth and Camera MotionCheng Chi, Qingjie Wang, Tianyu Hao, Peng Guo et al.CVPR 2021
- Unified Category-Level Object Detection and Pose Estimation from RGB Images Using 3D PrototypesTom Fischer, Xiaojie Zhang, Eddy IlgICCV 2025 · 2 citations
- Leveraging SE(3) Equivariance for Self-supervised Category-Level Object Pose Estimation from Point CloudsXiaolong Li, Yijia Weng, Li Yi, Leonidas J. Guibas et al.NeurIPS 2021 · 61 citations
- SANet: Scene Agnostic Network for Camera LocalizationLuwei Yang, Ziqian Bai, Chengzhou Tang, Honghua Li et al.ICCV 2019 · 105 citations
