Alligat0R: Pre-Training through Covisibility Segmentation for Relative Camera Pose Regression
Thibaut Loiseau, Guillaume Bourmaud, Vincent Lepetit
Abstract
Pre-training techniques have greatly advanced computer vision, with CroCo's cross-view completion approach yielding impressive results in tasks like 3D reconstruction and pose regression. However, cross-view completion is ill-posed in non-covisible regions, limiting its effectiveness. We introduce Alligat0R, a novel pre-training approach that replaces cross-view learning with a covisibility segmentation task. Our method predicts whether each pixel in one image is covisible in the second image, occluded, or outside the field of view, making the pre-training effective in both covisible and non-covisible regions, and provides interpretable predictions. To support this, we present Cub3, a large-scale dataset with 5M image pairs and dense covisibility annotations derived from the nuScenes and ScanNet datasets. Cub3 includes diverse scenarios with varying degrees of overlap. The experiments show that our novel pre-training method Alligat0R significantly outperforms CroCo in relative pose regression. Code is available at https://github.com/thibautloiseau/alligat0r.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 64323ab2-3deb-42fa-8570-9a896cd8bc2dCited by top-tier papers3
- TTT3R: 3D Reconstruction as Test-Time TrainingXingyu Chen, Yue Chen, Yuliang Xiu, Andreas Geiger et al.ICLR 2026 · 139 citations
- AMB3R: Accurate Feed-forward Metric-scale 3D Reconstruction with BackendHengyi Wang, Lourdes AgapitoCVPR 2026 · 17 citations
- MuM: Multi-View Masked Image Modeling for 3D VisionDavid Nordström, Johan Edstedt, Fredrik Kahl, Georg BökmanCVPR 2026 · 6 citations
Builds on37
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
Related papers
- CroCoDiLight: Repurposing Cross-View Completion Encoders for RelightingAlistair J. Foggin, William SmithICLR 2026
- Building a Strong Pre-Training Baseline for Universal 3D Large-Scale PerceptionHaoming Chen, Zhizhong Zhang, Yanyun Qu, Ruixin Zhang et al.CVPR 2024
- CroCo: Self-Supervised Pre-training for 3D Vision Tasks by Cross-View CompletionPhilippe Weinzaepfel, Vincent Leroy, Thomas Lucas, Romain Brégier et al.NeurIPS 2022 · 189 citations
- Multi-Modal Contrastive Masked Autoencoders: A Two-Stage Progressive Pre-training Approach for RGBD DatasetsMuhammad Abdullah Jamal, Omid MohareriCVPR 2025
- Scene-agnostic Pose Regression for Visual LocalizationJunwei Zheng, Ruiping Liu, Yufan Chen, Zhenfang Chen et al.CVPR 2025
