LoCo: Learning 3D Location-Consistent Image Features with a Memory-Efficient Ranking Loss
Dominik A. Kloepfer, João F. Henriques, Dylan Campbell
Abstract
Image feature extractors are rendered substantially more useful if different views of the same 3D location yield similar features while still being distinct from other locations. A feature extractor that achieves this goal even under significant viewpoint changes must recognise not just semantic categories in a scene, but also understand how different objects relate to each other in three dimensions. Existing work addresses this task by posing it as a patch retrieval problem, training the extracted features to facilitate retrieval of all image patches that project from the same 3D location. However, this approach uses a loss formulation that requires substantial memory and computation resources, limiting its applicability for large-scale training. We present a method for memory-efficient learning of location-consistent features that reformulates and approximates the smooth average precision objective. This novel loss function enables improvements in memory efficiency by three orders of magnitude, mitigating a key bottleneck of previous methods and allowing much larger models to be trained with the same computational resources. We showcase the improved location consistency of our trained feature extractor directly on a multi-view consistency task, as well as the downstream task of scene-stable panoptic segmentation, significantly outperforming previous state-of-the-art.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on14
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Contrastive Learning with Hard Negative SamplesJoshua David Robinson, Ching-Yao Chuang, Suvrit Sra, Stefanie JegelkaICLR 2021 · 999 citations
- Hard Negative Mixing for Contrastive LearningYannis Kalantidis, Mert Bülent Sariyildiz, Noé Pion, Philippe Weinzaepfel et al.NeurIPS 2020 · 805 citations
- Decomposing NeRF for Editing via Feature Field DistillationSosuke Kobayashi, Eiichi Matsumoto, Vincent SitzmannNeurIPS 2022 · 479 citations
Related papers
- LoCUS: Learning Multiscale 3D-consistent Features from Posed ImagesDominik A. Kloepfer, Dylan Campbell, João F. HenriquesICCV 2023 · 1 citation
- EPS3D: End-to-End Feed-Forward 3D Panoptic SegmentationRunsong Zhu, Jiaxin GUO, Xiaoyang Guo, Zhengzhe Liu et al.ICML 2026 · 3 citations
- CrOC: Cross-View Online Clustering for Dense Visual Representation LearningThomas Stegmüller, Tim Lebailly, Behzad Bozorgtabar, Tinne Tuytelaars et al.CVPR 2023
- Multi-view Consistent 3D Panoptic Scene UnderstandingXianzhu Liu, Xin Sun, Haozhe Xie, Zonglin Li et al.AAAI 2025 · 6 citations
- CoLoR: The Devil is in Scene Coordinate Regression for Large-Scale Visual LocalizationXindong Mao, Hang Li, Yuchen Wu, Jiahe Li et al.CVPR 2026
