LoCo: Learning 3D Location-Consistent Image Features with a Memory-Efficient Ranking Loss
Dominik A. Kloepfer, João F. Henriques, Dylan Campbell
摘要
Image feature extractors are rendered substantially more useful if different views of the same 3D location yield similar features while still being distinct from other locations. A feature extractor that achieves this goal even under significant viewpoint changes must recognise not just semantic categories in a scene, but also understand how different objects relate to each other in three dimensions. Existing work addresses this task by posing it as a patch retrieval problem, training the extracted features to facilitate retrieval of all image patches that project from the same 3D location. However, this approach uses a loss formulation that requires substantial memory and computation resources, limiting its applicability for large-scale training. We present a method for memory-efficient learning of location-consistent features that reformulates and approximates the smooth average precision objective. This novel loss function enables improvements in memory efficiency by three orders of magnitude, mitigating a key bottleneck of previous methods and allowing much larger models to be trained with the same computational resources. We showcase the improved location consistency of our trained feature extractor directly on a multi-view consistency task, as well as the downstream task of scene-stable panoptic segmentation, significantly outperforming previous state-of-the-art.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper14
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Contrastive Learning with Hard Negative SamplesJoshua David Robinson, Ching-Yao Chuang, Suvrit Sra, Stefanie JegelkaICLR 2021 · 被引用 999 次
- Hard Negative Mixing for Contrastive LearningYannis Kalantidis, Mert Bülent Sariyildiz, Noé Pion, Philippe Weinzaepfel 等NeurIPS 2020 · 被引用 805 次
- Decomposing NeRF for Editing via Feature Field DistillationSosuke Kobayashi, Eiichi Matsumoto, Vincent SitzmannNeurIPS 2022 · 被引用 479 次
相关 Paper
- LoCUS: Learning Multiscale 3D-consistent Features from Posed ImagesDominik A. Kloepfer, Dylan Campbell, João F. HenriquesICCV 2023 · 被引用 1 次
- EPS3D: End-to-End Feed-Forward 3D Panoptic SegmentationRunsong Zhu, Jiaxin GUO, Xiaoyang Guo, Zhengzhe Liu 等ICML 2026 · 被引用 3 次
- CrOC: Cross-View Online Clustering for Dense Visual Representation LearningThomas Stegmüller, Tim Lebailly, Behzad Bozorgtabar, Tinne Tuytelaars 等CVPR 2023
- Multi-view Consistent 3D Panoptic Scene UnderstandingXianzhu Liu, Xin Sun, Haozhe Xie, Zonglin Li 等AAAI 2025 · 被引用 6 次
- CoLoR: The Devil is in Scene Coordinate Regression for Large-Scale Visual LocalizationXindong Mao, Hang Li, Yuchen Wu, Jiahe Li 等CVPR 2026
