CrIBo: Self-Supervised Learning via Cross-Image Object-Level Bootstrapping
Tim Lebailly, Thomas Stegmüller, Behzad Bozorgtabar, Jean-Philippe Thiran, Tinne Tuytelaars
摘要
Leveraging nearest neighbor retrieval for self-supervised representation learning has proven beneficial with object-centric images. However, this approach faces limitations when applied to scene-centric datasets, where multiple objects within an image are only implicitly captured in the global representation. Such global bootstrapping can lead to undesirable entanglement of object representations. Furthermore, even object-centric datasets stand to benefit from a finer-grained bootstrapping approach. In response to these challenges, we introduce a novel Cross-Image Object-Level Bootstrapping method tailored to enhance dense visual representation learning. By employing object-level nearest neighbor bootstrapping throughout the training, CrIBo emerges as a notably strong and adequate candidate for in-context learning, leveraging nearest neighbor retrieval at test time. CrIBo shows state-of-the-art performance on the latter task while being highly competitive in more standard downstream segmentation tasks. Our code and pretrained models are publicly available at https://github.com/tileb1/CrIBo . * denotes equal contribution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Exploring Structural Degradation in Dense Representations for Self-supervised LearningSiran Dai, Qianqian Xu, Peisong Wen, Yang Liu 等NeurIPS 2025 · 被引用 5 次
- CG-SSL: Concept-Guided Self-Supervised LearningSara Atito, Josef Kittler, Imran Razzak, Muhammad AwaisNeurIPS 2025 · 被引用 1 次
- UniRefiner: Teaching Pre-trained ViTs to Self-Dispose Dross via Contrastive RegisterCongpei Qiu, Zhaoyu Hu, Wei Ke, Zhuotao Tian 等CVPR 2026
- Featurising Pixels from Dynamic 3D Scenes with Linear In-Context LearnersNikita Araslanov, Martin Sundermeyer, Hidenobu Matsuki, David Joseph Tan 等CVPR 2026
- Mosic: Optimal-Transport Motion Trajectory for Dense Self-Supervised LearningMohammadreza Salehi, Shashanka Venkataramanan, Ioana Simion, Efstratios Gavves 等ICCV 2025
它引用的顶会 Paper30
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal 等NeurIPS 2020 · 被引用 5,249 次
相关 Paper
- Unsupervised Object-Level Representation Learning from Scene ImagesJiahao Xie, Xiaohang Zhan, Ziwei Liu, Yew Soon Ong 等NeurIPS 2021 · 被引用 93 次
- Adaptive Similarity Bootstrapping for Self-Distillation based Representation LearningTim Lebailly, Thomas Stegmüller, Behzad Bozorgtabar, Jean-Philippe Thiran 等ICCV 2023 · 被引用 3 次
- Dense Semantic Contrast for Self-Supervised Visual Representation LearningXiaoni Li, Yu Zhou, Yifei Zhang, Aoting Zhang 等ACM MM 2021 · 被引用 35 次
- UniVIP: A Unified Framework for Self-Supervised Visual Pre-trainingZhaowen Li, Yousong Zhu, Fan Yang, Wei Li 等CVPR 2022 · 被引用 29 次
- Unsupervised Learning of Dense Visual RepresentationsPedro O. Pinheiro, Amjad Almahairi, Ryan Y. Benmalek, Florian Golemo 等NeurIPS 2020 · 被引用 227 次
