Learning Representations by Predicting Bags of Visual Words
Spyros Gidaris, Andrei Bursuc, Nikos Komodakis, Patrick Pérez, Matthieu Cord
摘要
Self-supervised representation learning targets to learn convnet-based image representations from unlabeled data. Inspired by the success of NLP methods in this area, in this work we propose a self-supervised approach based on spatially dense image descriptions that encode discrete visual concepts, here called visual words. To build such discrete representations, we quantize the feature maps of a first pretrained self-supervised convnet, over a k-means based vocabulary. Then, as a self-supervised task, we train another convnet to predict the histogram of visual words of an image (i.e., its Bag-of-Words representation) given as input a perturbed version of that image. The proposed task forces the convnet to learn perturbation-invariant and context-aware image features, useful for downstream image understanding tasks. We extensively evaluate our method and demonstrate very strong empirical results, e.g., our pre-trained self-supervised representations transfer better on detection task and similarly on classification over classes "unseen" during pre-training, when compared to the supervised case. This also shows that the process of image discretization into visual words can provide the basis for very powerful selfsupervised approaches in the image domain, thus allowing further connections to be made to related methods from the NLP domain that have been extremely successful so far. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper39
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal 等NeurIPS 2020 · 被引用 5,249 次
- VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised LearningAdrien Bardes, Jean Ponce, Yann LeCunICLR 2022 · 被引用 1,226 次
- With a Little Help from My Friends: Nearest-Neighbor Contrastive Learning of Visual RepresentationsDebidatta Dwibedi, Yusuf Aytar, Jonathan Tompson, Pierre Sermanet 等ICCV 2021 · 被引用 542 次
- Visual Prompting via Image InpaintingAmir Bar, Yossi Gandelsman, Trevor Darrell, Amir Globerson 等NeurIPS 2022 · 被引用 340 次
它引用的顶会 Paper11
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- Data-Efficient Image Recognition with Contrastive Predictive CodingOlivier J. HénaffICML 2020 · 被引用 1,553 次
- Self-labelling via simultaneous clustering and representation learningYuki Markus Asano, Christian Rupprecht, Andrea VedaldiICLR 2020 · 被引用 873 次
- S4L: Self-Supervised Semi-Supervised LearningLucas Beyer, Xiaohua Zhai, Avital Oliver, Alexander KolesnikovICCV 2019 · 被引用 854 次
- Local Aggregation for Unsupervised Learning of Visual EmbeddingsChengxu Zhuang, Alex Lin Zhai, Daniel YaminsICCV 2019 · 被引用 462 次
相关 Paper
- OBoW: Online Bag-of-Visual-Words Generation for Self-Supervised LearningSpyros Gidaris, Andrei Bursuc, Gilles Puy, Nikos Komodakis 等CVPR 2021
- VirTex: Learning Visual Representations From Textual AnnotationsKaran Desai, Justin JohnsonCVPR 2021
- Self-supervised Product Quantization for Deep Unsupervised Image RetrievalYoung Kyun Jang, Nam Ik ChoICCV 2021 · 被引用 90 次
- R-MAE: Regions Meet Masked AutoencodersDuy-Kien Nguyen, Yanghao Li, Vaibhav Aggarwal, Martin R. Oswald 等ICLR 2024 · 被引用 18 次
- Cross-Modal Discrete Representation LearningAlexander H. Liu, SouYoung Jin, Cheng-I Lai, Andrew Rouditchenko 等ACL 2022 · 被引用 57 次
