Steering Self-Supervised Feature Learning Beyond Local Pixel Statistics
Simon Jenni, Hailin Jin, Paolo Favaro
Abstract
We introduce a novel principle for self-supervised feature learning based on the discrimination of specific transformations of an image. We argue that the generalization capability of learned features depends on what image neighborhood size is sufficient to discriminate different image transformations: The larger the required neighborhood size and the more global the image statistics that the feature can describe. An accurate description of global image statistics allows to better represent the shape and configuration of objects and their context, which ultimately generalizes better to new tasks such as object classification and detection. This suggests a criterion to choose and design image transformations. Based on this criterion, we introduce a novel image transformation that we call limited context inpainting (LCI). This transformation inpaints an image patch conditioned only on a small rectangular pixel boundary (the limited context). Because of the limited boundary information, the inpainter can learn to match local pixel statistics, but is unlikely to match the global statistics of the image. We claim that the same principle can be used to justify the performance of transformations such as image rotations and warping. Indeed, we demonstrate experimentally that learning to discriminate transformations such as LCI, image warping and rotations, yields features with state of the art generalization capabilities on several datasets such as Pascal VOC, STL-10, CelebA, and ImageNet. Remarkably, our trained features achieve a performance on Places on par with features trained through supervised learning with ImageNet labels.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Self-supervised 3D Skeleton Action Representation Learning with Motion Consistency and ContinuityYukun Su, Guosheng Lin, Qingyao WuICCV 2021 · 86 citations
- Enhancing Unsupervised Video Representation Learning by Decoupling the Scene and the MotionJinpeng Wang, Yuting Gao, Ke Li, Jianguo Hu et al.AAAI 2021 · 70 citations
- Time-Equivariant Contrastive Video Representation LearningSimon Jenni, Hailin JinICCV 2021 · 64 citations
- Audio-Visual Contrastive Learning with Temporal Self-SupervisionSimon Jenni, Alexander Black, John P. CollomosseAAAI 2023 · 25 citations
- Self-Supervised Object Localization with Joint Graph PartitionYukun Su, Guosheng Lin, Yun Hao, Yiwen Cao et al.AAAI 2022 · 17 citations
Builds on3
- Data-Efficient Image Recognition with Contrastive Predictive CodingOlivier J. HénaffICML 2020 · 1,553 citations
- Local Aggregation for Unsupervised Learning of Visual EmbeddingsChengxu Zhuang, Alex Lin Zhai, Daniel YaminsICCV 2019 · 462 citations
- Detecting Photoshopped Faces by Scripting PhotoshopSheng-Yu Wang, Oliver Wang, Richard Zhang, Andrew Owens et al.ICCV 2019 · 147 citations
Related papers
- Towards Generic Image Manipulation Detection with Weakly-Supervised Self-Consistency LearningYuanhao Zhai, Tianyu Luan, David S. Doermann, Junsong YuanICCV 2023 · 35 citations
- VICRegL: Self-Supervised Learning of Local Visual FeaturesAdrien Bardes, Jean Ponce, Yann LeCunNeurIPS 2022 · 189 citations
- Instance Localization for Self-Supervised Detection PretrainingCeyuan Yang, Zhirong Wu, Bolei Zhou, Stephen LinCVPR 2021
- Self-Supervised Learning of Intertwined Content and Positional Features for Object DetectionKang-Jun Liu, Masanori Suganuma, Takayuki OkataniICML 2025
- Patch-level Representation Learning for Self-supervised Vision TransformersSukmin Yun, Hankook Lee, Jaehyung Kim, Jinwoo ShinCVPR 2022 · 52 citations
