Structuring Representation Geometry with Rotationally Equivariant Contrastive Learning
Sharut Gupta, Joshua Robinson, Derek Lim, Soledad Villar, Stefanie Jegelka
Abstract
Self-supervised learning converts raw perceptual data such as images to a compact space where simple Euclidean distances measure meaningful variations in data. In this paper, we extend this formulation by adding additional geometric structure to the embedding space by enforcing transformations of input space to correspond to simple (i.e., linear) transformations of embedding space. Specifically, in the contrastive learning setting, we introduce an equivariance objesctive and theoretically prove that its minima forces augmentations on input space to correspond to rotations on the spherical embedding space. We show that merely combining our equivariant loss with a non-collapse term results in non-trivial representations, without requiring invariance to data augmentations. Optimal performance is achieved by also encouraging approximate invariance, where input augmentations correspond to small rotations. Our method, Care: Contrastive Augmentation-induced Rotational Equivariance, leads to improved performance on downstream tasks, and ensures sensitivity in embedding space to important variations in data (e.g., color) that standard contrastive methods do not achieve. Code is available at https://github.com/Sharut/CARE . Recent contrastive self-supervised learning approaches have explored ways to close this gap by * Equal contribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers17
- Approximately Equivariant Graph NetworksNingyuan Huang, Ron Levie, Soledad VillarNeurIPS 2023 · 29 citations
- Self-supervised Transformation Learning for Equivariant RepresentationsJaemyung Yu, Jaehyun Choi, Dong-Jae Lee, Hyeong Gwon Hong et al.NeurIPS 2024 · 10 citations
- Understanding the Role of Equivariance in Self-supervised LearningYifei Wang, Kaiwen Hu, Sharut Gupta, Ziyu Ye et al.NeurIPS 2024 · 10 citations
- seq-JEPA: Autoregressive Predictive Learning of Invariant-Equivariant World ModelsHafez Ghaemi, Eilif B. Muller, Shahab BakhtiariNeurIPS 2025 · 8 citations
- In-Context Symmetries: Self-Supervised Learning through Contextual World ModelsSharut Gupta, Chenyu Wang, Yifei Wang, Tommi S. Jaakkola et al.NeurIPS 2024 · 8 citations
Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 2,360 citations
Related papers
- Improving Transformation Invariance in Contrastive Representation LearningAdam Foster, Rattana Pukdee, Tom RainforthICLR 2021 · 25 citations
- EquiAV: Leveraging Equivariance for Audio-Visual Contrastive LearningJongsuk Kim, Hyeongkeun Lee, Kyeongha Rho, Junmo Kim et al.ICML 2024 · 15 citations
- EquiMod: An Equivariance Module to Improve Visual Instance DiscriminationAlexandre Devillers, Mathieu LefortICLR 2023 · 2 citations
- Understanding Contrastive Learning Requires Incorporating Inductive BiasesNikunj Saunshi, Jordan T. Ash, Surbhi Goel, Dipendra Misra et al.ICML 2022 · 130 citations
- What Should Not Be Contrastive in Contrastive LearningTete Xiao, Xiaolong Wang, Alexei A. Efros, Trevor DarrellICLR 2021 · 338 citations
