Self-supervised Transformation Learning for Equivariant Representations
Jaemyung Yu, Jaehyun Choi, Dong-Jae Lee, Hyeong Gwon Hong, Junmo Kim
Abstract
Unsupervised representation learning has significantly advanced various machine learning tasks. In the computer vision domain, state-of-the-art approaches utilize transformations like random crop and color jitter to achieve invariant representations, embedding semantically the same inputs despite transformations. However, this can degrade performance in tasks requiring precise features, such as localization or flower classification. To address this, recent research incorporates equivariant representation learning, which captures transformation-sensitive information. However, current methods depend on transformation labels and thus struggle with interdependency and complex transformations. We propose Self-supervised Transformation Learning (STL), replacing transformation labels with transformation representations derived from image pairs. The proposed method ensures transformation representation is image-invariant and learns corresponding equivariant transformations, enhancing performance without increased batch complexity. We demonstrate the approach’s effectiveness across diverse classification and detection tasks, outperforming existing methods in 7 out of 11 benchmarks and excelling in detection. By integrating complex transformations like AugMix, unusable by prior equivariant methods, this approach enhances performance across tasks, underscoring its adaptability and resilience. Additionally, its compatibility with various base models highlights its flexibility and broad applicability. The code is available at https://github.com/jaemyung-u/stl .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2bc18da8-b840-4f85-815d-cf6407c0772bCited by top-tier papers5
- Equivariance by Contrast: Identifiable Equivariant Embeddings from Unlabeled Finite Group ActionsTobias Schmidt, Steffen Schneider, Matthias BethgeNeurIPS 2025 · 2 citations
- Soft Equivariance Regularization for Invariant Self-Supervised LearningJoohyung Lee, Changhun Kim, Hyunsu Kim, Kwanhyung Lee et al.ICLR 2026 · 1 citation
- Learning Without Augmenting: Unsupervised Time Series Representation Learning via Frame ProjectionsBerken Utku Demirel, Christian HolzNeurIPS 2025 · 1 citation
- Equicaps: Predictor-Free Pose-Aware Pre-Trained Capsule NetworksAthinoulla Konstantinou, Georgios Leontidis, Mamatha Thota, Aiden DurrantICCV 2025
- Soft Task-Aware Routing of Experts for Equivariant Representation LearningJaebyeong Jeon, Hyunseo Jang, Jy-yong Sohn, Kibok LeeNeurIPS 2025
Builds on14
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun et al.ICML 2021 · 2,942 citations
- AugMix: A Simple Data Processing Method to Improve Robustness and UncertaintyDan Hendrycks, Norman Mu, Ekin Dogus Cubuk, Barret Zoph et al.ICLR 2020 · 1,572 citations
Related papers
- Improving Transferability of Representations via Augmentation-Aware Self-SupervisionHankook Lee, Kibok Lee, Kimin Lee, Honglak Lee et al.NeurIPS 2021 · 66 citations
- Self-Supervised Learning of Pretext-Invariant RepresentationsIshan Misra, Laurens van der MaatenCVPR 2020
- Self-Supervised Learning Based on Transformed Image Reconstruction for Equivariance-Coherent Feature RepresentationQin Wang, Alessio Quercia, Benjamin Bruns, Abigail Morrison et al.AAAI 2026 · 2 citations
- Spatially Consistent Representation LearningByungseok Roh, Wuhyun Shin, Ildoo Kim, Sungwoong KimCVPR 2021
- EquiAV: Leveraging Equivariance for Audio-Visual Contrastive LearningJongsuk Kim, Hyeongkeun Lee, Kyeongha Rho, Junmo Kim et al.ICML 2024 · 15 citations
