Self-Supervised Learning Disentangled Group Representation as Feature
Tan Wang, Zhongqi Yue, Jianqiang Huang, Qianru Sun, Hanwang Zhang
Abstract
A good visual representation is an inference map from observations (images) to features (vectors) that faithfully reflects the hidden modularized generative factors (semantics). In this paper, we formulate the notion of "good" representation from a group-theoretic view using Higgins' definition of disentangled representation [42], and show that existing Self-Supervised Learning (SSL) only disentangles simple augmentation features such as rotation and colorization, thus unable to modularize the remaining semantics. To break the limitation, we propose an iterative SSL algorithm: Iterative Partition-based Invariant Risk Minimization (IP-IRM), which successfully grounds the abstract semantics and the group acting on them into concrete contrastive learning. At each iteration, IP-IRM first partitions the training samples into two subsets that correspond to an entangled group element. Then, it minimizes a subset-invariant contrastive loss, where the invariance guarantees to disentangle the group element. We prove that IP-IRM converges to a fully disentangled representation and show its effectiveness on various benchmarks. Codes are available at https://github.com/Wangt-CN/IP-IRM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1d266a71-ab1d-4533-af06-b282b17550f3Cited by top-tier papers28
- Discovering Invariant Rationales for Graph Neural NetworksYingxin Wu, Xiang Wang, An Zhang, Xiangnan He et al.ICLR 2022 · 313 citations
- Change is Hard: A Closer Look at Subpopulation ShiftYuzhe Yang, Haoran Zhang, Dina Katabi, Marzyeh GhassemiICML 2023 · 149 citations
- Equivariant Similarity for Vision-Language Foundation ModelsTan Wang, Kevin Lin, Linjie Li, Chung-Ching Lin et al.ICCV 2023 · 67 citations
- Identifiability of Label Noise Transition MatrixYang Liu, Hao Cheng, Kun ZhangICML 2023 · 58 citations
- Make the U in UDA Matter: Invariant Consistency Learning for Unsupervised Domain AdaptationZhongqi Yue, Qianru Sun, Hanwang ZhangNeurIPS 2023 · 39 citations
Builds on28
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- What Makes for Good Views for Contrastive Learning?Yonglong Tian, Chen Sun, Ben Poole, Dilip Krishnan et al.NeurIPS 2020 · 1,631 citations
Related papers
- Self-Supervised Learning of Pretext-Invariant RepresentationsIshan Misra, Laurens van der MaatenCVPR 2020
- An Investigation of Representation and Allocation Harms in Contrastive LearningSubha Maity, Mayank Agarwal, Mikhail Yurochkin, Yuekai SunICLR 2024 · 2 citations
- An Information Criterion for Controlled Disentanglement of Multimodal DataChenyu Wang, Sharut Gupta, Xinyi Zhang, Sana Tonekaboni et al.ICLR 2025
- STiL: Semi-supervised Tabular-Image Learning for Comprehensive Task-Relevant Information Exploration in Multimodal ClassificationSiyi Du, Xinzhe Luo, Declan P. O'Regan, Chen QinCVPR 2025
- Fair Contrastive Learning for Facial Attribute ClassificationSungho Park, Jewook Lee, Pilhyeon Lee, Sunhee Hwang et al.CVPR 2022 · 61 citations
