Harmony: A Generic Unsupervised Approach for Disentangling Semantic Content from Parameterized Transformations
Mostofa Rafid Uddin, Gregory Howe, Xiangrui Zeng, Min Xu
Abstract
In many real-life image analysis applications, particularly in biomedical research domains, the objects of interest undergo multiple transformations that alters their visual properties while keeping the semantic content unchanged. Disentangling images into semantic content factors and transformations can provide significant benefits into many domain-specific image analysis tasks. To this end, we propose a generic unsupervised framework, Harmony, that simultaneously and explicitly disentangles semantic content from multiple parameterized transformations. Harmony leverages a simple cross-contrastive learning framework with multiple explicitly parameterized latent representations to disentangle content from transformations. To demonstrate the efficacy of Harmony, we apply it to disentangle image semantic content from several parameterized transformations (rotation, translation, scaling, and contrast). Harmony achieves significantly improved disentanglement over the baseline models on several image datasets of diverse domains. With such disentanglement, Harmony is demonstrated to incentivize bioimage analysis research by modeling structural heterogeneity of macromolecules from cryo-ET images and learning transformation-invariant representations of protein particles from single-particle cryo-EM images. Harmony also performs very well in disentangling content from 3D transformations and can perform coarse and fast alignment of 3D cryo-ET subtomograms. Therefore, Harmony is generalizable to many other imaging domains and can potentially be extended to domains beyond imaging as well.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 69253d0b-2e14-46f1-a081-4a8d76a59cd2Cited by top-tier papers2
- Amortized Inference for Heterogeneous Reconstruction in Cryo-EMAxel Levy, Gordon Wetzstein, Julien N. P. Martel, Frédéric Poitevin et al.NeurIPS 2022 · 57 citations
- Unsupervised Identification of Protein Compositions and Conformations Via Implicit Content-Transformation DisentanglementMostofa Rafid Uddin, Jana Armouti, Min XuICCV 2025
Builds on4
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun et al.ICML 2021 · 2,942 citations
- Gum-Net: Unsupervised Geometric Matching for Fast and Accurate 3D Subtomogram Image Alignment and AveragingXiangrui Zeng, Min XuCVPR 2020
Related papers
- Rotation and Translation Invariant Representation Learning with Implicit Neural RepresentationsSehyun Kwon, Joo Young Choi, Ernest K. RyuICML 2023 · 5 citations
- Unsupervised Object Representation Learning using Translation and Rotation Group Equivariant VAEAlireza Nasiri, Tristan BeplerNeurIPS 2022 · 18 citations
- Image Harmonization with TransformerZonghui Guo, Dongsheng Guo, Haiyong Zheng, Zhaorui Gu et al.ICCV 2021 · 95 citations
- End-to-end robust joint unsupervised image alignment and clusteringXiangrui Zeng, Gregory Howe, Min XuICCV 2021 · 12 citations
- DRANet: Disentangling Representation and Adaptation Networks for Unsupervised Cross-Domain AdaptationSeunghun Lee, Sunghyun Cho, Sunghoon ImCVPR 2021
