End-to-end robust joint unsupervised image alignment and clustering
Xiangrui Zeng, Gregory Howe, Min Xu
Abstract
Computing dense pixel-to-pixel image correspondences is a fundamental task of computer vision. Often, the objective is to align image pairs from the same semantic category for manipulation or segmentation purposes. Despite achieving superior performance, existing deep learning alignment methods cannot cluster images; consequently, clustering and pairing images needed to be a separate laborious and expensive step. Given a dataset with diverse semantic categories, we propose a multi-task model, Jim-Net, that can directly learn to cluster and align images without any pixel-level or image-level annotations. We design a pair-matching alignment unsupervised training algorithm that selectively matches and aligns image pairs from the clustering branch. Our unsupervised Jim-Net achieves comparable accuracy with state-of-the-art supervised methods on benchmark 2D image alignment dataset PF-PASCAL. Specifically, we apply Jim-Net to cryo-electron tomography, a revolutionary 3D microscopy imaging technique of native subcellular structures. After extensive evaluation on seven datasets, we demonstrate that Jim-Net enables systematic discovery and recovery of representative macromolecular structures in situ, which is essential for revealing molecular mechanisms underlying cellular functions. To our knowledge, Jim-Net is the first end-to-end model that can simultaneously align and cluster images, which significantly improves the performance as compared to performing each task alone.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7c508daf-4c81-4e3c-9a25-8d204fbf88cbCited by top-tier papers2
- On the Robustness of Deep Clustering Models: Adversarial Attacks and DefensesAnshuman Chhabra, Ashwin Sekhari, Prasant MohapatraNeurIPS 2022 · 12 citations
- BOE-ViT: Boosting Orientation Estimation with Equivariance in Self-Supervised 3D Subtomogram AlignmentRunmin Jiang, Jackson Daggett, Shriya Pingulkar, Yizhou Zhao et al.CVPR 2025
Builds on8
- Invariant Information Clustering for Unsupervised Image Classification and SegmentationXu Ji, Andrea Vedaldi, João F. HenriquesICCV 2019 · 956 citations
- Dual-Resolution Correspondence NetworksXinghui Li, Kai Han, Shuda Li, Victor PrisacariuNeurIPS 2020 · 207 citations
- Hyperpixel Flow: Semantic Correspondence With Multi-Layer Neural FeaturesJuhong Min, Jongmin Lee, Jean Ponce, Minsu ChoICCV 2019 · 120 citations
- Dynamic Context Correspondence Network for Semantic AlignmentShuaiyi Huang, Qiuyue Wang, Songyang Zhang, Shipeng Yan et al.ICCV 2019 · 97 citations
- Gum-Net: Unsupervised Geometric Matching for Fast and Accurate 3D Subtomogram Image Alignment and AveragingXiangrui Zeng, Min XuCVPR 2020
Related papers
- CryoLVM: Self-supervised Learning from Cryo-EM Density Maps with Large Vision ModelsWeining Fu, Kai Shu, Kui Xu, Qiangfeng Cliff ZhangICLR 2026 · 14 citations
- Vox-UDA: Voxel-wise Unsupervised Domain Adaptation for Cryo-Electron Subtomogram Segmentation with Denoised Pseudo-LabelingHaoran Li, Xingjian Li, Jiahua Shi, Huaming Chen et al.AAAI 2025 · 4 citations
- Unsupervised Multi-Scale Segmentation of 3D Subcellular World with Stable Diffusion Foundation ModelMostofa Rafid Uddin, H. M. Shadman Tabib, Thanh-Huy Nguyen, Kashish Gandhi et al.CVPR 2026
- Self-Supervised Cryo-Electron Tomography Volumetric Image Restoration from Single Noisy Volume with Sparsity ConstraintZhidong Yang, Fa Zhang, Renmin HanICCV 2021 · 15 citations
- Weakly Supervised 3D Semantic Segmentation Using Cross-Image Consensus and Inter-Voxel Affinity RelationsXiaoyu Zhu, Jeffrey Chen, Xiangrui Zeng, Junwei Liang et al.ICCV 2021 · 20 citations
