Adaptive and Multi-scale Affinity Alignment for Hierarchical Contrastive Learning
Jiawei Huang, Minming Li, Hu Ding
Abstract
Contrastive self-supervised learning has emerged as a powerful paradigm for extracting meaningful representations without labels. While effective at capturing broad categorical distinctions, current methods often struggle to preserve the fine-grained and hierarchical relationships inherent in real-world data. From the perspective of semantic alignment, conventional contrastive learning aligns representations to semantic structure at a global level, treating the entire embedding space uniformly and frequently overlooking rich local structural information. In this paper, we propose Adaptive Multi-scale Affinity alignment (AMA-alignment) , a framework that introduces localized contrastive objectives and a dynamic multi-scale optimization strategy to adaptively identify and refine poorly aligned regions within the embedding space. Although our model is inherently more complex due to its multi-scale and adaptive design, we provide the theoretical guarantees indicating that its convergence rate remains comparable to that of standard smooth non-convex optimization. We conduct a set of experiments on diverse benchmarks to show that AMA-alignment can effectively preserve hierarchical structure; more-over, AMA-alignment also outperforms existing contrastive methods on a range of downstream tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d4c7fa31-df27-4fa8-be9f-ad4f7a4eaad7Builds on38
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
Related papers
- Use All The Labels: A Hierarchical Multi-Label Contrastive Learning FrameworkShu Zhang, Ran Xu, Caiming Xiong, Chetan RamaiahCVPR 2022 · 71 citations
- MaMiCo: Macro-to-Micro Semantic Correspondence for Self-supervised Video Representation LearningBo Fang, Wenhao Wu, Chang Liu, Yu Zhou et al.ACM MM 2022 · 6 citations
- Anatomical Invariance Modeling and Semantic Alignment for Self-supervised Learning in 3D Medical Image AnalysisYankai Jiang, Mingze Sun, Heng Guo, Xiaoyu Bai et al.ICCV 2023 · 38 citations
- Multi-Label Learning with Contrastive Cluster Self-Supervision for 3D Hierarchical Semantic SegmentationShuyu Cao, Chongshou Li, Jie Xu, Tianrui Li et al.ICML 2026
- Multi-Scale Similarity Aggregation for Dynamic Metric LearningDingyi Zhang, Yingming Li, Zhongfei ZhangACM MM 2023 · 2 citations
