Accept the Modality Gap: An Exploration in the Hyperbolic Space
Sameera Ramasinghe, Violetta Shevchenko, Gil Avraham, Thalaiyasingam Ajanthan
Abstract
Recent advancements in machine learning have spotlighted the potential of hyperbolic spaces as they effectively learn hierarchical feature representations. While there has been progress in leveraging hyperbolic spaces in single-modality contexts, its exploration in multimodal settings remains under explored. A recent work has sought to transpose Euclidean multimodal learning techniques to hyperbolic spaces, by adopting a geodesic distance based contrastive loss. However, we show both theoretically and empirically that such spatial proximity based contrastive loss significantly disrupts hierarchies in the latent space. To remedy this, we advocate that the cross-modal representations should accept the inherent modality gap between text and images, and introduce a novel approach to measure cross-modal similarity that does not enforce spatial proximity. Our approach shows remarkable capabilities in preserving unimodal hierarchies while aligning the two modalities. Our experiments on a series of downstream tasks demonstrate that a better latent structure emerges with our objective function while being superior in text-to-image and image-to-text retrieval tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 80741ca5-66c9-4ca5-a977-2e0d2f184da2Cited by top-tier papers27
- Hyperbolic Dataset DistillationWenyuan Li, Guang Li, Keisuke Maeda, Takahiro Ogawa et al.NeurIPS 2025 · 17 citations
- Closing the Modality Gap Aligns Group-Wise SemanticsEleonora Grassucci, Giordano Cicchetti, Emanuele Frasca, Aurelio Uncini et al.ICLR 2026 · 5 citations
- HyperET: Efficient Training in Hyperbolic Space for Multi-modal Large Language ModelsZelin Peng, Zhengqin Xu, Qingyang Liu, Xiaokang Yang et al.NeurIPS 2025 · 5 citations
- Learning Visual Hierarchies in Hyperbolic Space for Image RetrievalZiwei Wang, Sameera Ramasinghe, Chenchen Hu, Julien Monteil et al.ICCV 2025 · 4 citations
- Is the Modality Gap a Bug or a Feature? A Robustness PerspectiveRhea Chowers, Oshri Naparstek, Udi Barzelay, Yair WeissCVPR 2026 · 4 citations
Builds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
Related papers
- Hyperbolic Multimodal Continual LearningJiahong Liu, Ming Shen, Xiaohao Liu, ZHITAO YING et al.ICML 2026
- Hyperbolic Vision Transformers: Combining Improvements in Metric LearningAleksandr Ermolov, Leyla Mirvakhabova, Valentin Khrulkov, Nicu Sebe et al.CVPR 2022 · 97 citations
- Hyperbolic Contrastive Learning for Visual Representations beyond ObjectsSongwei Ge, Shlok Mishra, Simon Kornblith, Chun-Liang Li et al.CVPR 2023
- Hyperbolic Gramian Volumes for Multimodal AlignmentSaiyang Na, Feng Jiang, Qifeng Zhou, Wenliang Zhong et al.CVPR 2026
- Understanding and Constructing Latent Modality Structures in Multi-Modal Representation LearningQian Jiang, Changyou Chen, Han Zhao, Liqun Chen et al.CVPR 2023
