Intrinsic Dimension Correlation: uncovering nonlinear connections in multimodal representations
Lorenzo Basile, Santiago Acevedo, Luca Bortolussi, Fabio Anselmi, Alex Rodriguez
Abstract
To gain insight into the mechanisms behind machine learning methods, it is crucial to establish connections among the features describing data points. However, these correlations often exhibit a high-dimensional and strongly nonlinear nature, which makes them challenging to detect using standard methods. This paper exploits the entanglement between intrinsic dimensionality and correlation to propose a metric that quantifies the (potentially nonlinear) correlation between high-dimensional manifolds. We first validate our method on synthetic data in controlled environments, showcasing its advantages and drawbacks compared to existing techniques. Subsequently, we extend our analysis to large-scale applications in neural network representations. Specifically, we focus on latent representations of multimodal data, uncovering clear correlations between paired visual and textual embeddings, whereas existing methods struggle significantly in detecting similarity. Our results indicate the presence of highly nonlinear correlation patterns between latent manifolds.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0957a25f-9b92-46c7-a706-084ff506390aCited by top-tier papers3
- Large Vision-Language Models Get Lost in AttentionGongli Xi, Ye Tian, Mengyu Yang, Huahui Yi et al.ICML 2026 · 4 citations
- UniFast-HGR: Scalable and Efficient Maximal Correlation for Multimodal ModelsHongkang Zhang, Shao-Lun Huang, Yanlong Wang, Ercan KURUOGLUICML 2026
- Textural or Textual: How Vision-Language Models Read Text in ImagesHanzhang Wang, Qingyuan MaICML 2025
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- Do Vision Transformers See Like Convolutional Neural Networks?Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang et al.NeurIPS 2021 · 1,553 citations
Related papers
- The Shape of Data: Intrinsic Distance for Data DistributionsAnton Tsitsulin, Marina Munkhoeva, Davide Mottin, Panagiotis Karras et al.ICLR 2020 · 57 citations
- Evaluating the Disentanglement of Deep Generative Models through Manifold TopologySharon Zhou, Eric Zelikman, Fred Lu, Andrew Y. Ng et al.ICLR 2021 · 29 citations
- Disentanglement Analysis with Partial Information DecompositionSeiya Tokui, Issei SatoICLR 2022 · 16 citations
- The Effect of Manifold Entanglement and Intrinsic Dimensionality on LearningDaniel Kienitz, Ekaterina Komendantskaya, Michael A. LonesAAAI 2022 · 6 citations
- Connecting Neural Models Latent Geometries with Relative Geodesic RepresentationsHanlin Yu, Berfin Inal, Georgios Arvanitidis, Søren Hauberg et al.NeurIPS 2025 · 5 citations
