FiGuRO: Intrinsic Dimension Estimation for Multi-Modal Data
Viktoria Schuster, Sana Tonekaboni, Caroline Uhler
Abstract
Determining the complexity, or Intrinsic Dimension (ID), of data is fundamental to efficient and interpretable representation learning. This is particularly challenging in multi-modal settings when trying to learn disentangled representations for shared and private information. Existing techniques leave a critical gap: they are often static, uni-modal, or in the case of contrastive methods, adapt only to the shared ID implicitly. We introduce Fidelity-Guided Rank Optimization (FiGuRO), a framework for approximating the ID of uni- and multi-modal data under constraints of model capacity and hyperparameters. FiGuRO learns the dimensions of low-rank projections using truncated singular value decomposition and an algorithm that determines when to reduce or increase dimension and in which latent space. Disentanglement of shared and private information arises as an emergent property of this optimization, eliminating the need for complex auxiliary loss functions. We demonstrate that FiGuRO outperforms existing ID estimation techniques and is more robust to hyperparameter changes. Across simulations and real-world data, FiGuRO captures distinct ID scales and varying subspace ratios, and decomposes shared and private information successfully. Furthermore, we show that FiGuRO can be applied to modern uni-modal pretrained models, enabling efficient, post-hoc disentanglement of multi-modal representations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6e39bdbd-e279-4b5b-afba-ecd6e348edd4Builds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
- What Makes Multi-Modal Learning Better than Single (Provably)Yu Huang, Chenzhuang Du, Zihui Xue, Xuanyao Chen et al.NeurIPS 2021 · 404 citations
- The Intrinsic Dimension of Images and Its Impact on LearningPhillip Pope, Chen Zhu, Ahmed Abdelkader, Micah Goldblum et al.ICLR 2021 · 381 citations
Related papers
- MultiLoReFT: Decoupling Shared and Modality-Specific Subspaces in Multimodal Learning via Low-Rank Representation Fine-TuningSana Tonekaboni, Viktoria Schuster, Caroline UhlerICML 2026
- Partial Information Decomposition via Normalizing Flows in Latent Gaussian DistributionsWenyuan Zhao, Adithya Balachandran, Chao Tian, Paul Pu LiangNeurIPS 2025 · 5 citations
- LORE: Jointly Learning The Intrinsic Dimensionality and Relative Similarity Structure from Ordinal DataVivek Anand, Alec Helbling, Mark A. Davenport, Gordon J. Berman et al.ICLR 2026
- Disentanglement Analysis with Partial Information DecompositionSeiya Tokui, Issei SatoICLR 2022 · 16 citations
- Multi-View Causal Representation Learning with Partial ObservabilityDingling Yao, Danru Xu, Sébastien Lachapelle, Sara Magliacane et al.ICLR 2024 · 70 citations
