The Indra Representation Hypothesis for Multimodal Alignment
Jianglin Lu, Hailing Wang, Kuo Yang, Yitian Zhang, Simon Jenni, Yun Fu
Abstract
Recent studies have uncovered an interesting phenomenon: unimodal foundation models tend to learn convergent representations, regardless of differences in architecture, training objectives, or data modalities. However, these representations are essentially internal abstractions of samples that characterize samples independently, leading to limited expressiveness. In this paper, we propose The Indra Representation Hypothesis, inspired by the philosophical metaphor of Indra's Net. We argue that representations from unimodal foundation models are converging to implicitly reflect a shared relational structure underlying reality, akin to the relational ontology of Indra's Net. We formalize this hypothesis using the V-enriched Yoneda embedding from category theory, defining the Indra representation as a relational profile of each sample with respect to others. This formulation is shown to be unique, complete, and structure-preserving under a given cost function. We instantiate the Indra representation using angular distance and evaluate it in cross-model and cross-modal scenarios involving vision, language, and audio. Extensive experiments demonstrate that Indra representations consistently enhance robustness and alignment across architectures and modalities, providing a theoretically grounded and practical framework for training-free alignment of unimodal foundation models. Our code is available at https://github.com/Jianglin954/Indra .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dc9e6f7b-54e5-440c-98d6-90a733063872Cited by top-tier papers1
Ask how each one uses itBuilds on30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
Related papers
- Representation Potentials of Foundation Models for Multimodal Alignment: A SurveyJianglin Lu, Hailing Wang, Yi Xu, Yizhou Wang et al.EMNLP 2025
- It's a (Blind) Match! Towards Vision-Language Correspondence without Parallel DataDominik Schnaus, Nikita Araslanov, Daniel CremersCVPR 2025
- SOTAlign: Semi-Supervised Alignment of Unimodal Vision and Language Models via Optimal TransportSimon Roschmann, Paul KRZAKALA, Sonia Mazelet, Quentin Bouniot et al.ICML 2026 · 1 citation
- Unifying Vision-Language Representation Space with Single-Tower TransformerJiho Jang, Chaerin Kong, Donghyeon Jeon, Seonhoon Kim et al.AAAI 2023 · 34 citations
- Geometry-Preserving Unsupervised Alignment for Heterogeneous Foundation ModelsShuwen Yu, Zhanxuan Hu, Yi Zhao, Yonghang Tai et al.ICML 2026 · 1 citation
