Representation Potentials of Foundation Models for Multimodal Alignment: A Survey
Jianglin Lu, Hailing Wang, Yi Xu, Yizhou Wang, Kuo Yang, Yun Fu
Abstract
Foundation models learn highly transferable representations through large-scale pretraining on diverse data. An increasing body of research indicates that these representations exhibit a remarkable degree of similarity across architectures and modalities. In this survey, we investigate the representation potentials of foundation models, defined as the latent capacity of their learned representations to capture task-specific information within a single modality while also providing a transferable basis for alignment and unification across modalities. We begin by reviewing representative foundation models and the key metrics that make alignment measurable. We then synthesize empirical evidence of representation potentials from studies in vision, language, speech, multimodality, and neuroscience. The evidence suggests that foundation models often exhibit structural regularities and semantic consistencies in their representation spaces, positioning them as strong candidates for cross-modal transfer and alignment. We further analyze the key factors that foster representation potentials, discuss open questions, and highlight potential challenges.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cde45d55-fd3e-4398-a41a-30d18d36915cCited by top-tier papers3
- Ref-Adv: Exploring MLLM Visual Reasoning in Referring Expression TasksQihua Dong, Kuo Yang, Lin Ju, Handong Zhao et al.ICLR 2026 · 13 citations
- The Indra Representation Hypothesis for Multimodal AlignmentJianglin Lu, Hailing Wang, Kuo Yang, Yitian Zhang et al.NeurIPS 2025 · 8 citations
- Seeing Through Words: Controlling Visual Retrieval Quality with Language ModelsJianglin Lu, Simon Jenni, Kushal Kafle, Jing Shi et al.ICLR 2026 · 3 citations
Builds on47
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- Understanding Transfer Learning of RNA Foundation Models on Downstream TasksYuan Li, Heng Yang, Renzhi Chen, Ke LiICML 2026
- Foundation Models Defining A New Era In Sensor-based Human Activity Recognition: A Survey And OutlookSizhen Bian, Mengxi Liu, Lala Shakti Swarup Ray, Bo Zhou et al.UbiComp 2026 · 2 citations
- Brain encoding models based on multimodal transformers can transfer across language and visionJerry Tang, Meng Du, Vy A. Vo, Vasudev Lal et al.NeurIPS 2023 · 76 citations
- Objective drives the consistency of representational similarity across datasetsLaure Ciernik, Lorenz Linhardt, Marco Morik, Jonas Dippel et al.ICML 2025
- Probing the 3D Awareness of Visual Foundation ModelsMohamed El Banani, Amit Raj, Kevis-Kokitsi Maninis, Abhishek Kar et al.CVPR 2024
