Representation Potentials of Foundation Models for Multimodal Alignment: A Survey
Jianglin Lu, Hailing Wang, Yi Xu, Yizhou Wang, Kuo Yang, Yun Fu
摘要
Foundation models learn highly transferable representations through large-scale pretraining on diverse data. An increasing body of research indicates that these representations exhibit a remarkable degree of similarity across architectures and modalities. In this survey, we investigate the representation potentials of foundation models, defined as the latent capacity of their learned representations to capture task-specific information within a single modality while also providing a transferable basis for alignment and unification across modalities. We begin by reviewing representative foundation models and the key metrics that make alignment measurable. We then synthesize empirical evidence of representation potentials from studies in vision, language, speech, multimodality, and neuroscience. The evidence suggests that foundation models often exhibit structural regularities and semantic consistencies in their representation spaces, positioning them as strong candidates for cross-modal transfer and alignment. We further analyze the key factors that foster representation potentials, discuss open questions, and highlight potential challenges.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Ref-Adv: Exploring MLLM Visual Reasoning in Referring Expression TasksQihua Dong, Kuo Yang, Lin Ju, Handong Zhao 等ICLR 2026 · 被引用 13 次
- The Indra Representation Hypothesis for Multimodal AlignmentJianglin Lu, Hailing Wang, Kuo Yang, Yitian Zhang 等NeurIPS 2025 · 被引用 8 次
- Seeing Through Words: Controlling Visual Retrieval Quality with Language ModelsJianglin Lu, Simon Jenni, Kushal Kafle, Jing Shi 等ICLR 2026 · 被引用 3 次
它引用的顶会 Paper47
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
相关 Paper
- Understanding Transfer Learning of RNA Foundation Models on Downstream TasksYuan Li, Heng Yang, Renzhi Chen, Ke LiICML 2026
- Foundation Models Defining A New Era In Sensor-based Human Activity Recognition: A Survey And OutlookSizhen Bian, Mengxi Liu, Lala Shakti Swarup Ray, Bo Zhou 等UbiComp 2026 · 被引用 2 次
- Brain encoding models based on multimodal transformers can transfer across language and visionJerry Tang, Meng Du, Vy A. Vo, Vasudev Lal 等NeurIPS 2023 · 被引用 76 次
- Objective drives the consistency of representational similarity across datasetsLaure Ciernik, Lorenz Linhardt, Marco Morik, Jonas Dippel 等ICML 2025
- Probing the 3D Awareness of Visual Foundation ModelsMohamed El Banani, Amit Raj, Kevis-Kokitsi Maninis, Abhishek Kar 等CVPR 2024
