With Limited Data for Multimodal Alignment, Let the STRUCTURE Guide You
Fabian Gröger, Shuo Wen, Huyen Le, Maria Brbic
摘要
Multimodal models have demonstrated powerful capabilities in complex tasks requiring multimodal alignment, including zero-shot classification and cross-modal retrieval. However, existing models typically rely on millions of paired multimodal samples, which are prohibitively expensive or infeasible to obtain in many domains. In this work, we explore the feasibility of building multimodal models with limited amount of paired data by aligning pretrained unimodal foundation models. We show that high-quality alignment is possible with as few as tens of thousands of paired samplesless than of the data typically used in the field. To achieve this, we introduce STRUCTURE, an effective regularization technique that preserves the neighborhood geometry of the latent space of unimodal encoders. Additionally, we show that aligning last layers is often suboptimal and demonstrate the benefits of aligning the layers with the highest representational similarity across modalities. These two components can be readily incorporated into existing alignment methods, yielding substantial gains across 24 zero-shot image classification and retrieval benchmarks, with average relative improvement of in classification and in retrieval tasks. Our results highlight the effectiveness and broad applicability of our framework for limited-sample multimodal learning and offer a promising path forward for resource-constrained domains.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Revisiting the Platonic Representation Hypothesis: An Aristotelian ViewFabian Gröger, Shuo Wen, Maria BrbicICML 2026 · 被引用 27 次
- SOTAlign: Semi-Supervised Alignment of Unimodal Vision and Language Models via Optimal TransportSimon Roschmann, Paul KRZAKALA, Sonia Mazelet, Quentin Bouniot 等ICML 2026 · 被引用 1 次
它引用的顶会 Paper18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras 等EMNLP 2021 · 被引用 937 次
相关 Paper
- Understanding and Constructing Latent Modality Structures in Multi-Modal Representation LearningQian Jiang, Changyou Chen, Han Zhao, Liqun Chen 等CVPR 2023
- EmergentBridge: Improving Zero-Shot Cross-Modal Transfer in Unified Multimodal Embedding ModelsJincheng Xie, Xingchen Xiao, Runheng Liu, Zhongyi Huang 等KDD 2026 · 被引用 1 次
- Better Integrating Vision and Semantics for Improving Few-shot ClassificationZhuoling Li, Yong WangACM MM 2023 · 被引用 4 次
- Geometry-Preserving Unsupervised Alignment for Heterogeneous Foundation ModelsShuwen Yu, Zhanxuan Hu, Yi Zhao, Yonghang Tai 等ICML 2026 · 被引用 1 次
- Harnessing Frozen Unimodal Encoders for Flexible Multimodal AlignmentMayug Maniparambil, Raiymbek Akshulakov, Yasser Abdelaziz Dahou Djilali, Sanath Narayan 等CVPR 2025
