Cross-Modal Generalization: Learning in Low Resource Modalities via Meta-Alignment
Paul Pu Liang, Peter Wu, Liu Ziyin, Louis-Philippe Morency, Ruslan Salakhutdinov
Abstract
How can we generalize to a new prediction task at test time when it also uses a new modality as input? More importantly, how can we do this with as little annotated data as possible? This problem of cross-modal generalization is a new research milestone with concrete impact on real-world applications. For example, can an AI system start understanding spoken language from mostly written text? Or can it learn the visual steps of a new recipe from only text descriptions? In this work, we formalize cross-modal generalization as a learning paradigm to train a model that can (1) quickly perform new tasks (from new domains) while (2) being originally trained on a different input modality. Such a learning paradigm is crucial for generalization to low-resource modalities such as spoken speech in rare languages while utilizing a different high-resource modality such as text. One key technical challenge that makes it different from other learning paradigms such as meta-learning and domain adaptation is the presence of different source and target modalities which will require different encoders. We propose an effective solution based on meta-alignment, a novel method to align representation spaces using strongly and weakly paired cross-modal data while ensuring quick generalization to new tasks across different modalities. This approach uses key ideas from cross-modal learning and meta-learning, and presents strong results on the cross-modal generalization problem. We benchmark several approaches on 3 real-world classification tasks: few-shot recipe classification from text to images of recipes, object classification from images to audio of objects, and language classification from text to spoken speech across 100 languages spanning many rare languages. Our results demonstrate strong performance even when the new target modality has only a few (1-10) labeled samples and in the presence of noisy labels, a scenario particularly prevalent in low-resource modalities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5a923e70-c1a9-4d33-a979-5d55f1fe1002Cited by top-tier papers5
- I can't believe there's no images! : Learning Visual Tasks Using Only Language SupervisionSophia Gu, Christopher Clark, Aniruddha KembhaviICCV 2023 · 41 citations
- MultiViz: Towards Visualizing and Understanding Multimodal ModelsPaul Pu Liang, Yiwei Lyu, Gunjan Chhablani, Nihal Jain et al.ICLR 2023 · 15 citations
- Forecasting of 3D Whole-Body Human Poses with Grasping ObjectsHaitao Yan, Qiongjie Cui, Jiexin Xie, Shijie GuoCVPR 2024 · 6 citations
- Towards Out-of-Modal Generalization without Instance-level Modal CorrespondenceZhuo Huang, Gang Niu, Bo Han, Masashi Sugiyama et al.ICLR 2025
- Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?Simon Park, Abhishek Panigrahi, Yun Cheng, Dingli Yu et al.ICML 2025
Builds on2
Related papers
- Universal Algorithm-Implicit LearningStefano Woerner, Seong Joon Oh, Christian BaumgartnerICML 2026
- Meta Learning to Bridge Vision and Language Models for Multimodal Few-Shot LearningIvona Najdenkoska, Xiantong Zhen, Marcel WorringICLR 2023 · 8 citations
- Meta-Learning for Fast Cross-Lingual Adaptation in Dependency ParsingAnna Langedijk, Verna Dankers, Phillip Lippe, Sander Bos et al.ACL 2022 · 17 citations
- Multimodal Few-Shot Learning with Frozen Language ModelsMaria Tsimpoukelli, Jacob Menick, Serkan Cabi, S. M. Ali Eslami et al.NeurIPS 2021 · 1,020 citations
- Learning to Generalize across Domains on Single Test SamplesZehao Xiao, Xiantong Zhen, Ling Shao, Cees G. M. SnoekICLR 2022 · 40 citations
