Can LLMs Reason Over Non-Text Modalities in a Training-Free Manner? A Case Study with In-Context Representation Learning
Tianle Zhang, Wanlong Fang, Jonathan Woo, Paridhi Latawa, Deepak A. Subramanian, Alvin Chan
摘要
The remarkable performance of Large Language Models (LLMs) can be enhanced with test-time computation, which relies on external tools and even other deep learning models. However, existing approaches for integrating non-text modality representations into LLMs typically require additional costly supervised training, restricting on-the-fly adaptation to new domains and modalities. In this work, we explore the feasibility of integrating representations from non-text foundational models (FMs) into text-based LLMs in a training-free manner. We propose In-Context Representation Learning (ICRL) as a proof-of-concept to allow LLMs to adaptively utilize non-text modality representations with few-shot learning. Unlike traditional in-context learning, which incorporates text-label pairs, ICRL replaces text inputs with FM representations, enabling the LLM to perform multimodal inference without fine-tuning. We evaluate ICRL on a suite of tasks in the molecular domain, investigating three core research questions: (i) how to map FM representations into LLMs in a training-free manner, (ii) what factors influence ICRL performance, and (iii) what mechanisms underlie the effectiveness of ICRL. To the best of our knowledge, ICRL is the first training-free framework for integrating non-text modality representations into text-based LLMs, presenting a promising direction for adaptable, multi-modal generalization. 3
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Towards Understanding Modality Interaction in Multimodal Language Models via Partial Information DecompositionWanlong Fang, Tianle Zhang, Wen Tao, Alvin ChanICML 2026 · 被引用 17 次
- Rethinking Video-Language Model from the Language Input PerspectiveXiang Fang, Wanlong Fang, Changshuo Wang, Xiaoye Qu 等AAAI 2026 · 被引用 2 次
- Towards Unified Vision-Language Models with Incomplete Multi-Modal InputsXiang Fang, Wanlong Fang, Changshuo Wang, Keke Tang 等AAAI 2026 · 被引用 1 次
它引用的顶会 Paper24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
相关 Paper
- Why is a Bird's Caption a Good Demonstration? Towards Effective Multimodal In-Context Learning without Dedicated DataJunlin Fang, Wenya Wang, Lingli Zhang, Fengmao LvACM MM 2025 · 被引用 3 次
- VL-ICL Bench: The Devil in the Details of Multimodal In-Context LearningYongshuo Zong, Ondrej Bohdal, Timothy M. HospedalesICLR 2025
- Vector-ICL: In-context Learning with Continuous Vector RepresentationsYufan Zhuang, Chandan Singh, Liyuan Liu, Jingbo Shang 等ICLR 2025
- Why Multimodal In-Context Learning Lags Behind? Unveiling the Inner Mechanisms and BottlenecksYu Wang, Sharon LiACL 2026
- MMRL: Multi-Modal Representation Learning for Vision-Language ModelsYuncheng Guo, Xiaodong GuCVPR 2025
