Vision-Speech Models: Teaching Speech Models to Converse about Images
Amélie Royer, Moritz Böhle, Laurent Mazaré, Neil Zeghidour, Alexandre Défossez, Patrick Pérez
Abstract
The recent successes of Vision-Language models raise the question of how to equivalently imbue a pretrained speech model with visual understanding, an important milestone towards building a multimodal speech model able to freely converse about images. Building such a conversational Vision-Speech model brings its unique challenges: (i) paired image-speech datasets are much scarcer than their image-text counterparts, (ii) ensuring real-time latency at inference is crucial thus bringing compute and memory constraints, and (iii) the model should preserve prosodic features (e.g., speaker tone) which cannot be inferred from text alone. In this work, we introduce MoshiVis, augmenting a recent dialogue speech LLM, Moshi, with visual inputs through lightweight adaptation modules. An additional dynamic gating mechanism enables the model to more easily switch between the visual inputs and unrelated conversation topics. To reduce training costs, we design a simple one-stage, parameter-efficient fine-tuning pipeline in which we leverage a mixture of image-text (i.e., "speechless") and image-speech samples. We evaluate the model on downstream visual understanding tasks with both audio and text prompts, and report qualitative samples of interactions with MoshiVis. Our inference code, the image-speech data used for audio evaluation, as well as additional information are available at github.com/kyutai-labs/moshivis.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f6809abe-65e1-4dab-86ec-888b3e0b940fCited by top-tier papers1
Ask how each one uses itBuilds on8
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- What matters when building vision-language models?Hugo Laurençon, Léo Tronchon, Matthieu Cord, Victor SanhNeurIPS 2024 · 401 citations
- eP-ALM: Efficient Perceptual Augmentation of Language ModelsMustafa Shukor, Corentin Dancette, Matthieu CordICCV 2023 · 36 citations
- LLaMA-Omni: Seamless Speech Interaction with Large Language ModelsQingkai Fang, Shoutao Guo, Yan Zhou, Zhengrui Ma et al.ICLR 2025 · 2 citations
Related papers
- VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech InteractionChaoyou Fu, Haojia Lin, Xiong Wang, Yifan Zhang et al.NeurIPS 2025 · 234 citations
- MoRA: Missing Modality Low-Rank Adaptation for Visual RecognitionShu Zhao, Nilesh A. Ahuja, Tan Yu, Tianyi Shen et al.ICLR 2026 · 5 citations
- Modality-Specialized Synergizers for Interleaved Vision-Language GeneralistsZhiyang Xu, Minqian Liu, Ying Shen, Joy Rimchala et al.ICLR 2025
- Vision Transformers are Parameter-Efficient Audio-Visual LearnersYan-Bo Lin, Yi-Lin Sung, Jie Lei, Mohit Bansal et al.CVPR 2023
- Cheap and Quick: Efficient Vision-Language Instruction Tuning for Large Language ModelsGen Luo, Yiyi Zhou, Tianhe Ren, Shengxin Chen et al.NeurIPS 2023 · 157 citations
