Champagne: Learning Real-world Conversation from Large-Scale Web Videos
Seungju Han, Jack Hessel, Nouha Dziri, Yejin Choi, Youngjae Yu
Abstract
Visual information is central to conversation: body gestures and physical behaviour, for example, contribute to meaning that transcends words alone. To date, however, most neural conversational models are limited to just text. We introduce Champagne, a generative model of conversations that can account for visual contexts. To train Champagne, we collect and release YTD-18M, a large-scale corpus of 18M video-based dialogues. YTD-18M is constructed from web videos: crucial to our data collection pipeline is a pretrained language model that converts error-prone automatic transcripts to a cleaner dialogue format while maintaining meaning.Human evaluation reveals that YTD-18M is more sensible and specific than prior resources (MMDialog [17], 1M dialogues), while maintaining visual-groundedness. Experiments demonstrate that 1) Champagne learns to conduct conversation from YTD-18M; and 2) when fine-tuned, it achieves state-of-the-art results on four vision-language tasks focused on real-world conversations. We release data, models, and code at https://seungjuhan.me/champagne.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- TikTalk: A Video-Based Dialogue Dataset for Multi-Modal Chitchat in Real WorldHongpeng Lin, Ludan Ruan, Wenke Xia, Peiyu Liu et al.ACM MM 2023 · 10 citations
- From Natural Alignment to Conditional Controllability in Multimodal DialogueZeyu Jin, Songtao Zhou, Haoyu Wang, Minghao Tian et al.ICLR 2026 · 2 citations
- DIVE: Towards Descriptive and Diverse Visual Commonsense GenerationJun-Hyung Park, Hyuntae Park, Youjin Kang, Eojin Jeon et al.EMNLP 2023
- Reading Books is Great, But Not if You Are Driving! Visually Grounded Reasoning about Defeasible Commonsense NormsSeungju Han, Junhyeok Kim, Jack Hessel, Liwei Jiang et al.EMNLP 2023
- Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded DialoguesYoungmin Kim, Jiwan Chung, Jisoo Kim, Sunghyun Lee et al.ACL 2025
Builds on22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
Related papers
- Let's Go Real Talk: Spoken Dialogue Model for Face-to-Face ConversationSe Jin Park, Chae Won Kim, Hyeongseop Rha, Minsu Kim et al.ACL 2024
- Distilling Vision-Language Models on Millions of VideosYue Zhao, Long Zhao, Xingyi Zhou, Jialin Wu et al.CVPR 2024
- VideoGLaMM : A Large Multimodal Model for Pixel-Level Visual Grounding in VideosShehan Munasinghe, Hanan Gani, Wenqi Zhu, Jiale Cao et al.CVPR 2025
- Sequential Modeling Enables Scalable Learning for Large Vision ModelsYutong Bai, Xinyang Geng, Karttikeya Mangalam, Amir Bar et al.CVPR 2024
- Scaling Up Vision-Language Pretraining for Image CaptioningXiaowei Hu, Zhe Gan, Jianfeng Wang, Zhengyuan Yang et al.CVPR 2022 · 203 citations
