Multi-Modal Open-Domain Dialogue
Kurt Shuster, Eric Michael Smith, Da Ju, Jason Weston
摘要
Recent work in open-domain conversational agents has demonstrated that significant improvements in humanness and user preference can be achieved via massive scaling in both pre-training data and model size (Adiwardana et al., 2020; Roller et al., 2020) . However, if we want to build agents with human-like abilities, we must expand beyond handling just text. A particularly important topic is the ability to see images and communicate about what is perceived. With the goal of getting humans to engage in multi-modal dialogue, we investigate combining components from state-of-the-art open-domain dialogue agents with those from state-of-the-art vision models. We study incorporating different image fusion schemes and domain-adaptive pre-training and fine-tuning strategies, and show that our best resulting model outperforms strong existing models in multi-modal dialogue while simultaneously performing as well as its predecessor (text-only) BlenderBot (Roller et al., 2020) in text-based conversation. We additionally investigate and incorporate safety components in our final model, and show that such efforts do not diminish model performance with respect to human preference.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- "I'm sorry to hear that": Finding New Biases in Language Models with a Holistic Descriptor DatasetEric Michael Smith, Melissa Hall, Melanie Kambadur, Eleonora Presani 等EMNLP 2022 · 被引用 56 次
- TikTalk: A Video-Based Dialogue Dataset for Multi-Modal Chitchat in Real WorldHongpeng Lin, Ludan Ruan, Wenke Xia, Peiyu Liu 等ACM MM 2023 · 被引用 10 次
- ReSee: Responding through Seeing Fine-grained Visual Knowledge in Open-domain DialogueHaoqin Tu, Yitong Li, Fei Mi, Zhongliang YangEMNLP 2023 · 被引用 4 次
- A Framework for Vision-Language Warm-up Tasks in Multimodal Dialogue ModelsJaewook Lee, Seongsik Park, Seong-Heum Park, Hongjin Kim 等EMNLP 2023 · 被引用 1 次
- Making Visual Dialogue More Engaging: A New Task, Method, and MetricGuanghui Ye, Huan Zhao, Yingxue Gao, Zhixue Zhao 等AAAI 2026
它引用的顶会 Paper9
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong 等AAAI 2020 · 被引用 966 次
- Intermediate-Task Transfer Learning with Pretrained Language Models: When and Why Does It Work?Yada Pruksachatkun, Jason Phang, Haokun Liu, Phu Mon Htut 等ACL 2020 · 被引用 168 次
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo 等ACL 2020 · 被引用 93 次
- Language (Technology) is Power: A Critical Survey of "Bias" in NLPSu Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. WallachACL 2020 · 被引用 68 次
- Can You Put it All Together: Evaluating Conversational Agents' Ability to Blend SkillsEric Michael Smith, Mary Williamson, Kurt Shuster, Jason Weston 等ACL 2020 · 被引用 18 次
相关 Paper
- Structure-Aware Multimodal Sequential Learning for Visual DialogYoung-Jin Kim, Min-Jun Kim, Kyunghwan An, Jinwoo Ahn 等AAAI 2024 · 被引用 3 次
- HaploVL: A Single-Transformer Baseline for Multi-Modal UnderstandingRui Yang, Lin Song, Yicheng Xiao, Runhui Huang 等ICML 2025
- Let's Go Real Talk: Spoken Dialogue Model for Face-to-Face ConversationSe Jin Park, Chae Won Kim, Hyeongseop Rha, Minsu Kim 等ACL 2024
- Maria: A Visual Experience Powered Conversational AgentZujie Liang, Huang Hu, Can Xu, Chongyang Tao 等ACL 2021
- VD-BERT: A Unified Vision and Dialog Transformer with BERTYue Wang, Shafiq R. Joty, Michael R. Lyu, Irwin King 等EMNLP 2020 · 被引用 68 次
