Multimodal Dialogue Response Generation
Qingfeng Sun, Yujing Wang, Can Xu, Kai Zheng, Yaming Yang, Huang Hu, Fei Xu, Jessica Zhang, Xiubo Geng, Daxin Jiang
摘要
Responsing with image has been recognized as an important capability for an intelligent conversational agent. Yet existing works only focus on exploring the multimodal dialogue models which depend on retrieval-based methods, but neglecting generation methods. To fill in the gaps, we first present a new task: multimodal dialogue response generation (MDRG) - given the dialogue history, one model needs to generate a text sequence or an image as response. Learning such a MDRG model often requires multimodal dialogues containing both texts and images which are difficult to obtain. Motivated by the challenge in practice, we consider MDRG under a natural assumption that only limited training examples are available. In such a low-resource setting, we devise a novel conversational agent, Divter, in order to isolate parameters that depend on multimodal dialogues from the entire generation model. By this means, the major part of the model can be learned from a large number of text-only dialogues and text-image pairs respectively, then the whole parameters can be well fitted using the limited training examples. Extensive experiments demonstrate our method achieves state-of-the-art results in both automatic and human evaluation, and can generate informative text and high-resolution image responses.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- DreamLLM: Synergistic Multimodal Comprehension and CreationRunpei Dong, Chunrui Han, Yuang Peng, Zekun Qi 等ICLR 2024 · 被引用 315 次
- Champagne: Learning Real-world Conversation from Large-Scale Web VideosSeungju Han, Jack Hessel, Nouha Dziri, Yejin Choi 等ICCV 2023 · 被引用 22 次
- MMDialog: A Large-scale Multi-turn Dialogue Dataset Towards Multi-modal Open-domain ConversationJiazhan Feng, Qingfeng Sun, Can Xu, Pu Zhao 等ACL 2023 · 被引用 20 次
- The Zeno's Paradox of 'Low-Resource' LanguagesHellina Hailu Nigatu, Atnafu Lambebo Tonja, Benjamin Rosman, Thamar Solorio 等EMNLP 2024 · 被引用 10 次
- TikTalk: A Video-Based Dialogue Dataset for Multi-Modal Chitchat in Real WorldHongpeng Lin, Ludan Ruan, Wenke Xia, Peiyu Liu 等ACM MM 2023 · 被引用 10 次
它引用的顶会 Paper15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- CogView: Mastering Text-to-Image Generation via TransformersMing Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng 等NeurIPS 2021 · 被引用 1,026 次
- Sequential Latent Knowledge Selection for Knowledge-Grounded DialogueByeongchang Kim, Jaewoo Ahn, Gunhee KimICLR 2020 · 被引用 179 次
- X-LXMERT: Paint, Caption and Answer Questions with Multi-Modal TransformersJaemin Cho, Jiasen Lu, Dustin Schwenk, Hannaneh Hajishirzi 等EMNLP 2020 · 被引用 80 次
相关 Paper
- Open Domain Dialogue Generation with Latent ImagesZe Yang, Wei Wu, Huang Hu, Can Xu 等AAAI 2021 · 被引用 30 次
- Low-Resource Knowledge-Grounded Dialogue GenerationXueliang Zhao, Wei Wu, Chongyang Tao, Can Xu 等ICLR 2020 · 被引用 115 次
- Reflecting on Experiences for Response GenerationChenchen Ye, Lizi Liao, Suyu Liu, Tat-Seng ChuaACM MM 2022 · 被引用 12 次
- Structure-Aware Multimodal Sequential Learning for Visual DialogYoung-Jin Kim, Min-Jun Kim, Kyunghwan An, Jinwoo Ahn 等AAAI 2024 · 被引用 3 次
- ZRIGF: An Innovative Multimodal Framework for Zero-Resource Image-Grounded Dialogue GenerationBo Zhang, Jian Wang, Hui Ma, Bo Xu 等ACM MM 2023 · 被引用 4 次
