Text is NOT Enough: Integrating Visual Impressions into Open-domain Dialogue Generation
Lei Shen, Haolan Zhan, Xin Shen, Yonghao Song, Xiaofang Zhao
摘要
Open-domain dialogue generation in natural language processing (NLP) is by default a pure-language task, which aims to satisfy human need for daily communication on open-ended topics by producing related and informative responses. In this paper, we point out that hidden images, named as visual impressions (VIs), can be explored from the text-only data to enhance dialogue understanding and help generate better responses. Besides, the semantic dependency between an dialogue post and its response is complicated, e.g., few word alignments and some topic transitions. Therefore, the visual impressions of them are not shared, and it is more reasonable to integrate the response visual impressions (RVIs) into the decoder, rather than the post visual impressions (PVIs). However, both the response and its RVIs are not given directly in the test process. To handle the above issues, we propose a framework to explicitly construct VIs based on pure-language dialogue datasets and utilize them for better dialogue understanding and generation. Specifically, we obtain a group of images (PVIs) for each post based on a pre-trained word-image mapping model. These PVIs are used in a co-attention encoder to get a post representation with both visual and textual information. Since the RVIs are not provided during testing, we design a cascade decoder that consists of two sub-decoders. The first sub-decoder predicts the content words in response, and applies the word-image mapping model to get corresponding RVIs. Then, the second sub-decoder generates the response based on the post and RVIs. Experimental results on two open-domain dialogue datasets show that our proposed approach achieves superior performance over competitive baselines in terms of fluency, relatedness, and diversity.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- TikTalk: A Video-Based Dialogue Dataset for Multi-Modal Chitchat in Real WorldHongpeng Lin, Ludan Ruan, Wenke Xia, Peiyu Liu 等ACM MM 2023 · 被引用 10 次
- Learning to Imagine: Visually-Augmented Natural Language GenerationTianyi Tang, Yushuo Chen, Yifan Du, Junyi Li 等ACL 2023 · 被引用 7 次
- ReSee: Responding through Seeing Fine-grained Visual Knowledge in Open-domain DialogueHaoqin Tu, Yitong Li, Fei Mi, Zhongliang YangEMNLP 2023 · 被引用 4 次
- ZRIGF: An Innovative Multimodal Framework for Zero-Resource Image-Grounded Dialogue GenerationBo Zhang, Jian Wang, Hui Ma, Bo Xu 等ACM MM 2023 · 被引用 4 次
- Improving Personalized Explanation Generation through VisualizationShijie Geng, Zuohui Fu, Yingqiang Ge, Lei Li 等ACL 2022
它引用的顶会 Paper13
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li 等ICLR 2020 · 被引用 1,825 次
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong 等AAAI 2020 · 被引用 966 次
- Neural Machine Translation with Universal Visual RepresentationZhuosheng Zhang, Kehai Chen, Rui Wang, Masao Utiyama 等ICLR 2020 · 被引用 117 次
- CDL: Curriculum Dual Learning for Emotion-Controllable Response GenerationLei Shen, Yang FengACL 2020 · 被引用 82 次
- DualVD: An Adaptive Dual Encoding Model for Deep Visual Understanding in Visual DialogueXiaoze Jiang, Jing Yu, Zengchang Qin, Yingying Zhuang 等AAAI 2020 · 被引用 72 次
相关 Paper
- Open Domain Dialogue Generation with Latent ImagesZe Yang, Wei Wu, Huang Hu, Can Xu 等AAAI 2021 · 被引用 30 次
- A Framework for Vision-Language Warm-up Tasks in Multimodal Dialogue ModelsJaewook Lee, Seongsik Park, Seong-Heum Park, Hongjin Kim 等EMNLP 2023 · 被引用 1 次
- DialogVED: A Pre-trained Latent Variable Encoder-Decoder Model for Dialog Response GenerationWei Chen, Yeyun Gong, Song Wang, Bolun Yao 等ACL 2022
- PLATO: Pre-trained Dialogue Generation Model with Discrete Latent VariableSiqi Bao, Huang He, Fan Wang, Hua Wu 等ACL 2020 · 被引用 229 次
- History for Visual Dialog: Do we really need it?Shubham Agarwal, Trung Bui, Joon-Young Lee, Ioannis Konstas 等ACL 2020 · 被引用 8 次
