ReSee: Responding through Seeing Fine-grained Visual Knowledge in Open-domain Dialogue
Haoqin Tu, Yitong Li, Fei Mi, Zhongliang Yang
Abstract
Incorporating visual knowledge into text-only dialogue systems has become a potential direction to imitate the way humans think, imagine, and communicate. However, existing multimodal dialogue systems are either confined by the scale and quality of available datasets or the coarse concept of visual knowledge. To address these issues, we provide a new paradigm of constructing multimodal dialogues as well as two datasets extended from text-only dialogues under such paradigm (RESEE-WoW, RESEE-DD). We propose to explicitly split the visual knowledge into finer granularity ("turn-level" and "entity-level"). To further boost the accuracy and diversity of augmented visual information, we retrieve them from the Internet or a large image dataset. To demonstrate the superiority and universality of the provided visual knowledge, we propose a simple but effective framework RESEE to add visual representation into vanilla dialogue models by modality concatenations. We also conduct extensive experiments and ablations w.r.t. different model configurations and visual knowledge settings. Empirically, encouraging results not only demonstrate the effectiveness of introducing visual knowledge at both entity and turn level but also verify the proposed model RESEE outperforms several state-of-the-art methods on automatic and human evaluations. By leveraging text and vision knowledge, RESEE can produce informative responses with real-world visual concepts. Our code is available at https: //github.com/ImKeTT/ReSee .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- SimVLM: Simple Visual Language Model Pretraining with Weak SupervisionZirui Wang, Jiahui Yu, Adams Wei Yu, Zihang Dai et al.ICLR 2022 · 950 citations
- nocaps: novel object captioning at scaleHarsh Agrawal, Peter Anderson, Karan Desai, Yufei Wang et al.ICCV 2019 · 631 citations
Related papers
- DualVD: An Adaptive Dual Encoding Model for Deep Visual Understanding in Visual DialogueXiaoze Jiang, Jing Yu, Zengchang Qin, Yingying Zhuang et al.AAAI 2020 · 72 citations
- A Framework for Vision-Language Warm-up Tasks in Multimodal Dialogue ModelsJaewook Lee, Seongsik Park, Seong-Heum Park, Hongjin Kim et al.EMNLP 2023 · 1 citation
- Structure-Aware Multimodal Sequential Learning for Visual DialogYoung-Jin Kim, Min-Jun Kim, Kyunghwan An, Jinwoo Ahn et al.AAAI 2024 · 3 citations
- TikTalk: A Video-Based Dialogue Dataset for Multi-Modal Chitchat in Real WorldHongpeng Lin, Ludan Ruan, Wenke Xia, Peiyu Liu et al.ACM MM 2023 · 10 citations
- Multimodal Dialogue Systems via Capturing Context-aware Dependencies of Semantic ElementsWeidong He, Zhi Li, Dongcai Lu, Enhong Chen et al.ACM MM 2020 · 31 citations
