ZRIGF: An Innovative Multimodal Framework for Zero-Resource Image-Grounded Dialogue Generation
Bo Zhang, Jian Wang, Hui Ma, Bo Xu, Hongfei Lin
Abstract
Image-grounded dialogue systems benefit greatly from integrating visual information, resulting in high-quality response generation. However, current models struggle to effectively utilize such information in zero-resource scenarios, mainly due to the disparity between image and text modalities. To overcome this challenge, we propose an innovative multimodal framework, called ZRIGF, which assimilates image-grounded information for dialogue generation in zero-resource situations. ZRIGF implements a two-stage learning strategy, comprising contrastive pre-training and generative pretraining. Contrastive pre-training includes a text-image matching module that maps images and texts into a unified encoded vector space, along with a text-assisted masked image modeling module that preserves pre-training visual features and fosters further multimodal feature alignment. Generative pre-training employs a multimodal fusion module and an information transfer module to produce insightful responses based on harmonized multimodal representations. Comprehensive experiments conducted on both textbased and image-grounded dialogue datasets demonstrate ZRIGF's efficacy in generating contextually pertinent and informative responses. Furthermore, we adopt a fully zero-resource scenario in the image-grounded dialogue dataset to demonstrate our framework's robust generalization capabilities in novel domains. The code is available at https://github.com/zhangbo-nlp/ZRIGF.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7ba8d565-699a-4f45-a94e-dff86b7141feBuilds on11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Zero-Resource Knowledge-Grounded Dialogue GenerationLinxiao Li, Can Xu, Wei Wu, Yufan Zhao et al.NeurIPS 2020 · 75 citations
- MMConv: An Environment for Multimodal Conversational Search across Multiple DomainsLizi Liao, Le Hong Long, Zheng Zhang, Minlie Huang et al.SIGIR 2021 · 70 citations
- Multimodal Dialogue Response GenerationQingfeng Sun, Yujing Wang, Can Xu, Kai Zheng et al.ACL 2022 · 58 citations
Related papers
- Grounding Language Models to Images for Multimodal Inputs and OutputsJing Yu Koh, Ruslan Salakhutdinov, Daniel FriedICML 2023 · 160 citations
- Reflecting on Experiences for Response GenerationChenchen Ye, Lizi Liao, Suyu Liu, Tat-Seng ChuaACM MM 2022 · 12 citations
- Mario: Multimodal Graph Reasoning with Large Language ModelsYuanfu Sun, Kang Li, Pengkang Guo, Jiajin Liu et al.CVPR 2026 · 2 citations
- Open Domain Dialogue Generation with Latent ImagesZe Yang, Wei Wu, Huang Hu, Can Xu et al.AAAI 2021 · 30 citations
- HiVG: Hierarchical Multimodal Fine-grained Modulation for Visual GroundingLinhui Xiao, Xiaoshan Yang, Fang Peng, Yaowei Wang et al.ACM MM 2024 · 28 citations
