A Multi-view Meta-learning Approach for Multi-modal Response Generation
Zhiliang Tian, Zheng Xie, Fuqiang Lin, Yiping Song
Abstract
As massive conversation examples are easily accessible on the Internet, we are now able to organize large-scale conversation corpora to build chatbots in a data-driven manner. Multi-modal social chatbots produce conversational utterances according to both textual utterances and vision signals. Due to the difficulty of bridging different modalities, the dialogue generation model of chatbots falls into local minima that only capture the mapping between textual input and textual output, as a result, it almost ignores the non-textual signals. Further, similar to the dialogue model with plain text as input and output, the generated responses from multi-modal dialogue also lack diversity and informativeness. In this paper, to address the above issues, we propose a Multi-View Meta-Learning (MultiVML) algorithm that groups samples in multiple views and customizes generation models to different groups. We employ a multi-view clustering to group the training samples so as to attend more to the unique information in non-textual modality. Tailoring different sets of model parameters for each group boosts the genereation diversity via meta-learning. We evaluate MultiVML on two variants of the OpenViDial benchmark datasets. The experiments show that our model not only better explore the information from multiple modalities, but also excels baselines in both quality and diversity.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers3
- Enhancing Fairness in Meta-learned User Modeling via Adaptive SamplingZheng Zhang, Qi Liu, Zirui Hu, Yi Zhan et al.WWW 2024 · 14 citations
- Battling against Tough Resister: Strategy Planning with Adversarial Game for Non-collaborative DialoguesHaiyang Wang, Zhiliang Tian, Yuchen Pan, Xin Song et al.ACL 2025 · 3 citations
- Retrieval-Augmented Perception: High-resolution Image Perception Meets Visual RAGWenbin Wang, Yongcheng Jing, Liang Ding, Yingjie Wang et al.ICML 2025
Related papers
- A Multi-View Clustering Algorithm for Short TextMinkuan Lu, Jianhua Yin, Kaijun Wang, Liqiang NieICDE 2024 · 3 citations
- Learning to Customize Model Structures for Few-shot Dialogue Generation TasksYiping Song, Zequn Liu, Wei Bi, Rui Yan et al.ACL 2020 · 33 citations
- Let's Go Real Talk: Spoken Dialogue Model for Face-to-Face ConversationSe Jin Park, Chae Won Kim, Hyeongseop Rha, Minsu Kim et al.ACL 2024
- Unsupervised Multimodal Clustering for Semantics Discovery in Multimodal UtterancesHanlei Zhang, Hua Xu, Fei Long, Xin Wang et al.ACL 2024
- Alleviating Observation Bias via Causal-Invariant Meta-Learning for Unbalanced Incomplete Multi-view ClusteringJiaqi Jin, Siwei Wang, Taichun Zhou, Dong Zhibin et al.ICML 2026
