Image-Chat: Engaging Grounded Conversations
Kurt Shuster, Samuel Humeau, Antoine Bordes, Jason Weston
摘要
To achieve the long-term goal of machines being able to engage humans in conversation, our models should captivate the interest of their speaking partners. Communication grounded in images, whereby a dialogue is conducted based on a given photo, is a setup naturally appealing to humans (Hu et al., 2014) . In this work we study large-scale architectures and datasets for this goal. We test a set of neural architectures using state-of-the-art image and text representations, considering various ways to fuse the components. To test such models, we collect a dataset of grounded human-human conversations, where speakers are asked to play roles given a provided emotional mood or style, as the use of such traits is also a key factor in engagingness (Guo et al., 2019) . Our dataset, Image-Chat, consists of 202k dialogues over 202k images using 215 possible style traits. Automatic metrics and human evaluations of engagingness show the efficacy of our approach; in particular, we obtain state-of-the-art performance on the existing IGC task, and our best performing model is almost on par with humans on the Image-Chat test set (preferred 47.7% of the time).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper25
- Getting Meta: A Multimodal Approach for Detecting Unsafe Conversations within Instagram Direct Messages of YouthShiza Ali, Afsaneh Razi, Seunghyun Kim, Ashwaq Alsoubai 等CSCW 2023 · 被引用 32 次
- Open Domain Dialogue Generation with Latent ImagesZe Yang, Wei Wu, Huang Hu, Can Xu 等AAAI 2021 · 被引用 30 次
- Evaluating Very Long-Term Conversational Memory of LLM AgentsAdyasha Maharana, Dong-Ho Lee, Sergey Tulyakov, Mohit Bansal 等ACL 2024 · 被引用 30 次
- Champagne: Learning Real-world Conversation from Large-Scale Web VideosSeungju Han, Jack Hessel, Nouha Dziri, Yejin Choi 等ICCV 2023 · 被引用 22 次
- MMDialog: A Large-scale Multi-turn Dialogue Dataset Towards Multi-modal Open-domain ConversationJiazhan Feng, Qingfeng Sun, Can Xu, Pu Zhao 等ACL 2023 · 被引用 20 次
相关 Paper
- PhotoChat: A Human-Human Dialogue Dataset With Photo Sharing Behavior For Joint Image-Text ModelingXiaoxue Zang, Lijuan Liu, Maria Wang, Yang Song 等ACL 2021
- Chatting Makes Perfect: Chat-based Image RetrievalMatan Levy, Rami Ben-Ari, Nir Darshan, Dani LischinskiNeurIPS 2023 · 被引用 41 次
- BIG5-CHAT: Shaping LLM Personalities Through Training on Human-Grounded DataWenkai Li, Jiarui Liu, Andy Liu, Xuhui Zhou 等ACL 2025
- Making Visual Dialogue More Engaging: A New Task, Method, and MetricGuanghui Ye, Huan Zhao, Yingxue Gao, Zhixue Zhao 等AAAI 2026
- V-VAE: A Variational Auto Encoding Framework Towards Fine-Grained Control over Human-Like ChatQi Lin, Weikai Xu, Lisi Chen, Bin DaiEMNLP 2025
