Image-Chat: Engaging Grounded Conversations
Kurt Shuster, Samuel Humeau, Antoine Bordes, Jason Weston
Abstract
To achieve the long-term goal of machines being able to engage humans in conversation, our models should captivate the interest of their speaking partners. Communication grounded in images, whereby a dialogue is conducted based on a given photo, is a setup naturally appealing to humans (Hu et al., 2014) . In this work we study large-scale architectures and datasets for this goal. We test a set of neural architectures using state-of-the-art image and text representations, considering various ways to fuse the components. To test such models, we collect a dataset of grounded human-human conversations, where speakers are asked to play roles given a provided emotional mood or style, as the use of such traits is also a key factor in engagingness (Guo et al., 2019) . Our dataset, Image-Chat, consists of 202k dialogues over 202k images using 215 possible style traits. Automatic metrics and human evaluations of engagingness show the efficacy of our approach; in particular, we obtain state-of-the-art performance on the existing IGC task, and our best performing model is almost on par with humans on the Image-Chat test set (preferred 47.7% of the time).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ab79c79e-4657-47b4-a65f-b76a06d06d72Cited by top-tier papers25
- Getting Meta: A Multimodal Approach for Detecting Unsafe Conversations within Instagram Direct Messages of YouthShiza Ali, Afsaneh Razi, Seunghyun Kim, Ashwaq Alsoubai et al.CSCW 2023 · 32 citations
- Open Domain Dialogue Generation with Latent ImagesZe Yang, Wei Wu, Huang Hu, Can Xu et al.AAAI 2021 · 30 citations
- Evaluating Very Long-Term Conversational Memory of LLM AgentsAdyasha Maharana, Dong-Ho Lee, Sergey Tulyakov, Mohit Bansal et al.ACL 2024 · 30 citations
- Champagne: Learning Real-world Conversation from Large-Scale Web VideosSeungju Han, Jack Hessel, Nouha Dziri, Yejin Choi et al.ICCV 2023 · 22 citations
- MMDialog: A Large-scale Multi-turn Dialogue Dataset Towards Multi-modal Open-domain ConversationJiazhan Feng, Qingfeng Sun, Can Xu, Pu Zhao et al.ACL 2023 · 20 citations
Related papers
- PhotoChat: A Human-Human Dialogue Dataset With Photo Sharing Behavior For Joint Image-Text ModelingXiaoxue Zang, Lijuan Liu, Maria Wang, Yang Song et al.ACL 2021
- Chatting Makes Perfect: Chat-based Image RetrievalMatan Levy, Rami Ben-Ari, Nir Darshan, Dani LischinskiNeurIPS 2023 · 41 citations
- BIG5-CHAT: Shaping LLM Personalities Through Training on Human-Grounded DataWenkai Li, Jiarui Liu, Andy Liu, Xuhui Zhou et al.ACL 2025
- Making Visual Dialogue More Engaging: A New Task, Method, and MetricGuanghui Ye, Huan Zhao, Yingxue Gao, Zhixue Zhao et al.AAAI 2026
- V-VAE: A Variational Auto Encoding Framework Towards Fine-Grained Control over Human-Like ChatQi Lin, Weikai Xu, Lisi Chen, Bin DaiEMNLP 2025
