Integrating Stickers into Multimodal Dialogue Summarization: A Novel Dataset and Approach for Enhancing Social Media Interaction
Yuanchen Shi, Fang Kong
Abstract
With the popularity of social media, growing number of online chats and comments are presented in the form of multimodal dialogues containing stickers. Automatically summarizing these dialogues can effectively reduce content overload and save reading time. However, existing datasets and works are either text dialogue summarization, or articles with real photos that respectively perform text summaries and key image extraction, and have not simultaneously considered the multimodal dialogue automatic summarization tasks with sticker images and online chat scenarios. To compensate for the lack of datasets and researches in this field, we propose a brand-new Multimodal Chat Dialogue Summarization Containing Stickers (MCDSCS) task and dataset. It consists of 5,527 Chinese multimodal chat dialogues and 14,356 different sticker images, with each dialogue interspersed with stickers in the text to reflect the real social media chat scenario. MCDSCS can also contribute to filling the gap in Chinese multimodal dialogue data. We use the most advanced GPT4 model and carefully design Chain-of-Thoughts (COT) supplemented with manual review to generate dialogues and extract summaries. We also propose a novel method that integrates the visual information of stickers with the text descriptions of emotions and intentions (TEI). Experiments show that our method can effectively improve the performance of various mainstream summary generation models, even better than some other multimodal models, ChatGPT, and Vision Large Language Models (VLMs). Our data and code are publicly available at https://github.com/FakerBoom/MCDSCS.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get b5af4224-b311-4dfa-b2c0-7ff81b23cdbeCited by top-tier papers3
- Impact of Stickers on Multimodal Sentiment and Intent in Social Media: A New Task, Dataset and BaselineYuanchen Shi, Fang Kong, Longyin ZhangACM MM 2025 · 3 citations
- PerSRV: Personalized Sticker Retrieval with Vision-Language ModelHeng Er Metilda Chee, Jiayin Wang, Zhiqiang Guo, Weizhi Ma et al.WWW 2025 · 3 citations
- Emotion and Intention Guided Multi-Modal Learning for Sticker Response SelectionYuxuan Hu, Jian Chen, Yuhao Wang, Zixuan Li et al.AAAI 2026
Related papers
- A New Formula for Sticker Retrieval: Reply with Stickers in Multi-Modal and Multi-Session ConversationBingbing Wang, Yiming Du, Bin Liang, Zhixin Bai et al.AAAI 2025 · 5 citations
- STICKERCONV: Generating Multimodal Empathetic Responses from ScratchYiqun Zhang, Fanheng Kong, Peidong Wang, Shuang Sun et al.ACL 2024
- mRedditSum: A Multimodal Abstractive Summarization Dataset of Reddit Threads with ImagesKeighley Overbay, Jaewoo Ahn, Fatemeh Pesaran Zadeh, Joonsuk Park et al.EMNLP 2023 · 5 citations
- SER30K: A Large-Scale Dataset for Sticker Emotion RecognitionShengzhe Liu, Xin Zhang, Jufeng YangACM MM 2022 · 14 citations
- Observe before Generate: Emotion-Cause aware Video Caption for Multimodal Emotion Cause Generation in ConversationsFanfan Wang, Heqing Ma, Xiangqing Shen, Jianfei Yu et al.ACM MM 2024 · 6 citations
