Integrating Stickers into Multimodal Dialogue Summarization: A Novel Dataset and Approach for Enhancing Social Media Interaction
Yuanchen Shi, Fang Kong
摘要
With the popularity of social media, growing number of online chats and comments are presented in the form of multimodal dialogues containing stickers. Automatically summarizing these dialogues can effectively reduce content overload and save reading time. However, existing datasets and works are either text dialogue summarization, or articles with real photos that respectively perform text summaries and key image extraction, and have not simultaneously considered the multimodal dialogue automatic summarization tasks with sticker images and online chat scenarios. To compensate for the lack of datasets and researches in this field, we propose a brand-new Multimodal Chat Dialogue Summarization Containing Stickers (MCDSCS) task and dataset. It consists of 5,527 Chinese multimodal chat dialogues and 14,356 different sticker images, with each dialogue interspersed with stickers in the text to reflect the real social media chat scenario. MCDSCS can also contribute to filling the gap in Chinese multimodal dialogue data. We use the most advanced GPT4 model and carefully design Chain-of-Thoughts (COT) supplemented with manual review to generate dialogues and extract summaries. We also propose a novel method that integrates the visual information of stickers with the text descriptions of emotions and intentions (TEI). Experiments show that our method can effectively improve the performance of various mainstream summary generation models, even better than some other multimodal models, ChatGPT, and Vision Large Language Models (VLMs). Our data and code are publicly available at https://github.com/FakerBoom/MCDSCS.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper3
- Impact of Stickers on Multimodal Sentiment and Intent in Social Media: A New Task, Dataset and BaselineYuanchen Shi, Fang Kong, Longyin ZhangACM MM 2025 · 被引用 3 次
- PerSRV: Personalized Sticker Retrieval with Vision-Language ModelHeng Er Metilda Chee, Jiayin Wang, Zhiqiang Guo, Weizhi Ma 等WWW 2025 · 被引用 3 次
- Emotion and Intention Guided Multi-Modal Learning for Sticker Response SelectionYuxuan Hu, Jian Chen, Yuhao Wang, Zixuan Li 等AAAI 2026
相关 Paper
- A New Formula for Sticker Retrieval: Reply with Stickers in Multi-Modal and Multi-Session ConversationBingbing Wang, Yiming Du, Bin Liang, Zhixin Bai 等AAAI 2025 · 被引用 5 次
- STICKERCONV: Generating Multimodal Empathetic Responses from ScratchYiqun Zhang, Fanheng Kong, Peidong Wang, Shuang Sun 等ACL 2024
- mRedditSum: A Multimodal Abstractive Summarization Dataset of Reddit Threads with ImagesKeighley Overbay, Jaewoo Ahn, Fatemeh Pesaran Zadeh, Joonsuk Park 等EMNLP 2023 · 被引用 5 次
- SER30K: A Large-Scale Dataset for Sticker Emotion RecognitionShengzhe Liu, Xin Zhang, Jufeng YangACM MM 2022 · 被引用 14 次
- Observe before Generate: Emotion-Cause aware Video Caption for Multimodal Emotion Cause Generation in ConversationsFanfan Wang, Heqing Ma, Xiangqing Shen, Jianfei Yu 等ACM MM 2024 · 被引用 6 次
