Learning to Respond with Stickers: A Framework of Unifying Multi-Modality in Multi-Turn Dialog
Shen Gao, Xiuying Chen, Chang Liu, Li Liu, Dongyan Zhao, Rui Yan
摘要
Stickers with vivid and engaging expressions are becoming increasingly popular in online messaging apps, and some works are dedicated to automatically select sticker response by matching text labels of stickers with previous utterances. However, due to their large quantities, it is impractical to require text labels for the all stickers. Hence, in this paper, we propose to recommend an appropriate sticker to user based on multi-turn dialog context history without any external labels. Two main challenges are confronted in this task. One is to learn semantic meaning of stickers without corresponding text labels. Another challenge is to jointly model the candidate sticker with the multi-turn dialog context. To tackle these challenges, we propose a sticker response selector (SRS) model. Specifically, SRS first employs a convolutional based sticker image encoder and a self-attention based multi-turn dialog encoder to obtain the representation of stickers and utterances. Next, deep interaction network is proposed to conduct deep matching between the sticker with each utterance in the dialog history. SRS then learns the short-term and long-term dependency between all interaction results by a fusion network to output the the final matching score. To evaluate our proposed method, we collect a large-scale realworld dialog dataset with stickers from one of the most popular online chatting platform. Extensive experiments conducted on this dataset show that our model achieves the state-of-the-art performance for all commonly-used metrics. Experiments also verify the * Equal contribution. Ordering is decided by a coin flip. Work performed during an internship at IIAI. † WICT is the abbreviation of Wangxuan Institute of Computer Technology.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- VMSMO: Learning to Generate Multimodal Summary for Video-based News ArticlesMingzhe Li, Xiuying Chen, Shen Gao, Zhangming Chan 等EMNLP 2020 · 被引用 65 次
- Reasoning in Dialog: Improving Response Generation by Context Reading ComprehensionXiuying Chen, Zhi Cui, Jiayi Zhang, Chen Wei 等AAAI 2021 · 被引用 15 次
- Context-to-Session Matching: Utilizing Whole Session for Response Selection in Information-Seeking Dialogue SystemsZhenxin Fu, Shaobo Cui, Mingyue Shang, Feng Ji 等KDD 2020 · 被引用 13 次
- A New Formula for Sticker Retrieval: Reply with Stickers in Multi-Modal and Multi-Session ConversationBingbing Wang, Yiming Du, Bin Liang, Zhixin Bai 等AAAI 2025 · 被引用 5 次
- PerSRV: Personalized Sticker Retrieval with Vision-Language ModelHeng Er Metilda Chee, Jiayin Wang, Zhiqiang Guo, Weizhi Ma 等WWW 2025 · 被引用 3 次
它引用的顶会 Paper1
相关 Paper
- Perceive before Respond: Improving Sticker Response Selection by Emotion Distillation and Hard MiningWuyou Xia, Shengzhe Liu, Rong Qin, Guoli Jia 等ACM MM 2024 · 被引用 3 次
- Emotion and Intention Guided Multi-Modal Learning for Sticker Response SelectionYuxuan Hu, Jian Chen, Yuhao Wang, Zixuan Li 等AAAI 2026
- Deconfounded Emotion Guidance Sticker Selection with Causal InferenceJiali Chen, Yi Cai, Ruohang Xu, Jiexin Wang 等ACM MM 2024 · 被引用 5 次
- SER30K: A Large-Scale Dataset for Sticker Emotion RecognitionShengzhe Liu, Xin Zhang, Jufeng YangACM MM 2022 · 被引用 14 次
- TGCA-PVT: Topic-Guided Context-Aware Pyramid Vision Transformer for Sticker Emotion RecognitionJian Chen, Wei Wang, Yuzhu Hu, Junxin Chen 等ACM MM 2024 · 被引用 3 次
