Learning to Respond with Stickers: A Framework of Unifying Multi-Modality in Multi-Turn Dialog
Shen Gao, Xiuying Chen, Chang Liu, Li Liu, Dongyan Zhao, Rui Yan
Abstract
Stickers with vivid and engaging expressions are becoming increasingly popular in online messaging apps, and some works are dedicated to automatically select sticker response by matching text labels of stickers with previous utterances. However, due to their large quantities, it is impractical to require text labels for the all stickers. Hence, in this paper, we propose to recommend an appropriate sticker to user based on multi-turn dialog context history without any external labels. Two main challenges are confronted in this task. One is to learn semantic meaning of stickers without corresponding text labels. Another challenge is to jointly model the candidate sticker with the multi-turn dialog context. To tackle these challenges, we propose a sticker response selector (SRS) model. Specifically, SRS first employs a convolutional based sticker image encoder and a self-attention based multi-turn dialog encoder to obtain the representation of stickers and utterances. Next, deep interaction network is proposed to conduct deep matching between the sticker with each utterance in the dialog history. SRS then learns the short-term and long-term dependency between all interaction results by a fusion network to output the the final matching score. To evaluate our proposed method, we collect a large-scale realworld dialog dataset with stickers from one of the most popular online chatting platform. Extensive experiments conducted on this dataset show that our model achieves the state-of-the-art performance for all commonly-used metrics. Experiments also verify the * Equal contribution. Ordering is decided by a coin flip. Work performed during an internship at IIAI. † WICT is the abbreviation of Wangxuan Institute of Computer Technology.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 873e97cd-19ce-4de6-af8e-a0a918332ad4Cited by top-tier papers7
- VMSMO: Learning to Generate Multimodal Summary for Video-based News ArticlesMingzhe Li, Xiuying Chen, Shen Gao, Zhangming Chan et al.EMNLP 2020 · 65 citations
- Reasoning in Dialog: Improving Response Generation by Context Reading ComprehensionXiuying Chen, Zhi Cui, Jiayi Zhang, Chen Wei et al.AAAI 2021 · 15 citations
- Context-to-Session Matching: Utilizing Whole Session for Response Selection in Information-Seeking Dialogue SystemsZhenxin Fu, Shaobo Cui, Mingyue Shang, Feng Ji et al.KDD 2020 · 13 citations
- A New Formula for Sticker Retrieval: Reply with Stickers in Multi-Modal and Multi-Session ConversationBingbing Wang, Yiming Du, Bin Liang, Zhixin Bai et al.AAAI 2025 · 5 citations
- PerSRV: Personalized Sticker Retrieval with Vision-Language ModelHeng Er Metilda Chee, Jiayin Wang, Zhiqiang Guo, Weizhi Ma et al.WWW 2025 · 3 citations
Builds on1
Related papers
- Perceive before Respond: Improving Sticker Response Selection by Emotion Distillation and Hard MiningWuyou Xia, Shengzhe Liu, Rong Qin, Guoli Jia et al.ACM MM 2024 · 3 citations
- Emotion and Intention Guided Multi-Modal Learning for Sticker Response SelectionYuxuan Hu, Jian Chen, Yuhao Wang, Zixuan Li et al.AAAI 2026
- Deconfounded Emotion Guidance Sticker Selection with Causal InferenceJiali Chen, Yi Cai, Ruohang Xu, Jiexin Wang et al.ACM MM 2024 · 5 citations
- SER30K: A Large-Scale Dataset for Sticker Emotion RecognitionShengzhe Liu, Xin Zhang, Jufeng YangACM MM 2022 · 14 citations
- TGCA-PVT: Topic-Guided Context-Aware Pyramid Vision Transformer for Sticker Emotion RecognitionJian Chen, Wei Wang, Yuzhu Hu, Junxin Chen et al.ACM MM 2024 · 3 citations
