MGHFT: Multi-Granularity Hierarchical Fusion Transformer for Cross-Modal Sticker Emotion Recognition
Jian Chen, Yuxuan Hu, Haifeng Lu, Wei Wang, Min Yang, Chengming Li, Xiping Hu
Abstract
Although pre-trained visual models with text have demonstrated strong capabilities in visual feature extraction, sticker emotion understanding remains challenging due to its reliance on multi-view information, such as background knowledge and stylistic cues. To address this, we propose a novel multi-granularity hierarchical fusion transformer (MGHFT), with a multi-view sticker interpreter based on Multimodal Large Language Models. Specifically, inspired by the human ability to interpret sticker emotions from multiple views, we first use Multimodal Large Language Models to interpret stickers by providing rich textual context via multi-view descriptions. Then, we design a hierarchical fusion strategy to fuse the textual context into visual understanding, which builds upon a pyramid visual transformer to extract both global and local sticker features at multiple stages. Through contrastive learning and attention mechanisms, textual features are injected at different stages of the visual backbone, enhancing the fusion of global- and local-granularity visual semantics with textual guidance. Finally, we introduce a text-guided fusion attention mechanism to effectively integrate the overall multimodal features, enhancing semantic understanding. Extensive experiments on 2 public sticker emotion datasets demonstrate that MGHFT significantly outperforms existing sticker emotion recognition approaches, achieving higher accuracy and more fine-grained emotion recognition. Compared to the best pre-trained visual models, our MGHFT also obtains an obvious improvement, 5.4% on F1 and 4.0% on accuracy. The code is released at https://github.com/cccccj-03/MGHFT\_ACMMM2025.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f6689e3c-16fe-4fdb-9ff8-b0fd7e177616Cited by top-tier papers1
Ask how each one uses itBuilds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- Rethinking Spatial Dimensions of Vision TransformersByeongho Heo, Sangdoo Yun, Dongyoon Han, Sanghyuk Chun et al.ICCV 2021 · 733 citations
Related papers
- SER30K: A Large-Scale Dataset for Sticker Emotion RecognitionShengzhe Liu, Xin Zhang, Jufeng YangACM MM 2022 · 14 citations
- TGCA-PVT: Topic-Guided Context-Aware Pyramid Vision Transformer for Sticker Emotion RecognitionJian Chen, Wei Wang, Yuzhu Hu, Junxin Chen et al.ACM MM 2024 · 3 citations
- PerSRV: Personalized Sticker Retrieval with Vision-Language ModelHeng Er Metilda Chee, Jiayin Wang, Zhiqiang Guo, Weizhi Ma et al.WWW 2025 · 3 citations
- Enhancing Meme Emotion Understanding with Multi-Level Modality Enhancement and Dual-Stage Modal FusionYi Shi, Wenlong Meng, Zhenyuan Guo, Chengkun Wei et al.AAAI 2026
- Affective Image Filter: Reflecting Emotions from Text to ImagesShuchen Weng, Peixuan Zhang, Zheng Chang, Xinlong Wang et al.ICCV 2023 · 25 citations
