Cross-Media Keyphrase Prediction: A Unified Framework with Multi-Modality Multi-Head Attention and Image Wordings
Yue Wang, Jing Li, Michael R. Lyu, Irwin King
Abstract
Social media produces large amounts of contents every day. To help users quickly capture what they need, keyphrase prediction is receiving a growing attention. Nevertheless, most prior efforts focus on text modeling, largely ignoring the rich features embedded in the matching images. In this work, we explore the joint effects of texts and images in predicting the keyphrases for a multimedia post. To better align social media style texts and images, we propose: (1) a novel Multi-Modality Multi-Head Attention (M 3 H-Att) to capture the intricate cross-media interactions; (2) image wordings, in forms of optical characters and image attributes, to bridge the two modalities. Moreover, we design a novel unified framework to leverage the outputs of keyphrase classification and generation and couple their advantages. Extensive experiments on a large-scale dataset 1 newly collected from Twitter show that our model significantly outperforms the previous state of the art based on traditional co-attentions. Further analyses show that our multi-head attention is able to attend information from various aspects and boost classification or generation in diverse scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 456d55e1-10bd-4489-b8df-8c5c843e5656Cited by top-tier papers4
- Towards Better Multi-modal Keyphrase Generation via Visual Entity Enhancement and Multi-granularity Image Noise FilteringYifan Dong, Suhang Wu, Fandong Meng, Jie Zhou et al.ACM MM 2023 · 3 citations
- Point-of-Interest Type Prediction using Text and ImagesDanae Sánchez Villegas, Nikolaos AletrasEMNLP 2021 · 2 citations
- Boosting Multi-modal Keyphrase Prediction with Dynamic Chain-of-Thought in Vision-Language ModelsQihang Ma, Shengyu Li, Jie Tang, Dingkang Yang et al.EMNLP 2025
- Augmenting Intra-Modal Understanding in MLLMs for Robust Multimodal Keyphrase GenerationJiajun Cao, Qinggang Zhang, Yunbo Tang, Zhishang Xiang et al.AAAI 2026
Builds on2
- Unified Vision-Language Pre-Training for Image Captioning and VQALuowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu et al.AAAI 2020 · 1,047 citations
- Cross-media Structured Common Space for Multimedia Event ExtractionManling Li, Alireza Zareian, Qi Zeng, Spencer Whitehead et al.ACL 2020 · 87 citations
Related papers
- Alt-Text with Context: Improving Accessibility for Images on TwitterNikita Srivatsan, Sofía Samaniego, Omar Florez, Taylor Berg-KirkpatrickICLR 2024 · 9 citations
- Improving Multimodal Named Entity Recognition via Entity Span Detection with Unified Multimodal TransformerJianfei Yu, Jing Jiang, Li Yang, Rui XiaACL 2020 · 260 citations
- Multimodal Representation with Embedded Visual Guiding Objects for Named Entity Recognition in Social Media PostsZhiwei Wu, Changmeng Zheng, Yi Cai, Junying Chen et al.ACM MM 2020 · 139 citations
- Multi-Modality Cross Attention Network for Image and Sentence MatchingXi Wei, Tianzhu Zhang, Yan Li, Yongdong Zhang et al.CVPR 2020
- Diverse, Controllable, and Keyphrase-Aware: A Corpus and Method for News Multi-Headline GenerationDayiheng Liu, Yeyun Gong, Yu Yan, Jie Fu et al.EMNLP 2020 · 14 citations
