A Span-based Multimodal Variational Autoencoder for Semi-supervised Multimodal Named Entity Recognition
Baohang Zhou, Ying Zhang, Kehui Song, Wenya Guo, Guoqing Zhao, Hongbin Wang, Xiaojie Yuan
摘要
Multimodal named entity recognition (MNER) on social media is a challenging task which aims to extract named entities in free text and incorporate images to classify them into user-defined types. However, the annotation for named entities on social media demands a mount of human efforts. The existing semi-supervised named entity recognition methods focus on the text modal and are utilized to reduce labeling costs in traditional NER. However, the previous methods are not efficient for semi-supervised MNER. Because the MNER task is defined to combine the text information with image one and needs to consider the mismatch between the posted text and image. To fuse the text and image features for MNER effectively under semi-supervised setting, we propose a novel span-based multimodal variational autoencoder (SMVAE) model for semi-supervised MNER. The proposed method exploits modal-specific VAEs to model text and image latent features, and utilizes product-of-experts to acquire multimodal features. In our approach, the implicit relations between labels and multimodal features are modeled by multimodal VAE. Thus, the useful information of unlabeled data can be exploited in our method under semi-supervised setting. Experimental results on two benchmark datasets demonstrate that our approach not only outperforms baselines under supervised setting, but also improves MNER performance with less labeled data than existing semi-supervised methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper3
- Improving Multimodal Named Entity Recognition via Entity Span Detection with Unified Multimodal TransformerJianfei Yu, Jing Jiang, Li Yang, Rui XiaACL 2020 · 被引用 260 次
- Multi-modal Graph Fusion for Named Entity Recognition with Targeted Visual GuidanceDong Zhang, Suzhong Wei, Shoushan Li, Hanqian Wu 等AAAI 2021 · 被引用 240 次
- SeqVAT: Virtual Adversarial Training for Semi-Supervised Sequence LabelingLuoxin Chen, Weitong Ruan, Xinyue Liu, Jianhua LuACL 2020 · 被引用 118 次
相关 Paper
- Adaptive Transformer-Based Conditioned Variational Autoencoder for Incomplete Social Event ClassificationZhangming Li, Shengsheng Qian, Jie Cao, Quan Fang 等ACM MM 2022 · 被引用 19 次
- Multimodal Graph-Based Variational Mixture of Experts Network for Zero-Shot Multimodal Information ExtractionBaohang Zhou, Ying Zhang, Yu Zhao, Xuhui Sui 等WWW 2025 · 被引用 5 次
- Learning from Different text-image Pairs: A Relation-enhanced Graph Convolutional Network for Multimodal NERFei Zhao, Chunhui Li, Zhen Wu, Shangyu Xing 等ACM MM 2022 · 被引用 59 次
- Grounded Multimodal Named Entity Recognition on Social MediaJianfei Yu, Ziyan Li, Jieming Wang, Rui XiaACL 2023 · 被引用 32 次
- Learning Implicit Entity-object Relations by Bidirectional Generative Alignment for Multimodal NERFeng Chen, Jiajia Liu, Kaixiang Ji, Wang Ren 等ACM MM 2023 · 被引用 14 次
