Generative Multimodal Data Augmentation for Low-Resource Multimodal Named Entity Recognition
Ziyan Li, Jianfei Yu, Jia Yang, Wenya Wang, Li Yang, Rui Xia
摘要
As an important task in multimodal information extraction, Multimodal Named Entity Recognition (MNER) has recently attracted considerable attention. One key challenge of MNER lies in the lack of sufficient fine-grained annotated data, especially in low-resource scenarios. Although data augmentation is a widely used technique to tackle the above issue, it is challenging to simultaneously generate synthetic text-image pairs and their corresponding high-quality entity annotations. In this work, we propose a novel Generative Multimodal Data Augmentation (GMDA) framework for MNER, which contains two stages: Multimodal Text Generation and Multimodal Image Generation. Specifically, we first transform each annotated sentence into a linearized labeled sequence, and then train a Label-aware Multimodal Large Language Model (LMLLM) to generate the labeled sequence based on a label-aware prompt and its associated image. We further employ a Stable Diffusion model to generate the synthetic images that are semantically related to these sentences. Experimental results on three benchmark datasets demonstrate the effectiveness of the proposed GMDA framework, which consistently boosts the performance of several competitive methods for two subtasks of MNER in both full-supervision and low-resource settings. The low-resource dataset and source code are released at https://github.com/NUSTM/GMDA.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- Multi-Grained Query-Guided Set Prediction Network for Grounded Multimodal Named Entity RecognitionJielong Tang, Zhenxing Wang, Ziyang Gong, Jianxing Yu 等AAAI 2025 · 被引用 8 次
- SAKE: Self-aware Knowledge Exploitation-Exploration for Grounded Multimodal Named Entity RecognitionJielong Tang, Xujie Yuan, Jiayang Liu, Jianxing Yu 等KDD 2026 · 被引用 1 次
相关 Paper
- MELM: Data Augmentation with Masked Entity Language Modeling for Low-Resource NERRan Zhou, Xin Li, Ruidan He, Lidong Bing 等ACL 2022 · 被引用 114 次
- Exogenous and Endogenous Data Augmentation for Low-Resource Complex Named Entity RecognitionXinghua Zhang, Gaode Chen, Shiyao Cui, Jiawei Sheng 等SIGIR 2024 · 被引用 3 次
- RoPDA: Robust Prompt-Based Data Augmentation for Low-Resource Named Entity RecognitionSihan Song, Furao Shen, Jian ZhaoAAAI 2024 · 被引用 7 次
- DAGA: Data Augmentation with a Generation Approach forLow-resource Tagging TasksBosheng Ding, Linlin Liu, Lidong Bing, Canasai Kruengkrai 等EMNLP 2020 · 被引用 132 次
- Robust and Informative Text Augmentation (RITA) via Constrained Worst-Case Transformations for Low-Resource Named Entity RecognitionHyunwoo Sohn, Baekkwan ParkKDD 2022 · 被引用 3 次
