Rethinking Multimodal Entity and Relation Extraction from a Translation Point of View
Changmeng Zheng, Junhao Feng, Yi Cai, Xiaoyong Wei, Qing Li
Abstract
We revisit the multimodal entity and relation extraction from a translation point of view. Special attention is paid on the misalignment issue in text-image datasets which may mislead the learning. We are motivated by the fact that the cross-modal misalignment is a similar problem of cross-lingual divergence issue in machine translation. The problem can then be transformed and existing solutions can be borrowed by treating a text and its paired image as the translation to each other. We implement a multimodal back-translation using diffusion-based generative models for pseudo-paralleled pairs and a divergence estimator by constructing a high-resource corpora as a bridge for low-resource learners. Fine-grained confidence scores are generated to indicate both types and degrees of alignments with which better representations are obtained. The method has been validated in the experiments by outperforming 14 state-of-the-art methods in both entity and relation extraction tasks. The source code is available at https://github.com/thecharm/TMR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5d98dfc8-9536-4db8-8285-35ae7511378eCited by top-tier papers4
- Semantic Codebook Learning for Dynamic Recommendation ModelsZheqi Lv, Shaoxuan He, Tianyu Zhan, Shengyu Zhang et al.ACM MM 2024 · 8 citations
- Multi-Level Cross-Modal Alignment for Speech Relation ExtractionLiang Zhang, Zhen Yang, Biao Fu, Ziyao Lu et al.EMNLP 2024 · 2 citations
- REMOTE: A Unified Multimodal Relation Extraction Framework with Multilevel Optimal Transport and Mixture-of-ExpertsXinkui Lin, Yongxiu Xu, Minghao Tang, Shilong Zhang et al.ACM MM 2025 · 2 citations
- Retrieval over Classification: Integrating Relation Semantics for Multimodal Relation ExtractionLei Hei, Tingjing Liao, Peiyingxin, Yiyang Qi et al.EMNLP 2025
Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- Learning Implicit Entity-object Relations by Bidirectional Generative Alignment for Multimodal NERFeng Chen, Jiajia Liu, Kaixiang Ji, Wang Ren et al.ACM MM 2023 · 14 citations
- CL2CM: Improving Cross-Lingual Cross-Modal Retrieval via Cross-Lingual Knowledge TransferYabing Wang, Fan Wang, Jianfeng Dong, Hao LuoAAAI 2024 · 20 citations
- Bridging the Gap between Synthetic and Authentic Images for Multimodal Machine TranslationWenyu Guo, Qingkai Fang, Dong Yu, Yang FengEMNLP 2023 · 5 citations
- U-MERE: Unconstrained Multimodal Entity and Relation Extraction with Collaborative Modeling and Order-Sensitive OptimizationWei Jia, Li Jin, Kaiwen Wei, Yuying Shang et al.ACM MM 2025
- Prompt Me Up: Unleashing the Power of Alignments for Multimodal Entity and Relation ExtractionXuming Hu, Junzhe Chen, Aiwei Liu, Shiao Meng et al.ACM MM 2023 · 30 citations
