MuMuQA: Multimedia Multi-Hop News Question Answering via Cross-Media Knowledge Extraction and Grounding
Revanth Gangi Reddy, Xilin Rui, Manling Li, Xudong Lin, Haoyang Wen, Jaemin Cho, Lifu Huang, Mohit Bansal, Avirup Sil, Shih-Fu Chang, Alexander G. Schwing, Heng Ji
摘要
Recently, there has been an increasing interest in building question answering (QA) models that reason across multiple modalities, such as text and images. However, QA using images is often limited to just picking the answer from a pre-defined set of options. In addition, images in the real world, especially in news, have objects that are co-referential to the text, with complementary information from both modalities. In this paper, we present a new QA evaluation benchmark with 1,384 questions over news articles that require cross-media grounding of objects in images onto text. Specifically, the task involves multi-hop questions that require reasoning over image-caption pairs to identify the grounded visual object being referred to and then predicting a span from the news body text to answer the question. In addition, we introduce a novel multimedia data augmentation framework, based on cross-media knowledge extraction and synthetic question-answer generation, to automatically augment data that can provide weak supervision for this task. We evaluate both pipeline-based and end-to-end pretraining-based multimedia QA models on our benchmark, and show that they achieve promising performance, while considerably lagging behind human performance hence leaving large room for future work on this challenging new task.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- ViperGPT: Visual Inference via Python Execution for ReasoningDídac Surís, Sachit Menon, Carl VondrickICCV 2023 · 被引用 732 次
- MuRAG: Multimodal Retrieval-Augmented Generator for Open Question Answering over Images and TextWenhu Chen, Hexiang Hu, Xi Chen, Pat Verga 等EMNLP 2022 · 被引用 89 次
- Decompose, Analyze and Rethink: Solving Intricate Problems with Human-like Reasoning CycleShangzi Xue, Zhenya Huang, Jiayu Liu, Xin Lin 等NeurIPS 2024 · 被引用 55 次
- Enhancing Multi-modal Multi-hop Question Answering via Structured Knowledge and Unified Retrieval-GenerationQian Yang, Qian Chen, Wen Wang, Baotian Hu 等ACM MM 2023 · 被引用 14 次
- Manta: Enhancing Mamba for Few-Shot Action Recognition of Long Sub-SequenceWenbo Huang, Jinghui Zhang, Guang Li, Lei Zhang 等AAAI 2025 · 被引用 10 次
它引用的顶会 Paper8
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- nocaps: novel object captioning at scaleHarsh Agrawal, Peter Anderson, Karan Desai, Yufei Wang 等ICCV 2019 · 被引用 631 次
- MultiModalQA: complex question answering over text, tables and imagesAlon Talmor, Ori Yoran, Amnon Catav, Dan Lahav 等ICLR 2021 · 被引用 229 次
- Graph Optimal Transport for Cross-Domain AlignmentLiqun Chen, Zhe Gan, Yu Cheng, Linjie Li 等ICML 2020 · 被引用 193 次
- Cross-media Structured Common Space for Multimedia Event ExtractionManling Li, Alireza Zareian, Qi Zeng, Spencer Whitehead 等ACL 2020 · 被引用 87 次
相关 Paper
- CRIT: Graph-Based Automatic Data Synthesis to Enhance Cross-Modal Multi-Hop ReasoningJunyoung Sung, Seungwoo Lyu, Minjun Kim, Sumin An 等CVPR 2026 · 被引用 2 次
- VTQA: Visual Text Question Answering via Entity Alignment and Cross-Media ReasoningKang Chen, Xiangqian WuCVPR 2024
- Visual News: Benchmark and Challenges in News Image CaptioningFuxiao Liu, Yinghan Wang, Tianlu Wang, Vicente OrdonezEMNLP 2021 · 被引用 67 次
- Knowledge Completes the Vision: A Multimodal Entity-aware Retrieval-Augmented Generation Framework for News Image CaptioningXiaoxing You, Qiang Huang, Lingyu Li, Chi Zhang 等AAAI 2026 · 被引用 1 次
- Fine-tuning with Multi-modal Entity Prompts for News Image CaptioningJingjing Zhang, Shancheng Fang, Zhendong Mao, Zhiwei Zhang 等ACM MM 2022 · 被引用 16 次
