Disambiguating Reference in Visually Grounded Dialogues through Joint Modeling of Textual and Multimodal Semantic Structures
Shun Inadumi, Nobuhiro Ueda, Koichiro Yoshino
摘要
Multimodal reference resolution, including phrase grounding, aims to understand the semantic relations between mentions and realworld objects. Phrase grounding between images and their captions is a well-established task. In contrast, for real-world applications, it is essential to integrate textual and multimodal reference resolution to unravel the reference relations within dialogue, especially in handling ambiguities caused by pronouns and ellipses. This paper presents a framework that unifies textual and multimodal reference resolution by mapping mention embeddings to object embeddings and selecting mentions or objects based on their similarity. 1 Our experiments show that learning textual reference resolution, such as coreference resolution and predicateargument structure analysis, positively affects performance in multimodal reference resolution. In particular, our model with coreference resolution performs better in pronoun phrase grounding than representative models for this task, MDETR and GLIP. Our qualitative analysis demonstrates that incorporating textual reference relations strengthens the confidence scores between mentions, including pronouns and predicates, and objects, which can reduce the ambiguities that arise in visually grounded dialogues. * Currently at NEC Corporation. 1 The code is publicly available at https://github.com/ SInadumi/mmrr . Would you [Φ NOM ] この コップ を [Φ DAT ] 取っ て 頂けますか ? me DAT this cup ACC take you NOM FPV of the System … Dialogue Text P 2 : Person 2 P 1 : Person 1 はい。 コーヒー カップ ですよね。 Yes. the coffee cup correct ? Indirect (DAT) Direct (Coref.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve 等ICCV 2021 · 被引用 1,114 次
- Align2Ground: Weakly Supervised Phrase Grounding Guided by Image-Caption AlignmentSamyak Datta, Karan Sikka, Anirban Roy, Karuna Ahuja 等ICCV 2019 · 被引用 113 次
- SIMMC 2.0: A Task-oriented Dialog Dataset for Immersive Multimodal ConversationsSatwik Kottur, Seungwhan Moon, Alborz Geramifard, Babak DamavandiEMNLP 2021 · 被引用 54 次
相关 Paper
- Extending Phrase Grounding with Pronouns in Visual DialoguesPanzhong Lu, Xin Zhang, Meishan Zhang, Min ZhangEMNLP 2022 · 被引用 5 次
- Cross-Modal Omni Interaction Modeling for Phrase GroundingTianyu Yu, Tianrui Hui, Zhihao Yu, Yue Liao 等ACM MM 2020 · 被引用 14 次
- DQ-DETR: Dual Query Detection Transformer for Phrase Extraction and GroundingShilong Liu, Shijia Huang, Feng Li, Hao Zhang 等AAAI 2023 · 被引用 44 次
- Who are you referring to? Coreference resolution in image narrationsArushi Goel, Basura Fernando, Frank Keller, Hakan BilenICCV 2023 · 被引用 8 次
- Semi-supervised multimodal coreference resolution in image narrationsArushi Goel, Basura Fernando, Frank Keller, Hakan BilenEMNLP 2023 · 被引用 4 次
