Disambiguating Reference in Visually Grounded Dialogues through Joint Modeling of Textual and Multimodal Semantic Structures
Shun Inadumi, Nobuhiro Ueda, Koichiro Yoshino
Abstract
Multimodal reference resolution, including phrase grounding, aims to understand the semantic relations between mentions and realworld objects. Phrase grounding between images and their captions is a well-established task. In contrast, for real-world applications, it is essential to integrate textual and multimodal reference resolution to unravel the reference relations within dialogue, especially in handling ambiguities caused by pronouns and ellipses. This paper presents a framework that unifies textual and multimodal reference resolution by mapping mention embeddings to object embeddings and selecting mentions or objects based on their similarity. 1 Our experiments show that learning textual reference resolution, such as coreference resolution and predicateargument structure analysis, positively affects performance in multimodal reference resolution. In particular, our model with coreference resolution performs better in pronoun phrase grounding than representative models for this task, MDETR and GLIP. Our qualitative analysis demonstrates that incorporating textual reference relations strengthens the confidence scores between mentions, including pronouns and predicates, and objects, which can reduce the ambiguities that arise in visually grounded dialogues. * Currently at NEC Corporation. 1 The code is publicly available at https://github.com/ SInadumi/mmrr . Would you [Φ NOM ] この コップ を [Φ DAT ] 取っ て 頂けますか ? me DAT this cup ACC take you NOM FPV of the System … Dialogue Text P 2 : Person 2 P 1 : Person 1 はい。 コーヒー カップ ですよね。 Yes. the coffee cup correct ? Indirect (DAT) Direct (Coref.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f3d2ae98-7049-4fac-a9d0-264f8de8fe7eBuilds on10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve et al.ICCV 2021 · 1,114 citations
- Align2Ground: Weakly Supervised Phrase Grounding Guided by Image-Caption AlignmentSamyak Datta, Karan Sikka, Anirban Roy, Karuna Ahuja et al.ICCV 2019 · 113 citations
- SIMMC 2.0: A Task-oriented Dialog Dataset for Immersive Multimodal ConversationsSatwik Kottur, Seungwhan Moon, Alborz Geramifard, Babak DamavandiEMNLP 2021 · 54 citations
Related papers
- Extending Phrase Grounding with Pronouns in Visual DialoguesPanzhong Lu, Xin Zhang, Meishan Zhang, Min ZhangEMNLP 2022 · 5 citations
- Cross-Modal Omni Interaction Modeling for Phrase GroundingTianyu Yu, Tianrui Hui, Zhihao Yu, Yue Liao et al.ACM MM 2020 · 14 citations
- DQ-DETR: Dual Query Detection Transformer for Phrase Extraction and GroundingShilong Liu, Shijia Huang, Feng Li, Hao Zhang et al.AAAI 2023 · 44 citations
- Who are you referring to? Coreference resolution in image narrationsArushi Goel, Basura Fernando, Frank Keller, Hakan BilenICCV 2023 · 8 citations
- Semi-supervised multimodal coreference resolution in image narrationsArushi Goel, Basura Fernando, Frank Keller, Hakan BilenEMNLP 2023 · 4 citations
