Divide-and-Conquer: Tree-structured Strategy with Answer Distribution Estimator for Goal-Oriented Visual Dialogue
Shuo Cai, Xinzhe Han, Shuhui Wang
Abstract
Goal-oriented visual dialogue involves multi-round interaction between artificial agents, which has been of remarkable attention due to its wide applications. Given a visual scene, this task occurs when a Questioner asks an action-oriented question and an Answerer responds with the intent of letting the Questioner know the correct action to take. The quality of questions affects the accuracy and efficiency of the target search progress. However, existing methods lack a clear strategy to guide the generation of questions, resulting in the randomness in the search process and inconvergent results. We propose a Tree-Structured Strategy with Answer Distribution Estimator (TSADE) which guides the question generation by excluding half of the current candidate objects in each round. The above process is implemented by maximizing a binary reward inspired by the "divide-and-conquer" paradigm. We further design a candidate-minimization reward which encourages the model to narrow down the scope of candidate objects toward the end of the dialogue. We experimentally demonstrate that our method can enable the agents to achieve high task-oriented accuracy with fewer repeating questions and rounds compared to traditional ergodic question generation approaches. Qualitative results further show that TSADE facilitates agents to generate higher-quality questions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on6
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Object Goal Navigation using Goal-Oriented Semantic ExplorationDevendra Singh Chaplot, Dhiraj Gandhi, Abhinav Gupta, Ruslan SalakhutdinovNeurIPS 2020 · 857 citations
- Visual Dialogue State Tracking for Question GenerationWei Pang, Xiaojie WangAAAI 2020 · 34 citations
- What's "up" with vision-language models? Investigating their struggle with spatial reasoningAmita Kamath, Jack Hessel, Kai-Wei ChangEMNLP 2023 · 31 citations
- Answer-Driven Visual State Estimator for Goal-Oriented Visual DialogueZipeng Xu, Fangxiang Feng, Xiaojie Wang, Yushu Yang et al.ACM MM 2020 · 4 citations
Related papers
- Unified Questioner Transformer for Descriptive Question Generation in Goal-Oriented Visual DialogueShoya Matsumori, Kosuke Shingyouchi, Yuki Abe, Yosuke Fukuchi et al.ICCV 2021 · 19 citations
- Learning to Retrieve Videos by Asking QuestionsAvinash Madasu, Junier Oliva, Gedas BertasiusACM MM 2022 · 17 citations
- DualVD: An Adaptive Dual Encoding Model for Deep Visual Understanding in Visual DialogueXiaoze Jiang, Jing Yu, Zengchang Qin, Yingying Zhuang et al.AAAI 2020 · 72 citations
- Gold Seeker: Information Gain From Policy Distributions for Goal-Oriented Vision-and-Langauge ReasoningEhsan Abbasnejad, Iman Abbasnejad, Qi Wu, Javen Shi et al.CVPR 2020
- AI-VQA: Visual Question Answering based on Agent Interaction with InterpretabilityRengang Li, Cong Xu, Zhenhua Guo, Baoyu Fan et al.ACM MM 2022 · 7 citations
