Gold Seeker: Information Gain From Policy Distributions for Goal-Oriented Vision-and-Langauge Reasoning
Ehsan Abbasnejad, Iman Abbasnejad, Qi Wu, Javen Shi, Anton van den Hengel
Abstract
As Computer Vision moves from passive analysis of pixels to active analysis of semantics, the breadth of information algorithms need to reason over has expanded significantly. One of the key challenges in this vein is the ability to identify the information required to make a decision, and select an action that will recover it. We propose a reinforcement-learning approach that maintains a distribution over its internal information, thus explicitly representing the ambiguity in what it knows, and needs to know, towards achieving its goal. Potential actions are then generated according to this distribution. For each potential action a distribution of the expected outcomes is calculated, and the value of the potential information gain assessed. The action taken is that which maximizes the potential information gain. We demonstrate this approach applied to two vision-and-language problems that have attracted significant recent interest, visual dialog and visual query generation. In both cases the method actively selects actions that will best reduce its internal uncertainty, and outperforms its competitors in achieving the goal of the challenge.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Unified Questioner Transformer for Descriptive Question Generation in Goal-Oriented Visual DialogueShoya Matsumori, Kosuke Shingyouchi, Yuki Abe, Yosuke Fukuchi et al.ICCV 2021 · 19 citations
- Bayesian Low-Rank Learning (Bella): A Practical Approach to Bayesian Neural NetworksBao Gia Doan, Afshar Shamsi, Xiao-Yu Guo, Arash Mohammadi et al.AAAI 2025 · 1 citation
Builds on1
Related papers
- VisPlay: Self-Evolving Vision-Language ModelsYicheng He, Chengsong Huang, Zongxia Li, Jiaxin Huang et al.CVPR 2026 · 3 citations
- Escaping the Mode: Multi-Answer Reinforcement Learning in LMsIsha Puri, Mehul Damani, Idan Shenfeld, Marzyeh Ghassemi et al.ICML 2026
- Unsupervised and Pseudo-Supervised Vision-Language Alignment in Visual DialogFeilong Chen, Duzhen Zhang, Xiuyi Chen, Jing Shi et al.ACM MM 2022 · 10 citations
- DMRM: A Dual-Channel Multi-Hop Reasoning Model for Visual DialogFeilong Chen, Fandong Meng, Jiaming Xu, Peng Li et al.AAAI 2020 · 35 citations
- Dialog Policy Learning for Joint Clarification and Active Learning QueriesAishwarya Padmakumar, Raymond J. MooneyAAAI 2021 · 12 citations
