Learning to Caption Images Through a Lifetime by Asking Questions
Tingke Shen, Amlan Kar, Sanja Fidler
Abstract
In order to bring artificial agents into our lives, we will need to go beyond supervised learning on closed datasets to having the ability to continuously expand knowledge. Inspired by a student learning in a classroom, we present an agent that can continuously learn by posing natural language questions to humans. Our agent is composed of three interacting modules, one that performs captioning, another that generates questions and a decision maker that learns when to ask questions by implicitly reasoning about the uncertainty of the agent and expertise of the teacher. As compared to current active learning methods which query images for full captions, our agent is able to ask pointed questions to improve the generated captions. The agent trains on the improved captions, expanding its knowledge. We show that our approach achieves better performance using less human supervision than the baselines on the challenging MSCOCO [15] dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4b6ca6d0-9a12-48f3-85b2-7182ae81bf1fCited by top-tier papers4
- Self-Motivated Communication Agent for Real-World Vision-Dialog NavigationYi Zhu, Yue Weng, Fengda Zhu, Xiaodan Liang et al.ICCV 2021 · 41 citations
- Few-Shot Continual Active Learning by a RobotAli Ayub, Carter FendleyNeurIPS 2022 · 36 citations
- Unsupervised Commonsense Question Answering with Self-TalkVered Shwartz, Peter West, Ronan Le Bras, Chandra Bhagavatula et al.EMNLP 2020 · 25 citations
- CapWAP: Image Captioning with a PurposeAdam Fisch, Kenton Lee, Ming-Wei Chang, Jonathan H. Clark et al.EMNLP 2020 · 17 citations
Related papers
- Just Ask: An Interactive Learning Framework for Vision and Language NavigationTa-Chung Chi, Minmin Shen, Mihail Eric, Seokhwan Kim et al.AAAI 2020 · 88 citations
- Where to Go for the Holidays: Towards Mixed-Type Dialogs for Clarification of User GoalsZeming Liu, Jun Xu, Zeyang Lei, Haifeng Wang et al.ACL 2022 · 18 citations
- Gold Seeker: Information Gain From Policy Distributions for Goal-Oriented Vision-and-Langauge ReasoningEhsan Abbasnejad, Iman Abbasnejad, Qi Wu, Javen Shi et al.CVPR 2020
- Dialog Policy Learning for Joint Clarification and Active Learning QueriesAishwarya Padmakumar, Raymond J. MooneyAAAI 2021 · 12 citations
- Intra-agent speech permits zero-shot task acquisitionChen Yan, Federico Carnevale, Petko Georgiev, Adam Santoro et al.NeurIPS 2022 · 10 citations
