Learning to Generate Visual Questions with Noisy Supervision
Kai Shen, Lingfei Wu, Siliang Tang, Yueting Zhuang, Zhen He, Zhuoye Ding, Yun Xiao, Bo Long
Abstract
The task of visual question generation (VQG) aims to generate human-like neural questions from an image and potentially other side information (e.g., answer type or the answer itself). Existing works often suffer from the severe one image to many questions mapping problem, which generates uninformative and non-referential questions. Recent work has demonstrated that by leveraging double visual and answer hints, a model can faithfully generate much better quality questions. However, visual hints are not available naturally. Despite they proposed a simple rule-based similarity matching method to obtain candidate visual hints, they could be very noisy practically and thus restrict the quality of generated questions. In this paper, we present a novel learning approach for double-hints based VQG, which can be cast as a weakly supervised learning problem with noises. The key rationale is that the salient visual regions of interest can be viewed as a constraint to improve the generation procedure for producing high-quality questions. As a result, given the predicted salient visual regions of interest, we can focus on estimating the probability of being ground-truth questions, which in turn implicitly measures the quality of predicted visual hints. Experimental results on two benchmark datasets show that our proposed method outperforms the state-of-the-art approaches by a large margin on a variety of metrics, including both automatic machine metrics and human evaluation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- End-to-End Modeling via Information Tree for One-Shot Natural Language Spatial Video GroundingMengze Li, Tianbao Wang, Haoyu Zhang, Shengyu Zhang et al.ACL 2022 · 46 citations
- Towards Interactivity and Interpretability: A Rationale-based Legal Judgment Prediction FrameworkYiquan Wu, Yifei Liu, Weiming Lu, Yating Zhang et al.EMNLP 2022 · 33 citations
- ML-LJP: Multi-Law Aware Legal Judgment PredictionYifei Liu, Yiquan Wu, Yating Zhang, Changlong Sun et al.SIGIR 2023 · 31 citations
- De-biased Attention Supervision for Text Classification with CausalityYiquan Wu, Yifei Liu, Ziyu Zhao, Weiming Lu et al.AAAI 2024 · 10 citations
- Dr.Academy: A Benchmark for Evaluating Questioning Capability in Education for Large Language ModelsYuyan Chen, Songzhou Yan, Panjun Liu, Yanghua XiaoACL 2024 · 8 citations
Builds on9
- Attacks Which Do Not Kill Training Make Adversarial Learning StrongerJingfeng Zhang, Xilie Xu, Bo Han, Gang Niu et al.ICML 2020 · 452 citations
- Domain Generalization via Entropy RegularizationShanshan Zhao, Mingming Gong, Tongliang Liu, Huan Fu et al.NeurIPS 2020 · 327 citations
- Geometry-aware Instance-reweighted Adversarial TrainingJingfeng Zhang, Jianing Zhu, Gang Niu, Bo Han et al.ICLR 2021 · 316 citations
- Reinforcement Learning Based Graph-to-Sequence Model for Natural Question GenerationYu Chen, Lingfei Wu, Mohammed J. ZakiICLR 2020 · 167 citations
- De-Biased Court's View Generation with CausalityYiquan Wu, Kun Kuang, Yating Zhang, Xiaozhong Liu et al.EMNLP 2020 · 71 citations
Related papers
- Inferential Visual Question GenerationChao Bi, Shuhui Wang, Zhe Xue, Shengbo Chen et al.ACM MM 2022 · 8 citations
- ConVQG: Contrastive Visual Question Generation with Multimodal GuidanceLi Mi, Syrielle Montariol, Javiera Castillo Navarro, Xianjie Dai et al.AAAI 2024 · 8 citations
- Deconfounded Visual Question Generation with Causal InferenceJiali Chen, Zhenjun Guo, Jiayuan Xie, Yi Cai et al.ACM MM 2023 · 8 citations
- Multiple Objects-Aware Visual Question GenerationJiayuan Xie, Yi Cai, Qingbao Huang, Tao WangACM MM 2021 · 23 citations
- Re-Attention for Visual Question AnsweringWenya Guo, Ying Zhang, Xiaoping Wu, Jufeng Yang et al.AAAI 2020 · 90 citations
