Learning to Generate Visual Questions with Noisy Supervision
Kai Shen, Lingfei Wu, Siliang Tang, Yueting Zhuang, Zhen He, Zhuoye Ding, Yun Xiao, Bo Long
摘要
The task of visual question generation (VQG) aims to generate human-like neural questions from an image and potentially other side information (e.g., answer type or the answer itself). Existing works often suffer from the severe one image to many questions mapping problem, which generates uninformative and non-referential questions. Recent work has demonstrated that by leveraging double visual and answer hints, a model can faithfully generate much better quality questions. However, visual hints are not available naturally. Despite they proposed a simple rule-based similarity matching method to obtain candidate visual hints, they could be very noisy practically and thus restrict the quality of generated questions. In this paper, we present a novel learning approach for double-hints based VQG, which can be cast as a weakly supervised learning problem with noises. The key rationale is that the salient visual regions of interest can be viewed as a constraint to improve the generation procedure for producing high-quality questions. As a result, given the predicted salient visual regions of interest, we can focus on estimating the probability of being ground-truth questions, which in turn implicitly measures the quality of predicted visual hints. Experimental results on two benchmark datasets show that our proposed method outperforms the state-of-the-art approaches by a large margin on a variety of metrics, including both automatic machine metrics and human evaluation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- End-to-End Modeling via Information Tree for One-Shot Natural Language Spatial Video GroundingMengze Li, Tianbao Wang, Haoyu Zhang, Shengyu Zhang 等ACL 2022 · 被引用 46 次
- Towards Interactivity and Interpretability: A Rationale-based Legal Judgment Prediction FrameworkYiquan Wu, Yifei Liu, Weiming Lu, Yating Zhang 等EMNLP 2022 · 被引用 33 次
- ML-LJP: Multi-Law Aware Legal Judgment PredictionYifei Liu, Yiquan Wu, Yating Zhang, Changlong Sun 等SIGIR 2023 · 被引用 31 次
- De-biased Attention Supervision for Text Classification with CausalityYiquan Wu, Yifei Liu, Ziyu Zhao, Weiming Lu 等AAAI 2024 · 被引用 10 次
- Dr.Academy: A Benchmark for Evaluating Questioning Capability in Education for Large Language ModelsYuyan Chen, Songzhou Yan, Panjun Liu, Yanghua XiaoACL 2024 · 被引用 8 次
它引用的顶会 Paper9
- Attacks Which Do Not Kill Training Make Adversarial Learning StrongerJingfeng Zhang, Xilie Xu, Bo Han, Gang Niu 等ICML 2020 · 被引用 452 次
- Domain Generalization via Entropy RegularizationShanshan Zhao, Mingming Gong, Tongliang Liu, Huan Fu 等NeurIPS 2020 · 被引用 327 次
- Geometry-aware Instance-reweighted Adversarial TrainingJingfeng Zhang, Jianing Zhu, Gang Niu, Bo Han 等ICLR 2021 · 被引用 316 次
- Reinforcement Learning Based Graph-to-Sequence Model for Natural Question GenerationYu Chen, Lingfei Wu, Mohammed J. ZakiICLR 2020 · 被引用 167 次
- De-Biased Court's View Generation with CausalityYiquan Wu, Kun Kuang, Yating Zhang, Xiaozhong Liu 等EMNLP 2020 · 被引用 71 次
相关 Paper
- Inferential Visual Question GenerationChao Bi, Shuhui Wang, Zhe Xue, Shengbo Chen 等ACM MM 2022 · 被引用 8 次
- ConVQG: Contrastive Visual Question Generation with Multimodal GuidanceLi Mi, Syrielle Montariol, Javiera Castillo Navarro, Xianjie Dai 等AAAI 2024 · 被引用 8 次
- Deconfounded Visual Question Generation with Causal InferenceJiali Chen, Zhenjun Guo, Jiayuan Xie, Yi Cai 等ACM MM 2023 · 被引用 8 次
- Multiple Objects-Aware Visual Question GenerationJiayuan Xie, Yi Cai, Qingbao Huang, Tao WangACM MM 2021 · 被引用 23 次
- Re-Attention for Visual Question AnsweringWenya Guo, Ying Zhang, Xiaoping Wu, Jufeng Yang 等AAAI 2020 · 被引用 90 次
