Multi-VQG: Generating Engaging Questions for Multiple Images
Min-Hsuan Yeh, Vincent Chen, Ting-Hao 'Kenneth' Huang, Lun-Wei Ku
摘要
Generating engaging content has drawn much recent attention in the NLP community. Asking questions is a natural way to respond to photos and promote awareness. However, most answers to questions in traditional questionanswering (QA) datasets are factoids, which reduce individuals' willingness to answer. Furthermore, traditional visual question generation (VQG) confines the source data for question generation to single images, resulting in a limited ability to comprehend time-series information of the underlying event. In this paper, we propose generating engaging questions from multiple images. We present MVQG 1 , a new dataset, and establish a series of baselines, including both end-to-end and dual-stage architectures. Results show that building stories behind the image sequence enables models to generate engaging questions, which confirms our assumption that people typically construct a picture of the event in their minds before asking questions. These results open up an exciting challenge for visual-and-language models to implicitly construct a story behind a series of photos to allow for creativity and experience sharing and hence draw attention to downstream applications. How would you act if you found yourself in a room filled with cans of free drinks? Have you ever gone to beer tastings and where would that be at? How long did the cat lounge around in the book room? What would this cat sit on next?
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Chain-of-Exemplar: Enhancing Distractor Generation for Multimodal Educational Question GenerationHaohao Luo, Yang Deng, Ying Shen, See-Kiong Ng 等ACL 2024 · 被引用 6 次
- Location-Aware Visual Question Generation with Lightweight ModelsNicholas Collin Suwono, Justin Chih-Yao Chen, Tun-Min Hung, Ting-Hao 'Kenneth' Huang 等EMNLP 2023 · 被引用 3 次
它引用的顶会 Paper13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Unifying Vision-and-Language Tasks via Text GenerationJaemin Cho, Jie Lei, Hao Tan, Mohit BansalICML 2021 · 被引用 624 次
- How Much Can CLIP Benefit Vision-and-Language Tasks?Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal 等ICLR 2022 · 被引用 503 次
- Visually Grounded Reasoning across Languages and CulturesFangyu Liu, Emanuele Bugliarello, Edoardo Maria Ponti, Siva Reddy 等EMNLP 2021 · 被引用 87 次
相关 Paper
- CommVQA: Situating Visual Question Answering in Communicative ContextsNandita Naik, Christopher Potts, Elisa KreissEMNLP 2024
- Inferential Visual Question GenerationChao Bi, Shuhui Wang, Zhe Xue, Shengbo Chen 等ACM MM 2022 · 被引用 8 次
- Multi-Question Learning for Visual Question AnsweringChenyi Lei, Lei Wu, Dong Liu, Zhao Li 等AAAI 2020 · 被引用 9 次
- Explicitly Guided Difficulty-Controllable Visual Question GenerationJiayuan Xie, Mengqiu Cheng, Xinting Zhang, Yi Cai 等AAAI 2025 · 被引用 2 次
- Multiple Objects-Aware Visual Question GenerationJiayuan Xie, Yi Cai, Qingbao Huang, Tao WangACM MM 2021 · 被引用 23 次
