Multi-VQG: Generating Engaging Questions for Multiple Images
Min-Hsuan Yeh, Vincent Chen, Ting-Hao 'Kenneth' Huang, Lun-Wei Ku
Abstract
Generating engaging content has drawn much recent attention in the NLP community. Asking questions is a natural way to respond to photos and promote awareness. However, most answers to questions in traditional questionanswering (QA) datasets are factoids, which reduce individuals' willingness to answer. Furthermore, traditional visual question generation (VQG) confines the source data for question generation to single images, resulting in a limited ability to comprehend time-series information of the underlying event. In this paper, we propose generating engaging questions from multiple images. We present MVQG 1 , a new dataset, and establish a series of baselines, including both end-to-end and dual-stage architectures. Results show that building stories behind the image sequence enables models to generate engaging questions, which confirms our assumption that people typically construct a picture of the event in their minds before asking questions. These results open up an exciting challenge for visual-and-language models to implicitly construct a story behind a series of photos to allow for creativity and experience sharing and hence draw attention to downstream applications. How would you act if you found yourself in a room filled with cans of free drinks? Have you ever gone to beer tastings and where would that be at? How long did the cat lounge around in the book room? What would this cat sit on next?
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b69be641-8656-4141-808a-e1174207f883Cited by top-tier papers2
- Chain-of-Exemplar: Enhancing Distractor Generation for Multimodal Educational Question GenerationHaohao Luo, Yang Deng, Ying Shen, See-Kiong Ng et al.ACL 2024 · 6 citations
- Location-Aware Visual Question Generation with Lightweight ModelsNicholas Collin Suwono, Justin Chih-Yao Chen, Tun-Min Hung, Ting-Hao 'Kenneth' Huang et al.EMNLP 2023 · 3 citations
Builds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Unifying Vision-and-Language Tasks via Text GenerationJaemin Cho, Jie Lei, Hao Tan, Mohit BansalICML 2021 · 624 citations
- How Much Can CLIP Benefit Vision-and-Language Tasks?Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal et al.ICLR 2022 · 503 citations
- Visually Grounded Reasoning across Languages and CulturesFangyu Liu, Emanuele Bugliarello, Edoardo Maria Ponti, Siva Reddy et al.EMNLP 2021 · 87 citations
Related papers
- CommVQA: Situating Visual Question Answering in Communicative ContextsNandita Naik, Christopher Potts, Elisa KreissEMNLP 2024
- Inferential Visual Question GenerationChao Bi, Shuhui Wang, Zhe Xue, Shengbo Chen et al.ACM MM 2022 · 8 citations
- Multi-Question Learning for Visual Question AnsweringChenyi Lei, Lei Wu, Dong Liu, Zhao Li et al.AAAI 2020 · 9 citations
- Explicitly Guided Difficulty-Controllable Visual Question GenerationJiayuan Xie, Mengqiu Cheng, Xinting Zhang, Yi Cai et al.AAAI 2025 · 2 citations
- Multiple Objects-Aware Visual Question GenerationJiayuan Xie, Yi Cai, Qingbao Huang, Tao WangACM MM 2021 · 23 citations
