ECIS-VQG: Generation of Entity-centric Information-seeking Questions from Videos
Arpan Phukan, Manish Gupta, Asif Ekbal
Abstract
Previous studies on question generation from videos have mostly focused on generating questions about common objects and attributes and hence are not entity-centric.In this work, we focus on the generation of entity-centric information-seeking questions from videos.Such a system could be useful for videobased learning, recommending "People Also Ask" questions, video-based chatbots, and factchecking.Our work addresses three key challenges: identifying question-worthy information, linking it to entities, and effectively utilizing multimodal signals.Further, to the best of our knowledge, there does not exist a large-scale dataset for this task.Most video question generation datasets are on TV shows, movies, or human activities or lack entitycentric information-seeking questions.Hence, we contribute a diverse dataset of YouTube videos, VIDEOQUESTIONS, consisting of 411 videos with 2265 manually annotated questions.We further propose a model architecture combining Transformers, rich context signals (titles, transcripts, captions, embeddings), and a combination of cross-entropy and contrastive loss function to encourage entity-centric question generation.Our best method yields BLEU, ROUGE, CIDEr, and METEOR scores of 71.3, 78.6, 7.31, and 81.9, respectively, demonstrating practical usability.We make the code and dataset publicly available 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Just Ask: Learning to Answer Questions from Millions of Narrated VideosAntoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev et al.ICCV 2021 · 345 citations
- Reinforcement Learning Based Graph-to-Sequence Model for Natural Question GenerationYu Chen, Lingfei Wu, Mohammed J. ZakiICLR 2020 · 167 citations
- Semantic Graphs for Generating Deep QuestionsLiangming Pan, Yuxi Xie, Yansong Feng, Tat-Seng Chua et al.ACL 2020 · 79 citations
Related papers
- Dense-Caption Matching and Frame-Selection Gating for Temporal Localization in VideoQAHyounghun Kim, Zineng Tang, Mohit BansalACL 2020 · 31 citations
- Detecting Moments and Highlights in Videos via Natural Language QueriesJie Lei, Tamara L. Berg, Mohit BansalNeurIPS 2021 · 425 citations
- YTCommentQA: Video Question Answerability in Instructional VideosSaelyne Yang, Sunghyun Park, Yunseok Jang, Moontae LeeAAAI 2024 · 6 citations
- InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and GenerationYi Wang, Yinan He, Yizhuo Li, Kunchang Li et al.ICLR 2024 · 467 citations
- Visual News: Benchmark and Challenges in News Image CaptioningFuxiao Liu, Yinghan Wang, Tianlu Wang, Vicente OrdonezEMNLP 2021 · 67 citations
