ECIS-VQG: Generation of Entity-centric Information-seeking Questions from Videos
Arpan Phukan, Manish Gupta, Asif Ekbal
摘要
Previous studies on question generation from videos have mostly focused on generating questions about common objects and attributes and hence are not entity-centric.In this work, we focus on the generation of entity-centric information-seeking questions from videos.Such a system could be useful for videobased learning, recommending "People Also Ask" questions, video-based chatbots, and factchecking.Our work addresses three key challenges: identifying question-worthy information, linking it to entities, and effectively utilizing multimodal signals.Further, to the best of our knowledge, there does not exist a large-scale dataset for this task.Most video question generation datasets are on TV shows, movies, or human activities or lack entitycentric information-seeking questions.Hence, we contribute a diverse dataset of YouTube videos, VIDEOQUESTIONS, consisting of 411 videos with 2265 manually annotated questions.We further propose a model architecture combining Transformers, rich context signals (titles, transcripts, captions, embeddings), and a combination of cross-entropy and contrastive loss function to encourage entity-centric question generation.Our best method yields BLEU, ROUGE, CIDEr, and METEOR scores of 71.3, 78.6, 7.31, and 81.9, respectively, demonstrating practical usability.We make the code and dataset publicly available 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Just Ask: Learning to Answer Questions from Millions of Narrated VideosAntoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev 等ICCV 2021 · 被引用 345 次
- Reinforcement Learning Based Graph-to-Sequence Model for Natural Question GenerationYu Chen, Lingfei Wu, Mohammed J. ZakiICLR 2020 · 被引用 167 次
- Semantic Graphs for Generating Deep QuestionsLiangming Pan, Yuxi Xie, Yansong Feng, Tat-Seng Chua 等ACL 2020 · 被引用 79 次
相关 Paper
- Dense-Caption Matching and Frame-Selection Gating for Temporal Localization in VideoQAHyounghun Kim, Zineng Tang, Mohit BansalACL 2020 · 被引用 31 次
- Detecting Moments and Highlights in Videos via Natural Language QueriesJie Lei, Tamara L. Berg, Mohit BansalNeurIPS 2021 · 被引用 425 次
- YTCommentQA: Video Question Answerability in Instructional VideosSaelyne Yang, Sunghyun Park, Yunseok Jang, Moontae LeeAAAI 2024 · 被引用 6 次
- InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and GenerationYi Wang, Yinan He, Yizhuo Li, Kunchang Li 等ICLR 2024 · 被引用 467 次
- Visual News: Benchmark and Challenges in News Image CaptioningFuxiao Liu, Yinghan Wang, Tianlu Wang, Vicente OrdonezEMNLP 2021 · 被引用 67 次
