EduVidQA: Generating and Evaluating Long-form Answers to Student Questions based on Lecture Videos
Sourjyadip Ray, Shubham Sharma, Somak Aditya, Pawan Goyal
摘要
As digital platforms redefine educational paradigms, ensuring interactivity remains vital for effective learning. This paper explores using Multimodal Large Language Models (MLLMs) to automatically respond to student questions from online lectures -a novel question answering task of real world significance. We introduce the EduVidQA Dataset with 5252 question-answer pairs (both synthetic and realworld) from 296 computer science videos covering diverse topics and difficulty levels. To understand the needs of the dataset and task evaluation, we empirically study the qualitative preferences of students, which we provide as an important contribution to this line of work. Our benchmarking experiments consist of 6 stateof-the-art MLLMs, through which we study the effectiveness of our synthetic data for finetuning, as well as showing the challenging nature of the task. We evaluate the models using both text-based and qualitative metrics, thus showing a nuanced perspective of the models' performance, which is paramount to future work. This work not only sets a benchmark for this important problem, but also opens exciting avenues for future research in the field of Natural Language Processing for Education. * indicates equal supervision Code and data: https://github.com/sourjyadip/ eduvidqa-emnlp25 0 5 10 15 20 25 30 2 Related Work 2.1 Educational Video QA Datasets Dataset Video Type Answer Type Avg Vid Length Reasoning Type TutorialVQA Tutorial Open Ended 1488 secs Comprehension How2QA Tutorial MCQ 17.45 secs Comprehension HowToVQA Tutorial MCQ 12.1 secs Comprehension YTCommentQA Tutorial Open Ended 524 secs Knowledge EduVidQA Lecture Open Ended 4054 secs Evaluation https://pypi.org/project/python-youtube/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- The Flan Collection: Designing Data and Methods for Effective Instruction TuningShayne Longpre, Le Hou, Tu Vu, Albert Webson 等ICML 2023 · 被引用 908 次
- LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic AlignmentBin Zhu, Bin Lin, Munan Ning, Yang Yan 等ICLR 2024 · 被引用 403 次
相关 Paper
- Lecture Presentations Multimodal Dataset: Towards Understanding Multimodality in Educational VideosDong Won Lee, Chaitanya Ahuja, Paul Pu Liang, Sanika Natu 等ICCV 2023 · 被引用 20 次
- VQAGuider: Guiding Multimodal Large Language Models to Answer Complex Video QuestionsYuyan Chen, Jiyuan Jia, Jiaxin Lu, Siyue Li 等ACL 2025
- ExpVid: A Benchmark for Experiment Video Understanding & ReasoningYicheng Xu, Yue Wu, Jiashuo Yu, Ziang Yan 等ICLR 2026 · 被引用 2 次
- EduDiag: A Benchmark for Educational Diagnostic Reasoning with Error Tracing and Correction on Large Multimodal ModelsJiali Chen, Yuqi Xue, Xusen Hei, DingBa Fu 等CVPR 2026 · 被引用 1 次
- YTCommentQA: Video Question Answerability in Instructional VideosSaelyne Yang, Sunghyun Park, Yunseok Jang, Moontae LeeAAAI 2024 · 被引用 6 次
