EduVidQA: Generating and Evaluating Long-form Answers to Student Questions based on Lecture Videos
Sourjyadip Ray, Shubham Sharma, Somak Aditya, Pawan Goyal
Abstract
As digital platforms redefine educational paradigms, ensuring interactivity remains vital for effective learning. This paper explores using Multimodal Large Language Models (MLLMs) to automatically respond to student questions from online lectures -a novel question answering task of real world significance. We introduce the EduVidQA Dataset with 5252 question-answer pairs (both synthetic and realworld) from 296 computer science videos covering diverse topics and difficulty levels. To understand the needs of the dataset and task evaluation, we empirically study the qualitative preferences of students, which we provide as an important contribution to this line of work. Our benchmarking experiments consist of 6 stateof-the-art MLLMs, through which we study the effectiveness of our synthetic data for finetuning, as well as showing the challenging nature of the task. We evaluate the models using both text-based and qualitative metrics, thus showing a nuanced perspective of the models' performance, which is paramount to future work. This work not only sets a benchmark for this important problem, but also opens exciting avenues for future research in the field of Natural Language Processing for Education. * indicates equal supervision Code and data: https://github.com/sourjyadip/ eduvidqa-emnlp25 0 5 10 15 20 25 30 2 Related Work 2.1 Educational Video QA Datasets Dataset Video Type Answer Type Avg Vid Length Reasoning Type TutorialVQA Tutorial Open Ended 1488 secs Comprehension How2QA Tutorial MCQ 17.45 secs Comprehension HowToVQA Tutorial MCQ 12.1 secs Comprehension YTCommentQA Tutorial Open Ended 524 secs Knowledge EduVidQA Lecture Open Ended 4054 secs Evaluation https://pypi.org/project/python-youtube/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 88d24270-cecc-41d3-96db-a93d5136be09Builds on15
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- The Flan Collection: Designing Data and Methods for Effective Instruction TuningShayne Longpre, Le Hou, Tu Vu, Albert Webson et al.ICML 2023 · 908 citations
- LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic AlignmentBin Zhu, Bin Lin, Munan Ning, Yang Yan et al.ICLR 2024 · 403 citations
Related papers
- Lecture Presentations Multimodal Dataset: Towards Understanding Multimodality in Educational VideosDong Won Lee, Chaitanya Ahuja, Paul Pu Liang, Sanika Natu et al.ICCV 2023 · 20 citations
- VQAGuider: Guiding Multimodal Large Language Models to Answer Complex Video QuestionsYuyan Chen, Jiyuan Jia, Jiaxin Lu, Siyue Li et al.ACL 2025
- ExpVid: A Benchmark for Experiment Video Understanding & ReasoningYicheng Xu, Yue Wu, Jiashuo Yu, Ziang Yan et al.ICLR 2026 · 2 citations
- EduDiag: A Benchmark for Educational Diagnostic Reasoning with Error Tracing and Correction on Large Multimodal ModelsJiali Chen, Yuqi Xue, Xusen Hei, DingBa Fu et al.CVPR 2026 · 1 citation
- YTCommentQA: Video Question Answerability in Instructional VideosSaelyne Yang, Sunghyun Park, Yunseok Jang, Moontae LeeAAAI 2024 · 6 citations
