Lecture Presentations Multimodal Dataset: Towards Understanding Multimodality in Educational Videos
Dong Won Lee, Chaitanya Ahuja, Paul Pu Liang, Sanika Natu, Louis-Philippe Morency
Abstract
Many educational videos use slide presentations, a sequence of visual pages that contain text and figures accompanied by spoken language, which are constructed and presented carefully in order to optimally transfer knowledge to students. Previous studies in multimedia and psychology attribute the effectiveness of lecture presentations to their multimodal nature. As a step toward developing vision-language models to aid in student learning as intelligent teacher assistants, we introduce the Lecture Presentations Multimodal (LPM) Dataset as a large-scale benchmark testing the capabilities of vision-and-language models in multimodal understanding of educational videos. Our dataset contains aligned slides and spoken language, for 180+ hours of video and 9000+ slides, with 10 lecturers from various subjects (e.g., computer science, dentistry, biology). We introduce three research tasks, (1) figure-to-text retrieval, (2) text-to-figure retrieval, and (3) generation of slide explanations, which are grounded in multimedia learning and psychology principles to test a vision-language model’s understanding of multimodal content. We provide manual annotations to help implement these tasks and establish baselines on them. Comparing baselines and human student performances, we find that state-of-the-art vision-language models (zero-shot and fine-tuned) struggle in (1) weak crossmodal alignment between slides and spoken text, (2) learning novel visual mediums, (3) technical language, and (4) long-range sequences. We introduce PolyViLT, a novel multimodal transformer trained with a multi-instance learning loss that is more effective than current approaches for retrieval. We conclude by shedding light on the challenges and opportunities in multimodal understanding of educational presentation videos.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dae0fffe-14b7-40dc-9cb9-cd8ba5d28924Cited by top-tier papers4
- Attention Bootstrapping for Multi-Modal Test-Time AdaptationYusheng Zhao, Junyu Luo, Xiao Luo, Jinsheng Huang et al.AAAI 2025 · 5 citations
- Towards Comprehensive Lecture Slides Understanding: Large-Scale Dataset and Effective MethodEnming Zhang, Yuzhe Li, Yuliang Liu, Yingying Zhu et al.ICCV 2025 · 4 citations
- M³AV: A Multimodal, Multigenre, and Multipurpose Audio-Visual Academic Lecture DatasetZhe Chen, Heyang Liu, Wenyi Yu, Guangzhi Sun et al.ACL 2024 · 2 citations
- SlideAudit: A Dataset and Taxonomy for Automated Evaluation of Presentation SlidesZhuohao Jerry Zhang, Ruiqi Chen, Mingyuan Zhong, Jacob O. WobbrockUIST 2025
Builds on8
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu et al.NeurIPS 2022 · 2,727 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
- VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language ResearchXin Wang, Jiawei Wu, Jun-Kun Chen, Lei Li et al.ICCV 2019 · 688 citations
Related papers
- MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in VideosXuehai He, Weixi Feng, Kaizhi Zheng, Yujie Lu et al.ICLR 2025
- Vision-Language Models Do Not Understand NegationKumail Alhamoud, Shaden Alshammari, Yonglong Tian, Guohao Li et al.CVPR 2025
- MMICL: Empowering Vision-language Model with Multi-Modal In-Context LearningHaozhe Zhao, Zefan Cai, Shuzheng Si, Xiaojian Ma et al.ICLR 2024 · 206 citations
- Semantic Document Derendering: SVG Reconstruction via Vision-Language ModelingAdam Hazimeh, Ke Wang, Mark Collier, Gilles Baechler et al.AAAI 2026
- IPFormer: Instance Prompt-guided Transformer for Multi-modal Multi-shot Video UnderstandingYujia Liang, Jile Jiao, Xuetao Feng, Xinchen Liu et al.AAAI 2026
