DOC2PPT: Automatic Presentation Slides Generation from Scientific Documents
Tsu-Jui Fu, William Yang Wang, Daniel McDuff, Yale Song
Abstract
Creating presentation materials requires complex multimodal reasoning skills to summarize key concepts and arrange them in a logical and visually pleasing manner. Can machines learn to emulate this laborious process? We present a novel task and approach for document-to-slide generation. Solving this involves document summarization, image and text retrieval, slide structure and layout prediction to arrange key elements in a form suitable for presentation. We propose a hierarchical sequence-to-sequence approach to tackle our task in an end-to-end manner. Our approach exploits the inherent structures within documents and slides and incorporates paraphrasing and layout prediction modules to generate slides. To help accelerate research in this domain, we release a dataset about 6K paired documents and slide decks used in our experiments. We show that our approach outperforms strong baselines and produces slides with rich content and aligned imagery.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7750caea-9e93-447b-8ba6-908ae6eb0dc3Cited by top-tier papers18
- LayoutGPT: Compositional Visual Planning and Generation with Large Language ModelsWeixi Feng, Wanrong Zhu, Tsu-Jui Fu, Varun Jampani et al.NeurIPS 2023 · 462 citations
- SlideVQA: A Dataset for Document Visual Question Answering on Multiple ImagesRyota Tanaka, Kyosuke Nishida, Kosuke Nishida, Taku Hasegawa et al.AAAI 2023 · 178 citations
- Slide4N: Creating Presentation Slides from Computational Notebooks with Human-AI CollaborationFengjie Wang, Xuye Liu, Oujing Liu, Ali Neshati et al.CHI 2023 · 37 citations
- LessonPlanner: Assisting Novice Teachers to Prepare Pedagogy-Driven Lesson Plans with Large Language ModelsHaoxiang Fan, Guanzheng Chen, Xingbo Wang, Zhenhui PengUIST 2024 · 35 citations
- From Paper to Card: Transforming Design Implications with Generative AIDonghoon Shin, Lucy Lu Wang, Gary HsiehCHI 2024 · 28 citations
Builds on7
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 2,453 citations
- Semi-Supervised Learning with Normalizing FlowsPavel Izmailov, Polina Kirichenko, Marc Finzi, Andrew Gordon WilsonICML 2020 · 134 citations
- Multimodal Summarization with Guidance of Multimodal ReferenceJunnan Zhu, Yu Zhou, Jiajun Zhang, Haoran Li et al.AAAI 2020 · 113 citations
- An Effective Transition-based Model for Discontinuous NERXiang Dai, Sarvnaz Karimi, Ben Hachey, Cécile ParisACL 2020 · 78 citations
- VMSMO: Learning to Generate Multimodal Summary for Video-based News ArticlesMingzhe Li, Xiuying Chen, Shen Gao, Zhangming Chan et al.EMNLP 2020 · 65 citations
Related papers
- Deep Submodular Optimization and LLM for Multimodal Content Extraction and Automatic Poster Generation from Long DocumentVijay Jaisankar, Sambaran Bandyopadhyay, Kalp Vyas, Varre Suman Chaitanya et al.AAAI 2025 · 1 citation
- Semantic Document Derendering: SVG Reconstruction via Vision-Language ModelingAdam Hazimeh, Ke Wang, Mark Collier, Gilles Baechler et al.AAAI 2026
- Telling Stories from Computational Notebooks: AI-Assisted Presentation Slides Creation for Presenting Data Science WorkChengbo Zheng, Dakuo Wang, April Yi Wang, Xiaojuan MaCHI 2022 · 53 citations
- SlideCoder: Layout-aware RAG-enhanced Hierarchical Slide Generation from DesignWenxin Tang, Jingyu Xiao, Wenxuan Jiang, Xi Xiao et al.EMNLP 2025 · 2 citations
- SlideTailor: Personalized Presentation Slide Generation for Scientific PapersWenzheng Zeng, Mingyu Ouyang, Langyuan Cui, Hwee Tou NgAAAI 2026 · 3 citations
