ACL2026

mPresenter: An Agentic Framework for Generating Multilingual Presentation Videos from Scientific Papers

Wenhan Han, Xiao Xiao, Mykola Pechenizkiy, Meng Fang

摘要

Generating presentation videos from academic papers is challenging due to the need for longdocument discourse planning and cross-lingual grounding. Existing Paper2Video systems are largely monolingual and often rely on singlepass pipelines, which can limit the coherence and informativeness of the resulting presentations. We present MPRESENTER, a multilingual agentic Paper2Video system that decomposes the task into planning, audience-oriented critique, layout-aware slide generation, and multilingual figure interpretation, enabling iterative refinement at the discourse level. To facilitate reproducible evaluation, we also introduce MPREBENCH, a multilingual benchmark that evaluates presentation videos via question answering as a proxy for effective information transfer. Experimental results indicate that MP-RESENTER improves question-answering accuracy relative to prior systems, while maintaining affordable cost and latency. We release both the system and benchmark to support further research on multilingual Paper2Video 1 . Inputs Reviewer Planner 1. Assets Extraction