Code2Video: A Code-centric Paradigm for Educational Video Creation
Yanzhe Chen, Kevin Qinghong Lin, Mike Zheng Shou
摘要
While recent generative models can synthesize videos in pixel space, they often fail to produce educational videos with precise structures, domain knowledge, and coherent transitions. We argue that this setting is better served by operating in a renderable environment that is explicitly controlled by code. We propose Code2Video , a code-centric agent framework that generates educational videos by writing executable Python programs. Code2Video includes three agents: a Planner that converts lecture content into a temporal storyboard, a Coder that turns the storyboard into runnable code with scope-guided auto-fix, and a Critic that refines layout using a VLM guided by visual anchor prompting , i.e. , mappings from target visual outcomes to code edits. For evaluation, we build MMMC , a benchmark of professionally produced, discipline-specific educational videos. We assess Code2Video using aesthetic scores (VLM-as-a-Judge), code efficiency, and TeachQuiz , an end-to-end metric that measures how well an unlearned VLM can recover knowledge after watching generated videos. Code2Video improves performance by 40% over direct code generation and produces videos comparable to human-crafted tutorials. The code and datasets are available at https://github.com/showlab/Code2Video .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
- Gorilla: Large Language Model Connected with Massive APIsShishir G. Patil, Tianjun Zhang, Xin Wang, Joseph E. GonzalezNeurIPS 2024 · 被引用 1,715 次
- StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video GenerationYupeng Zhou, Daquan Zhou, Ming-Ming Cheng, Jiashi Feng 等NeurIPS 2024 · 被引用 291 次
- In-Context Unlearning: Language Models as Few-Shot UnlearnersMartin Pawelczyk, Seth Neel, Himabindu LakkarajuICML 2024 · 被引用 217 次
相关 Paper
- Synthesis-Assisted Video Prototyping From a DocumentPeggy Chi, Tao Dong, Christian Früh, Brian Colonna 等UIST 2022 · 被引用 18 次
- OctoBench: Benchmarking Scaffold-Aware Instruction Following in Repository-Grounded Agentic CodingDeming Ding, Shichun Liu, Enhui Yang, Jiahang Lin 等ACL 2026 · 被引用 10 次
- PlayCoder: Making LLM-Generated GUI Code PlayableZhiyuan Peng, Wei Tao, Xin Yin, Chenhao Ying 等FSE 2026
- PlotCoder: Hierarchical Decoding for Synthesizing Visualization Code in Programmatic ContextXinyun Chen, Linyuan Gong, Alvin Cheung, Dawn SongACL 2021
- FullStack-Agent: Enhancing Agentic Full-Stack Web Coding via Development-Oriented Testing and Repository Back-TranslationZimu Lu, Houxing Ren, Yunqiao Yang, Ke Wang 等ICML 2026 · 被引用 1 次
