Is Your World Simulator a Good Story Presenter? A Consecutive Events-Based Benchmark for Future Long Video Generation
Yiping Wang, Xuehai He, Kuan Wang, Luyao Ma, Jianwei Yang, Shuohang Wang, Simon Shaolei Du, Yelong Shen
摘要
The current state-of-the-art video generative models can produce commercial-grade videos with highly realistic details. However, they still struggle to coherently present multiple sequential events in the stories specified by the prompts, which is foreseeable an essential capability for future long video generation scenarios. For example, top T2V generative models still fail to generate a video of the short simple story "how to put an elephant into a refrigerator." While existing detail-oriented benchmarks primarily focus on fine-grained metrics like aesthetic quality and spatial-temporal consistency, they fall short of evaluating models’ abilities to handle event-level story presentation. To address this gap, we introduce StoryEval, a story-oriented benchmark specifically designed to assess text-to-video (T2V) models’ story-completion capabilities. StoryEval features 423 prompts spanning 7 classes, each representing short stories composed of 2–4 consecutive events. We employ Vision-Language Models, such as GPT-4o and LLaVA-OV-Chat-72B, to verify the completion of each event in the generated videos, applying a unanimous voting method to enhance reliability. Our methods ensure high alignment with human evaluations, and the evaluation of 11 models reveals its challenge, with none exceeding an average story-completion rate of 50%. StoryEval provides a new benchmark for advancing T2V models and highlights the challenges and opportunities in developing next-generation solutions for coherent story-driven video generation. Project website is available at https://ypwang61.github.io/project/StoryEval."The universe is made of stories, not of atoms."— Muriel Rukeyser
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Rethinking Video Generation Model for the Embodied WorldYufan Deng, Zilin Pan, Hongyu Zhang, Xiaojie Li 等ICML 2026 · 被引用 24 次
- WorldScore: A Unified Evaluation Benchmark for World GenerationHaoyi Duan, Hong-Xing Yu, Sirui Chen, Li Fei-Fei 等ICCV 2025 · 被引用 14 次
- NarrLV: Towards a Comprehensive Narrative-Centric Evaluation for Long Video GenerationXiaokun Feng, Haiming Yu, Meiqi Wu, Shiyu Hu 等ICLR 2026 · 被引用 13 次
- 4DWorldBench: A Comprehensive Evaluation Framework for 3D/4D World Generation ModelsYiting Lu, Wei Luo, Peiyan Tu, Haoran Li 等CVPR 2026 · 被引用 10 次
- Stable Cinemetrics : Structured Taxonomy and Evaluation for Professional Video GenerationAgneet Chatterjee, Rahim Entezari, Maksym Zhuravinskyi, Maksim Lapin 等NeurIPS 2025 · 被引用 6 次
它引用的顶会 Paper15
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language BenchmarkDongping Chen, Ruoxi Chen, Shilin Zhang, Yaochen Wang 等ICML 2024 · 被引用 345 次
- Scaling Autoregressive Video ModelsDirk Weissenborn, Oscar Täckström, Jakob UszkoreitICLR 2020 · 被引用 252 次
- FreeLong: Training-Free Long Video Generation with SpectralBlend Temporal AttentionYu Lu, Yuanzhi Liang, Linchao Zhu, Yi YangNeurIPS 2024 · 被引用 101 次
- DisenStudio: Customized Multi-Subject Text-to-Video Generation with Disentangled Spatial ControlHong Chen, Xin Wang, Yipeng Zhang, Yuwei Zhou 等ACM MM 2024 · 被引用 10 次
相关 Paper
- OSCBench: Benchmarking Object State Change in Text-to-Video GenerationXianjing Han, Bin Zhu, Shiqi Hu, Franklin Mingzhe Li 等ACL 2026 · 被引用 2 次
- LoCoT2V-Bench: Benchmarking Long-Form and Complex Text-to-Video GenerationXiangqing Zheng, CHENGYUE WU, Kehai Chen, Min zhangICML 2026 · 被引用 3 次
- AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video GenerationZiwei Zhou, Zeyuan Lai, Rui Wang, Yifan Yang 等ICML 2026 · 被引用 8 次
- T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video GenerationZhe Cao, Tao Wang, Jiaming Wang, Yanghai Wang 等ICML 2026 · 被引用 13 次
- ViStoryBench: Comprehensive Benchmark Suite for Story VisualizationCailin Zhuang, Ailin Huang, Hu Yaoqi, Jingwei Wu 等CVPR 2026 · 被引用 37 次
