Impossible Videos
Zechen Bai, Hai Ci, Mike Zheng Shou
摘要
Synthetic videos nowadays is widely used to complement data scarcity and diversity of realworld videos. Current synthetic datasets primarily replicate real-world scenarios, leaving impossible, counterfactual and anti-reality video concepts underexplored. This work aims to answer two questions: 1) Can today's video generation models effectively follow prompts to create impossible video content? 2) Are today's video understanding models good enough for understanding impossible videos? To this end, we introduce IPV-BENCH, a novel benchmark designed to evaluate and foster progress in video understanding and generation. IPV-BENCH is underpinned by a comprehensive taxonomy, encompassing 4 domains, 14 categories. It features diverse scenes that defy physical, biological, geographical, or social laws. Based on the taxonomy, a prompt suite is constructed to evaluate video generation models, challenging their prompt following and creativity capabilities. In addition, a video benchmark is curated to assess Video-LLMs on their ability of understanding impossible videos, which particularly requires reasoning on temporal dynamics and world knowledge. Comprehensive evaluations reveal limitations and insights for future directions of video models, paving the way for next-generation video models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Hierarchical Fine-grained Preference Optimization for Physically Plausible Video GenerationHarold Haodong Chen, Haojian Huang, Qifeng Chen, Harry Yang 等NeurIPS 2025 · 被引用 25 次
- LeMiCa: Lexicographic Minimax Path Caching for Efficient Diffusion-Based Video GenerationHuanlin Gao, Ping Chen, Fuyuan Shi, Chao Tan 等NeurIPS 2025 · 被引用 9 次
- Show, Don't Tell: Morphing Latent Reasoning into Image GenerationHarold Haodong Chen, Xinxiang Yin, Wenjie Shu, Hongfei (Faye) Zhang 等ICML 2026 · 被引用 7 次
- ScalingAR: Scaling Confidence for Autoregressive Image GenerationHarold Haodong Chen, Xianfeng Wu, Wenjie Shu, Rongjin Guo 等ICML 2026 · 被引用 3 次
- VFXMaster: Unlocking Dynamic Visual Effect Generation via In-Context LearningBaolu Li, Yiming Zhang, Qinghe Wang, Liqian Ma 等SIGGRAPH 2026
它引用的顶会 Paper20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video GenerationJay Zhangjie Wu, Yixiao Ge, Xintao Wang, Stan Weixian Lei 等ICCV 2023 · 被引用 1,113 次
- Video-LLaVA: Learning United Visual Representation by Alignment Before ProjectionBin Lin, Yang Ye, Bin Zhu, Jiaxi Cui 等EMNLP 2024 · 被引用 231 次
- UniVTG: Towards Unified Video-Language Temporal GroundingKevin Qinghong Lin, Pengchuan Zhang, Joya Chen, Shraman Pramanick 等ICCV 2023 · 被引用 221 次
相关 Paper
- PhyWorldBench: A Comprehensive Evaluation of Physical Realism in Text-to-Video ModelsJing Gu, Xian Liu, Yu Zeng, Ashwin Nagarajan 等ICLR 2026 · 被引用 29 次
- PAI-Bench: A Comprehensive Benchmark For Physical AIFengzhe Zhou, Jiannan Huang, Jialuo Li, Deva Ramanan 等CVPR 2026 · 被引用 32 次
- MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in VideosXuehai He, Weixi Feng, Kaizhi Zheng, Yujie Lu 等ICLR 2025
- VBench: Comprehensive Benchmark Suite for Video Generative ModelsZiqi Huang, Yinan He, Jiashuo Yu, Fan Zhang 等CVPR 2024
- Your One-Stop Solution for AI-Generated Video DetectionLong Ma, Zihao Xue, Yan Wang, Zhiyuan Yan 等CVPR 2026 · 被引用 13 次
