OSCBench: Benchmarking Object State Change in Text-to-Video Generation
Xianjing Han, Bin Zhu, Shiqi Hu, Franklin Mingzhe Li, Patrick Carrington, Roger Zimmermann, Jingjing Chen
摘要
Text-to-video (T2V) generation models have made rapid progress in producing visually highquality and temporally coherent videos. However, existing benchmarks primarily focus on perceptual quality, text-video alignment, or physical plausibility, leaving a critical aspect of action understanding largely unexplored: object state change (OSC) explicitly specified in the text prompt. OSC refers to the transformation of an object's state induced by an action, such as peeling a potato or slicing a lemon. In this paper, we introduce OSCBench, a benchmark specifically designed to assess OSC performance in T2V models. OSCBench is constructed from instructional cooking data and systematically organizes action-object interactions into regular, novel, and compositional scenarios to probe both in-distribution performance and generalization. We evaluate six representative open-source and proprietary T2V models using both human user study and multimodal large language model (MLLM)-based automatic evaluation. Our results show that, despite strong performance on semantic and scene alignment, current T2V models consistently struggle with accurate and temporally consistent object state changes, especially in novel and compositional settings. These findings position OSC as a key bottleneck in textto-video generation and establish OSCBench as a diagnostic benchmark for advancing stateaware video generation models. Project page: https://hanxjing.github.io/OSCBench .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi 等ICCV 2019 · 被引用 1,437 次
- VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video GenerationXuan He, Dongfu Jiang, Ge Zhang, Max Ku 等EMNLP 2024 · 被引用 20 次
- Look for the Change: Learning Object States and State-Modifying Actions from Untrimmed Web VideosTomás Soucek, Jean-Baptiste Alayrac, Antoine Miech, Ivan Laptev 等CVPR 2022 · 被引用 19 次
- Redundancy Principles for MLLMs BenchmarksZicheng Zhang, Xiangyu Zhao, Xinyu Fang, Chunyi Li 等ACL 2025 · 被引用 15 次
- Learning Object State Changes in Videos: An Open-World PerspectiveZihui Xue, Kumar Ashutosh, Kristen GraumanCVPR 2024 · 被引用 12 次
相关 Paper
- AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video GenerationZiwei Zhou, Zeyuan Lai, Rui Wang, Yifan Yang 等ICML 2026 · 被引用 8 次
- MVBench: A Comprehensive Multi-modal Video Understanding BenchmarkKunchang Li, Yali Wang, Yinan He, Yizhuo Li 等CVPR 2024
- STATUS Bench: A Rigorous Benchmark for Evaluating Object State Understanding in Vision-Language ModelsMahiro Ukai, Shuhei Kurita, Nakamasa InoueACM MM 2025
- Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture AssemblyAditya Chetan, Eric Cai, Peeyush Kushwaha, Bharath Raj Nagoor Kani 等CVPR 2026
- MOSCATO: Predicting Multiple Object State Change through ActionsParnian Zameni, Yuhan Shen, Ehsan ElhamifarICCV 2025 · 被引用 4 次
