Neuro-Symbolic Evaluation of Text-to-Video Models using Formal Verification
S. P. Sharan, Minkyu Choi, Sahil Shah, Harsh Goel, Mohammad Omama, Sandeep Chinchali
Abstract
Recent advancements in text-to-video models such as Sora, Gen-3, MovieGen, and CogVideoX are pushing the boundaries of synthetic video generation, with adoption seen in fields like robotics, autonomous driving, and entertainment. As these models become prevalent, various metrics and benchmarks have emerged to evaluate the quality of the generated videos. However, these metrics emphasize visual quality and smoothness, neglecting temporal fidelity and text-to-video alignment, which are crucial for safety-critical applications. To address this gap, we introduce NeuS-V, a novel synthetic video evaluation metric that rigorously assesses text-to-video alignment using neuro-symbolic formal verification techniques. Our approach first converts the prompt into a formally defined Temporal Logic (TL) specification and translates the generated video into an automaton representation. Then, it evaluates the text-to-video alignment by formally checking the video automaton against the TL specification. Furthermore, we present a dataset of temporally extended prompts to evaluate state-of-the-art video generation models against our benchmark. We find that NeuS-V demonstrates a higher correlation by over 5× with human evaluations when compared to existing metrics. Our evaluation further reveals that current video generation models perform poorly on these temporally complex prompts, highlighting the need for future work in improving text-to-video generation capabilities. We open-source our benchmark, code, and dataset at utaustin-swarmlab.github.io/neusv.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6e688d1e-6aa5-4857-b867-8a8c3a66099cCited by top-tier papers2
- NeuS-QA: Grounding Long-Form Video Understanding in Temporal Logic and Neuro-Symbolic ReasoningSahil Shah, S. P. Sharan, Harsh Goel, Minkyu Choi et al.AAAI 2026 · 4 citations
- Ref4D-VideoBench: Four-Dimensional Reference-Based Evaluation of Text-to-Video Generative ModelsJiajia Wei, YuJia He, Yuhan Hou, Hang Qi et al.CVPR 2026
Builds on17
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Structure and Content-Guided Video Synthesis with Diffusion ModelsPatrick Esser, Johnathan Chiu, Parmida Atighehchian, Jonathan Granskog et al.ICCV 2023 · 733 citations
- Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language ModelsMuhammad Maaz, Hanoona Abdul Rasheed, Salman Khan, Fahad KhanACL 2024 · 279 citations
- T2V-Turbo: Breaking the Quality Bottleneck of Video Consistency Model with Mixed Reward FeedbackJiachen Li, Weixi Feng, Tsu-Jui Fu, Xinyi Wang et al.NeurIPS 2024 · 97 citations
Related papers
- EvalCrafter: Benchmarking and Evaluating Large Video Generation ModelsYaofang Liu, Xiaodong Cun, Xuebo Liu, Xintao Wang et al.CVPR 2024
- LoCoT2V-Bench: Benchmarking Long-Form and Complex Text-to-Video GenerationXiangqing Zheng, CHENGYUE WU, Kehai Chen, Min zhangICML 2026 · 3 citations
- Evaluation of Text-to-Video Generation Models: A Dynamics PerspectiveMingxiang Liao, Hannan Lu, Qixiang Ye, Wangmeng Zuo et al.NeurIPS 2024 · 89 citations
- MoReGen: Multi-Agent Motion-Reasoning Engine for Code-based Text-to-Video SynthesisXiangyu Bai, He Liang, Bishoy Galoaa, Utsav Nandi et al.CVPR 2026 · 6 citations
- Subjective-Aligned Dataset and Metric for Text-to-Video Quality AssessmentTengchuan Kou, Xiaohong Liu, Zicheng Zhang, Chunyi Li et al.ACM MM 2024 · 29 citations
