Q-Bench-Video: Benchmark the Video Quality Understanding of LMMs
Zicheng Zhang, Ziheng Jia, Haoning Wu, Chunyi Li, Zijian Chen, Yingjie Zhou, Wei Sun, Xiaohong Liu, Xiongkuo Min, Weisi Lin, Guangtao Zhai
Abstract
With the rising interest in research on Large Multi-modal Models (LMMs) for video understanding, many studies have emphasized general video comprehension capabilities, neglecting the systematic exploration into video quality understanding. To address this oversight, we introduce Q-Bench-Video in this paper, a new benchmark specifically designed to evaluate LMMs' proficiency in discerning video quality. a) To ensure video source diversity, Q-Bench-Video encompasses videos from natural scenes, AIgenerated content (AIGC), and computer graphics (CG). b) Building on the traditional multiple-choice questions format with the Yes-or-No and What-How categories, we include Open-ended questions to better evaluate complex scenarios. Additionally, we incorporate the video pair quality comparison question to enhance comprehensiveness. c) Beyond the traditional Technical, Aesthetic, and Temporal distortions, we have expanded our evaluation aspects to include the dimension of AIGC distortions, which addresses the increasing demand for video generation. Finally, we collect a total of 2,378 question-answer pairs and test them on 12 open-source & 5 proprietary LMMs. Our findings indicate that while LMMs have a foundational understanding of perceptual video quality, their performance remains incomplete and imprecise, with a notable discrepancy compared to the performance of human beings. Through Q-Bench-Video, we seek to catalyze community interest, stimulate further research, and unlock the untapped potential of LMMs to close the gap in video quality understanding. Q: Does the fabric in this video exhibit good clarity and proper lighting? A. Yes B. No (a) Question Type Yes-or-No What-How Open-ended (b) Quality Concern (c) Single Video vs. Video Pairs Q: What quality issues do not exist in this video? A. None of the options B. Blurriness C. Overexposure D. Underexposure Q: Why is it difficult for viewers to identify the dish the man is cooking in the video? Open-ended Response: The video has severe compression blur and block artifacts, significantly reducing the discernibility of objects in the video. Q: What feelings does this video evoke? Open-ended Response: This video depicts a castle in a unnatural and twisted jungle, where the plants have bizarre, sharp structures and dull colors. The atmosphere is eerie and terrifying, creating a sense of horror. Technical Q: As the camera moves away in this video, is there a noticeable increase in the clarity of the person's face? A. No B. Yes AIGC Global Referring Temporal Aesthetic Q: Does this video have severe camera shake? A. No B. Yes Q: What is the most impactful quality issue of this video? A. Noise B. Incorrect human structure C. Blurriness D. Overexposure Q: How is the overall lighting level in this game video? A. Very poor B. Poor C. Average D. Good Q: Is the main character in this game video rendered in high details but with relatively low clarity? A. Yes B. No Joint Compare Q: Is the exposure of the first video more balanced than the second video? A. No B. Yes
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2dc613b9-f0e0-42e8-aef7-00d875271d37Cited by top-tier papers4
- Seeing Beyond 8bits: Subjective and Objective Quality Assessment of HDR-UGC VideosShreshth Saini, Bowen Chen, Yilin Wang, Neil Birkbeck et al.CVPR 2026 · 1 citation
- MetaSpace: Metamorphic Testing for Spatial Cognition in Embodied AgentsGengyang Xu, Dongwei Xiao, Yiteng Peng, Shuai WangOOPSLA 2026
- Score2Instruct: Scaling Up Video Quality-Centric Instructions via Automated Dimension ScoringQizhi Xie, Kun Yuan, Yunpeng Qu, Jiachao Gong et al.CVPR 2026
- Ref4D-VideoBench: Four-Dimensional Reference-Based Evaluation of Text-to-Video Generative ModelsJiajia Wei, YuJia He, Yuhan Hou, Hang Qi et al.CVPR 2026
Builds on18
- Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined LevelsHaoning Wu, Zicheng Zhang, Weixia Zhang, Chaofeng Chen et al.ICML 2024 · 499 citations
- Exploring Video Quality Assessment on User Generated Contents from Aesthetic and Technical PerspectivesHaoning Wu, Erli Zhang, Liang Liao, Chaofeng Chen et al.ICCV 2023 · 371 citations
- Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level VisionHaoning Wu, Zicheng Zhang, Erli Zhang, Chaofeng Chen et al.ICLR 2024 · 258 citations
- A Deep Learning based No-reference Quality Assessment Model for UGC VideosWei Sun, Xiongkuo Min, Wei Lu, Guangtao ZhaiACM MM 2022 · 239 citations
- Video-LLaVA: Learning United Visual Representation by Alignment Before ProjectionBin Lin, Yang Ye, Bin Zhu, Jiaxi Cui et al.EMNLP 2024 · 231 citations
Related papers
- CG-Bench: Clue-grounded Question Answering Benchmark for Long Video UnderstandingGuo Chen, Yicheng Liu, Yifei Huang, Baoqi Pei et al.ICLR 2025
- AIGV-Assessor: Benchmarking and Evaluating the Perceptual Quality of Text-to-Video Generation with LMMJiarui Wang, Huiyu Duan, Guangtao Zhai, Juntong Wang et al.CVPR 2025
- VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC VideosTingyu Song, Tongyan Hu, Guo Gan, Yilun ZhaoACL 2025 · 1 citation
- VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality EvaluationLongteng Jiang, Dandan Zheng, Qianqian Qiao, Heng Huang et al.CVPR 2026 · 2 citations
- MESH - Understanding Videos Like Human: Measuring Hallucinations in Large Video ModelsGarry Yang, Zizhe Chen, Man Hon Wong, Haoyu Lei et al.ACM MM 2025 · 1 citation
