Scaling-up Perceptual Video Quality Assessment
Ziheng Jia, Zicheng Zhang, Xiaorong Zhu, Chunyi Li, Jinliang Han, Xiaohong Liu, Guangtao Zhai, Xiongkuo Min
Abstract
The data scaling law has significantly enhanced large multi-modal models (LMMs) performance across various downstream tasks. However, in the domain of perceptual video quality assessment (VQA), the potential of data scaling remains unprecedented due to the scarcity of labeled resources and the insufficient scale of datasets. To address this, we propose OmniVQA, a framework designed to efficiently build high-quality, machine-dominated synthetic multi-modal instruction databases (MIDBs) for VQA. We then scale up to create OmniVQA-Chat-400K, the largest dataset in the VQA field concurrently. Our focus is on the technical and aesthetic quality dimensions, with abundant in-context instruction data to provide fine-grained VQA knowledge. Additionally, we build the OmniVQA-MOS-20K dataset to enhance the model's quantitative quality rating capabilities. We then introduce a complementary training strategy that effectively leverages the knowledge from datasets for different tasks. Furthermore, we propose the OmniVQA-FG (fine-grain)-Benchmark to evaluate the fine-grained performance of models. Our results demonstrate that our models achieve state-of-the-art performance in both tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1b9f4bff-0ab2-4e16-acd2-5789f18223baCited by top-tier papers4
- Generalizable Video Quality Assessment via Weak-to-Strong LearningLinhan Cao, Wei Sun, Xiangyang Zhu, Kaiwei Zhang et al.CVPR 2026 · 9 citations
- VITAL: Vision-Encoder-centered Pre-training for LMMs in Visual Quality AssessmentZiheng Jia, Linhan Cao, Jinliang Han, Zicheng Zhang et al.CVPR 2026 · 1 citation
- LiViBench: An Omnimodal Benchmark for Interactive Livestream Video UnderstandingXiaodong Wang, Langling Huang, Zhirong Wu, Xu Zhao et al.AAAI 2026 · 1 citation
- VisualScore: Learning Holistic Visual Quality Scores via Multi-Task ReasoningYiting Lu, Fengbin Guan, Yixin Gao, Yan Zhong et al.ICML 2026
Builds on12
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Exploring Video Quality Assessment on User Generated Contents from Aesthetic and Technical PerspectivesHaoning Wu, Erli Zhang, Liang Liao, Chaofeng Chen et al.ICCV 2023 · 371 citations
- When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning MethodBiao Zhang, Zhongtao Liu, Colin Cherry, Orhan FiratICLR 2024 · 271 citations
- A Deep Learning based No-reference Quality Assessment Model for UGC VideosWei Sun, Xiongkuo Min, Wei Lu, Guangtao ZhaiACM MM 2022 · 239 citations
- Towards Explainable In-the-Wild Video Quality Assessment: A Database and a Language-Prompted ApproachHaoning Wu, Erli Zhang, Liang Liao, Chaofeng Chen et al.ACM MM 2023 · 51 citations
Related papers
- AIGV-Assessor: Benchmarking and Evaluating the Perceptual Quality of Text-to-Video Generation with LMMJiarui Wang, Huiyu Duan, Guangtao Zhai, Juntong Wang et al.CVPR 2025
- Score2Instruct: Scaling Up Video Quality-Centric Instructions via Automated Dimension ScoringQizhi Xie, Kun Yuan, Yunpeng Qu, Jiachao Gong et al.CVPR 2026
- Subjective-Aligned Dataset and Metric for Text-to-Video Quality AssessmentTengchuan Kou, Xiaohong Liu, Zicheng Zhang, Chunyi Li et al.ACM MM 2024 · 29 citations
- VQA2: Visual Question Answering for Video Quality AssessmentZiheng Jia, Zicheng Zhang, Jiaying Qian, Haoning Wu et al.ACM MM 2025 · 13 citations
- Distilling Vision-Language Models on Millions of VideosYue Zhao, Long Zhao, Xingyi Zhou, Jialin Wu et al.CVPR 2024
