Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels
Haoning Wu, Zicheng Zhang, Weixia Zhang, Chaofeng Chen, Liang Liao, Chunyi Li, Yixuan Gao, Annan Wang, Erli Zhang, Wenxiu Sun, Qiong Yan, Xiongkuo Min
摘要
The explosion of visual content available online underscores the requirement for an accurate machine assessor to robustly evaluate scores across diverse types of visual contents. While recent studies have demonstrated the exceptional potentials of large multi-modality models (LMMs) on a wide range of related fields, in this work, we explore how to teach them for visual rating aligned with human opinions. Observing that human raters only learn and judge discrete text-defined levels in subjective studies, we propose to emulate this subjective process and teach LMMs with text-defined rating levels instead of scores. The proposed Q-ALIGN achieves state-of-the-art performance on image quality assessment (IQA), image aesthetic assessment (IAA), as well as video quality assessment (VQA) tasks under the original LMM structure. With the syllabus, we further unify the three tasks into one model, termed the ONEALIGN. In our experiments, we demonstrate the advantage of the discrete-level-based syllabus over direct-score-based variants for LMMs. Our code and the pre-trained weights are released at https://github.com/Q-Future/Q-Align .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper187
- Q-Insight: Understanding Image Quality via Visual Reinforcement LearningWeiqi Li, Xuanyu Zhang, Shijie Zhao, Yabin Zhang 等NeurIPS 2025 · 被引用 117 次
- Adaptive Image Quality Assessment via Teaching Large Multimodal Model to CompareHanwei Zhu, Haoning Wu, Yixuan Li, Zicheng Zhang 等NeurIPS 2024 · 被引用 108 次
- VisualQuality-R1: Reasoning-Induced Image Quality Assessment via Reinforcement Learning to RankTianhe Wu, Jian Zou, Jie Liang, Lei Zhang 等NeurIPS 2025 · 被引用 92 次
- Autoregressive Adversarial Post-Training for Real-Time Interactive Video GenerationShanchuan Lin, Ceyuan Yang, Hao He, Jianwen Jiang 等NeurIPS 2025 · 被引用 89 次
- Stochastic Self-Guidance for Training-Free Enhancement of Diffusion ModelsChubin Chen, Jiashu Zhu, Xiaokun Feng, Nisha Huang 等ICLR 2026 · 被引用 44 次
它引用的顶会 Paper13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
- Exploring CLIP for Assessing the Look and Feel of ImagesJianyi Wang, Kelvin C. K. Chan, Chen Change LoyAAAI 2023 · 被引用 1,208 次
- Exploring Video Quality Assessment on User Generated Contents from Aesthetic and Technical PerspectivesHaoning Wu, Erli Zhang, Liang Liao, Chaofeng Chen 等ICCV 2023 · 被引用 371 次
相关 Paper
- Revisiting MLLM Based Image Quality Assessment: Errors and RemedyZhenchen Tang, Songlin Yang, Bo Peng, Zichuan Wang 等AAAI 2026 · 被引用 2 次
- Teaching Large Language Models to Regress Accurate Image Quality Scores Using Score DistributionZhiyuan You, Xin Cai, Jinjin Gu, Tianfan Xue 等CVPR 2025
- LMM-PCQA: Assisting Point Cloud Quality Assessment with LMMZicheng Zhang, Haoning Wu, Yingjie Zhou, Chunyi Li 等ACM MM 2024 · 被引用 38 次
- CoFiVLA: Synergistic Coarse-Fine Vision-Language Alignment for Image Aesthetic AssessmentYuzhen Niu, Siling Chen, Yuzhong Chen, Fusheng Li 等ACM MM 2025 · 被引用 1 次
- Q-Eval-100K: Evaluating Visual Quality and Alignment Level for Text-to-Vision ContentZicheng Zhang, Tengchuan Kou, Shushi Wang, Chunyi Li 等CVPR 2025
