Q-Ground: Image Quality Grounding with Large Multi-modality Models
Chaofeng Chen, Sensen Yang, Haoning Wu, Liang Liao, Zicheng Zhang, Annan Wang, Wenxiu Sun, Qiong Yan, Weisi Lin
摘要
Recent advances of large multi-modality models (LMM) have greatly improved the ability of image quality assessment (IQA) method to evaluate and explain the quality of visual content. However, these advancements are mostly focused on overall quality assessment, and the detailed examination of local quality, which is crucial for comprehensive visual understanding, is still largely unexplored. In this work, we introduce Q-Ground, the first framework aimed at tackling fine-scale visual quality grounding by combining large multi-modality models with detailed visual quality analysis. Cen- tral to our contribution is the introduction of the QGround-100K dataset, a novel resource containing 100k triplets of (image, quality text, distortion segmentation) to facilitate deep investigations into visual quality. The dataset comprises two parts: one with human- labeled annotations for accurate quality assessment, and another la- beled automatically by LMMs such as GPT4V, which helps improve the robustness of model training while also reducing the costs of data collection. With the QGround-100K dataset, we propose a LMM-based method equipped with multi-scale feature learning to learn models capable of performing both image quality answer- ing and distortion segmentation based on text prompts. This dual- capability approach not only refines the model’s understanding of region-aware image quality but also enables it to interactively re- spond to complex, text-based queries about image quality and spe- cific distortions. Q-Ground takes a step towards sophisticated vi- sual quality analysis in a finer scale, establishing a new benchmark for future research in the area. Codes and dataset are available at https://github.com/Q-Future/Q-Ground.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Q-Insight: Understanding Image Quality via Visual Reinforcement LearningWeiqi Li, Xuanyu Zhang, Shijie Zhao, Yabin Zhang 等NeurIPS 2025 · 被引用 117 次
- Reasoning as Representation: Rethinking Visual Reinforcement Learning in Image Quality AssessmentShijie Zhao, Xuanyu Zhang, Weiqi Li, Junlin Li 等ICLR 2026 · 被引用 20 次
- Grounding-IQA: Grounding Multimodal Language Model for Image Quality AssessmentZheng Chen, Xun Zhang, Wenbo Li, Renjing Pei 等ICLR 2026 · 被引用 12 次
- UARE: A Unified Vision-Language Model for Image Quality Assessment, Restoration, and EnhancementWeiqi Li, Xuanyu Zhang, Bin Chen, Jingfen Xie 等CVPR 2026 · 被引用 5 次
- Mitigating Perception Bias: A Training-Free Approach to Enhance LMM for Image Quality AssessmentBaoliang Chen, Siyi Pan, Dongxu Wu, Liang Xie 等AAAI 2026 · 被引用 5 次
它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- MUSIQ: Multi-scale Image Quality TransformerJunjie Ke, Qifei Wang, Yilin Wang, Peyman Milanfar 等ICCV 2021 · 被引用 1,325 次
- ImageReward: Learning and Evaluating Human Preferences for Text-to-Image GenerationJiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong 等NeurIPS 2023 · 被引用 1,310 次
- Exploring CLIP for Assessing the Look and Feel of ImagesJianyi Wang, Kelvin C. K. Chan, Chen Change LoyAAAI 2023 · 被引用 1,208 次
相关 Paper
- VQA2: Visual Question Answering for Video Quality AssessmentZiheng Jia, Zicheng Zhang, Jiaying Qian, Haoning Wu 等ACM MM 2025 · 被引用 13 次
- VideoGLaMM : A Large Multimodal Model for Pixel-Level Visual Grounding in VideosShehan Munasinghe, Hanan Gani, Wenqi Zhu, Jiale Cao 等CVPR 2025
- Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level VisionHaoning Wu, Zicheng Zhang, Erli Zhang, Chaofeng Chen 等ICLR 2024 · 被引用 258 次
- IQA-Spider: Unifying Multi-Granularity Image Quality Assessment with Reasoning, Grounding and ReferringXinge Peng, Yiting Lu, Xin Li, Zhibo ChenICML 2026
- Score2Instruct: Scaling Up Video Quality-Centric Instructions via Automated Dimension ScoringQizhi Xie, Kun Yuan, Yunpeng Qu, Jiachao Gong 等CVPR 2026
