Beyond Cosine Similarity: Magnitude-Aware CLIP for No-Reference Image Quality Assessment
Zhicheng Liao, Dongxu Wu, Zhenshan Shi, Sijie Mai, Hanwei Zhu, Lingyu Zhu, Yuncheng Jiang, Baoliang Chen
Abstract
Recent efforts have repurposed the Contrastive Language-Image Pre-training (CLIP) model for No-Reference Image Quality Assessment (NR-IQA) by measuring the cosine similarity between the image embedding and textual prompts such as "a good photo" or "a bad photo." However, this semantic similarity overlooks a critical yet underexplored cue: the magnitude of the CLIP image features, which we empirically find to exhibit a strong correlation with perceptual quality. In this work, we introduce a novel adaptive fusion framework that complements cosine similarity with a magnitude-aware quality cue. Specifically, we first extract the absolute CLIP image features and apply a Box-Cox transformation to statistically normalize the feature distribution and mitigate semantic sensitivity. The resulting scalar summary serves as a semantically-normalized auxiliary cue that complements cosine-based prompt matching. To integrate both cues effectively, we further design a confidence-guided fusion scheme that adaptively weighs each term according to its relative strength. Extensive experiments on multiple benchmark IQA datasets demonstrate that our method consistently outperforms standard CLIP-based IQA and state-of-the-art baselines, without any task-specific training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8e7bb2a7-cda9-4bab-a6e7-46143c54b917Cited by top-tier papers2
- Bridging the Perceptual Gap: Residual-Enhanced Downscaling and Manifold-Aware Perception Alignment Adaptation for NR-IQAYu Li, Zhengran Shen, Yachun Mi, Puchao Zhou et al.ICML 2026
- Saving Foundation Flow-Matching Priors for Inverse ProblemsYuxiang Wan, Ryan Devera, Wenjie Zhang, Ju SunICML 2026
Builds on6
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Exploring CLIP for Assessing the Look and Feel of ImagesJianyi Wang, Kelvin C. K. Chan, Chen Change LoyAAAI 2023 · 1,208 citations
- Adaptive Image Quality Assessment via Teaching Large Multimodal Model to CompareHanwei Zhu, Haoning Wu, Yixuan Li, Zicheng Zhang et al.NeurIPS 2024 · 108 citations
- Re-IQA: Unsupervised Learning for Image Quality Assessment in the WildAvinab Saha, Sandeep Mishra, Alan C. BovikCVPR 2023
- MetaIQA: Deep Meta-Learning for No-Reference Image Quality AssessmentHancheng Zhu, Leida Li, Jinjian Wu, Weisheng Dong et al.CVPR 2020
Related papers
- CLIP-PCQA: Exploring Subjective-Aligned Vision-Language Modeling for Point Cloud Quality AssessmentYating Liu, Yujie Zhang, Ziyu Shan, Yiling XuAAAI 2025 · 9 citations
- CgT-GAN: CLIP-guided Text GAN for Image CaptioningJiarui Yu, Haoran Li, Yanbin Hao, Bin Zhu et al.ACM MM 2023 · 26 citations
- Few-Shot Image Quality Assessment via Adaptation of Vision-Language ModelsXudong Li, Zihao Huang, Yan Zhang, Yunhang Shen et al.ICCV 2025 · 2 citations
- ALIP: Adaptive Language-Image Pre-training with Synthetic CaptionKaicheng Yang, Jiankang Deng, Xiang An, Jiawei Li et al.ICCV 2023 · 93 citations
- Learning Where to Look and How to Judge: Resolution-agnostic Image Quality Assessment with Quality-aware SaliencyHakan Emre Gedik, Shashank Gupta, Alan BovikCVPR 2026 · 2 citations
