CLIP-PCQA: Exploring Subjective-Aligned Vision-Language Modeling for Point Cloud Quality Assessment
Yating Liu, Yujie Zhang, Ziyu Shan, Yiling Xu
Abstract
In recent years, No-Reference Point Cloud Quality Assessment (NR-PCQA) research has achieved significant progress. However, existing methods mostly seek a direct mapping function from visual data to the Mean Opinion Score (MOS), which is contradictory to the mechanism of practical subjective evaluation. To address this, we propose a novel language-driven PCQA method named CLIP-PCQA. Considering that human beings prefer to describe visual quality using discrete quality descriptions (e.g., "excellent" and "poor") rather than specific scores, we adopt a retrieval-based mapping strategy to simulate the process of subjective assessment. More specifically, based on the philosophy of CLIP, we calculate the cosine similarity between the visual features and multiple textual features corresponding to different quality descriptions, in which process an effective contrastive loss and learnable prompts are introduced to enhance the feature extraction. Meanwhile, given the personal limitations and bias in subjective experiments, we further covert the feature similarities into probabilities and consider the Opinion Score Distribution (OSD) rather than a single MOS as the final target. Experimental results show that our CLIP-PCQA outperforms other State-Of-The-Art (SOTA) approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b6cde124-8de3-42b8-b19e-e282badfe197Cited by top-tier papers3
- Perceptual Quality Assessment of 3D Gaussian Splatting: A Subjective Dataset and Prediction MetricZhaolin Wan, Yining Diao, Jingqi Xu, Hao Wang et al.AAAI 2026 · 1 citation
- DSP-PCQA: Integrating Multiple Perception Preferences for Point Cloud Quality AssessmentMingxuan Li, Fazhan Zhang, Zhenzhe Hou, Zihao Huang et al.AAAI 2026
- Points Meet Pixels: Bridging 2D Vision-Language Model and 3D Perception Gaps for Point Cloud Quality AssessmentMingxuan Li, Zihao Huang, Xiaohui Chu, Fazhan Zhang et al.AAAI 2026
Builds on7
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 1,274 citations
- Exploring CLIP for Assessing the Look and Feel of ImagesJianyi Wang, Kelvin C. K. Chan, Chen Change LoyAAAI 2023 · 1,208 citations
- Image Quality Assessment: From Mean Opinion Score to Opinion Score DistributionYixuan Gao, Xiongkuo Min, Yucheng Zhu, Jing Li et al.ACM MM 2022 · 53 citations
Related papers
- CLiF-VQA: Enhancing Video Quality Assessment by Incorporating High-Level Semantic Information related to Human FeelingsYachun Mi, Yan Shu, Yu Li, Chen Hui et al.ACM MM 2024 · 6 citations
- Contrastive Pre-Training with Multi-View Fusion for No-Reference Point Cloud Quality AssessmentZiyu Shan, Yujie Zhang, Qi Yang, Haichen Yang et al.CVPR 2024 · 21 citations
- Beyond Cosine Similarity: Magnitude-Aware CLIP for No-Reference Image Quality AssessmentZhicheng Liao, Dongxu Wu, Zhenshan Shi, Sijie Mai et al.AAAI 2026 · 3 citations
- Open-Set Fine-Grained Retrieval via Prompting Vision-Language EvaluatorShijie Wang, Jianlong Chang, Haojie Li, Zhihui Wang et al.CVPR 2023
- No-Reference Point Cloud Quality Assessment via Domain AdaptationQi Yang, Yipeng Liu, Siheng Chen, Yiling Xu et al.CVPR 2022 · 108 citations
