CLIP-PCQA: Exploring Subjective-Aligned Vision-Language Modeling for Point Cloud Quality Assessment
Yating Liu, Yujie Zhang, Ziyu Shan, Yiling Xu
摘要
In recent years, No-Reference Point Cloud Quality Assessment (NR-PCQA) research has achieved significant progress. However, existing methods mostly seek a direct mapping function from visual data to the Mean Opinion Score (MOS), which is contradictory to the mechanism of practical subjective evaluation. To address this, we propose a novel language-driven PCQA method named CLIP-PCQA. Considering that human beings prefer to describe visual quality using discrete quality descriptions (e.g., "excellent" and "poor") rather than specific scores, we adopt a retrieval-based mapping strategy to simulate the process of subjective assessment. More specifically, based on the philosophy of CLIP, we calculate the cosine similarity between the visual features and multiple textual features corresponding to different quality descriptions, in which process an effective contrastive loss and learnable prompts are introduced to enhance the feature extraction. Meanwhile, given the personal limitations and bias in subjective experiments, we further covert the feature similarities into probabilities and consider the Opinion Score Distribution (OSD) rather than a single MOS as the final target. Experimental results show that our CLIP-PCQA outperforms other State-Of-The-Art (SOTA) approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Perceptual Quality Assessment of 3D Gaussian Splatting: A Subjective Dataset and Prediction MetricZhaolin Wan, Yining Diao, Jingqi Xu, Hao Wang 等AAAI 2026 · 被引用 1 次
- DSP-PCQA: Integrating Multiple Perception Preferences for Point Cloud Quality AssessmentMingxuan Li, Fazhan Zhang, Zhenzhe Hou, Zihao Huang 等AAAI 2026
- Points Meet Pixels: Bridging 2D Vision-Language Model and 3D Perception Gaps for Point Cloud Quality AssessmentMingxuan Li, Zihao Huang, Xiaohui Chu, Fazhan Zhang 等AAAI 2026
它引用的顶会 Paper7
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 被引用 1,274 次
- Exploring CLIP for Assessing the Look and Feel of ImagesJianyi Wang, Kelvin C. K. Chan, Chen Change LoyAAAI 2023 · 被引用 1,208 次
- Image Quality Assessment: From Mean Opinion Score to Opinion Score DistributionYixuan Gao, Xiongkuo Min, Yucheng Zhu, Jing Li 等ACM MM 2022 · 被引用 53 次
相关 Paper
- CLiF-VQA: Enhancing Video Quality Assessment by Incorporating High-Level Semantic Information related to Human FeelingsYachun Mi, Yan Shu, Yu Li, Chen Hui 等ACM MM 2024 · 被引用 6 次
- Contrastive Pre-Training with Multi-View Fusion for No-Reference Point Cloud Quality AssessmentZiyu Shan, Yujie Zhang, Qi Yang, Haichen Yang 等CVPR 2024 · 被引用 21 次
- Beyond Cosine Similarity: Magnitude-Aware CLIP for No-Reference Image Quality AssessmentZhicheng Liao, Dongxu Wu, Zhenshan Shi, Sijie Mai 等AAAI 2026 · 被引用 3 次
- Open-Set Fine-Grained Retrieval via Prompting Vision-Language EvaluatorShijie Wang, Jianlong Chang, Haojie Li, Zhihui Wang 等CVPR 2023
- No-Reference Point Cloud Quality Assessment via Domain AdaptationQi Yang, Yipeng Liu, Siheng Chen, Yiling Xu 等CVPR 2022 · 被引用 108 次
