Points Meet Pixels: Bridging 2D Vision-Language Model and 3D Perception Gaps for Point Cloud Quality Assessment
Mingxuan Li, Zihao Huang, Xiaohui Chu, Fazhan Zhang, Bohan Fu, Runze Hu
摘要
Vision-Language Models (VLMs) have demonstrated significant progress in quality assessment tasks. However, a fundamental paradox arises when their application to Point Cloud Quality Assessment (PCQA). Existing VLMs, designed for image-text pairs, are inherently incompatible with 3D point cloud data due to the modality gap. While some PCQA research attempts to adapt point clouds to VLMs by 2D projection, this approach inevitably sacrifices crucial spatial structure information essential for accurate quality assessment. Conversely, directly integrating a dedicated 3D branch into a VLM-based PCQA framework introduces feature space misalignment and an influx of quality-insensitive information. To bridge these fundamental conflicts hindering VLMs' adaptation to PCQA, we propose the PMP-PCQA framework, which leverages the inherent mapping relationship between points and pixels to seamlessly apply VLMs to PCQA. Our approach introduces three key innovations: a Spatial Awareness Enhancer(SAE) module that enriches the image features with spatial coordinate clues to reinforce geometric awareness in 2D visual representations; a Fine-to-coarse Consistency Alignment(FCA) module that bridges the gap between 2D and 3D modalities by leveraging point-pixel correspondences to construct bridging features; and a Text-Guided Adaptive Miner(TAM) module that dynamically suppresses quality-insensitive features to mine discriminative visual clues for PCQA. Extensive evaluations demonstrate that PMP-PCQA consistently outperforms state-of-the-art methods across multiple benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Exploring CLIP for Assessing the Look and Feel of ImagesJianyi Wang, Kelvin C. K. Chan, Chen Change LoyAAAI 2023 · 被引用 1,208 次
- Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined LevelsHaoning Wu, Zicheng Zhang, Weixia Zhang, Chaofeng Chen 等ICML 2024 · 被引用 499 次
- Towards Explainable In-the-Wild Video Quality Assessment: A Database and a Language-Prompted ApproachHaoning Wu, Erli Zhang, Liang Liao, Chaofeng Chen 等ACM MM 2023 · 被引用 51 次
- LMM-PCQA: Assisting Point Cloud Quality Assessment with LMMZicheng Zhang, Haoning Wu, Yingjie Zhou, Chunyi Li 等ACM MM 2024 · 被引用 38 次
相关 Paper
- MT-DPCQA: A Multimodal Time-aware Learning Approach for No-Reference Dynamic Point Cloud Quality AssessmentSwarna Chakraborty, Mylène C. Q. FariasACM MM 2025 · 被引用 2 次
- R3-PCQA: Ray-Reprojection-Reinforcement for No-Reference 3D Point Cloud Quality AssessmentJunhyuk Seo, Sanghyuk SEO, Dawoon Kim, Heeseok OhCVPR 2026
- Adaptive Prompt Learning for Blind Image Quality Assessment with Multi-modal Mixed-datasets TrainingYan Zhong, Xinping Zhao, Li Zhang, Xinyuan Song 等ACM MM 2025
- Point Cloud Projection and Multi-Scale Feature Fusion Network Based Blind Quality Assessment for Colored Point CloudsWenxu Tao, Gangyi Jiang, Zhidi Jiang, Mei YuACM MM 2021 · 被引用 51 次
- PointAlign: Feature-Level Alignment Regularization for 3D Vision-Language ModelsYuanhao Su, Shaofeng Zhang, Xiaosong Jia, Qi FanCVPR 2026 · 被引用 1 次
