Learning Disentangled Representations for Perceptual Point Cloud Quality Assessment via Mutual Information Minimization
Ziyu Shan, Yujie Zhang, Yipeng Liu, Yiling Xu
Abstract
No-Reference Point Cloud Quality Assessment (NR-PCQA) aims to objectively assess the human perceptual quality of point clouds without relying on pristine-quality point clouds for reference. It is becoming increasingly significant with the rapid advancement of immersive media applications such as virtual reality (VR) and augmented reality (AR). However, current NR-PCQA models attempt to indiscriminately learn point cloud content and distortion representations within a single network, overlooking their distinct contributions to quality information. To address this issue, we propose DisPA, a novel disentangled representation learning framework for NR-PCQA. The framework trains a dual-branch disentanglement network to minimize mutual information (MI) between representations of point cloud content and distortion. Specifically, to fully disentangle representations, the two branches adopt different philosophies: the content-aware encoder is pretrained by a masked auto-encoding strategy, which can allow the encoder to capture semantic information from rendered images of distorted point clouds; the distortion-aware encoder takes a mini-patch map as input, which forces the encoder to focus on low-level distortion patterns. Furthermore, we utilize an MI estimator to estimate the tight upper bound of the actual MI and further minimize it to achieve explicit representation disentanglement. Extensive experimental results demonstrate that DisPA outperforms state-of-the-art methods on multiple PCQA datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on18
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- CLUB: A Contrastive Log-ratio Upper Bound of Mutual InformationPengyu Cheng, Weituo Hao, Shuyang Dai, Jiachang Liu et al.ICML 2020 · 512 citations
- Exploring Video Quality Assessment on User Generated Contents from Aesthetic and Technical PerspectivesHaoning Wu, Erli Zhang, Liang Liao, Chaofeng Chen et al.ICCV 2023 · 371 citations
- PointGPT: Auto-regressively Generative Pre-training from Point CloudsGuangyan Chen, Meiling Wang, Yi Yang, Kai Yu et al.NeurIPS 2023 · 219 citations
Related papers
- Contrastive Pre-Training with Multi-View Fusion for No-Reference Point Cloud Quality AssessmentZiyu Shan, Yujie Zhang, Qi Yang, Haichen Yang et al.CVPR 2024 · 21 citations
- MT-DPCQA: A Multimodal Time-aware Learning Approach for No-Reference Dynamic Point Cloud Quality AssessmentSwarna Chakraborty, Mylène C. Q. FariasACM MM 2025 · 2 citations
- R3-PCQA: Ray-Reprojection-Reinforcement for No-Reference 3D Point Cloud Quality AssessmentJunhyuk Seo, Sanghyuk SEO, Dawoon Kim, Heeseok OhCVPR 2026
- QD-PCQA: Quality-Aware Domain Adaptation for Point Cloud Quality AssessmentGuohua Zhang, Jian Jin, Meiqin Liu, Chao Yao et al.CVPR 2026 · 3 citations
- No-Reference Point Cloud Quality Assessment via Domain AdaptationQi Yang, Yipeng Liu, Siheng Chen, Yiling Xu et al.CVPR 2022 · 108 citations
