Ponder: Point Cloud Pre-training via Neural Rendering
Di Huang, Sida Peng, Tong He, Honghui Yang, Xiaowei Zhou, Wanli Ouyang
Abstract
We propose a novel approach to self-supervised learning of point cloud representations by differentiable neural rendering. Motivated by the fact that informative point cloud features should be able to encode rich geometry and appearance cues and render realistic images, we train a point-cloud encoder within a devised point-based neural renderer by comparing the rendered images with real images on massive RGB-D data. The learned point-cloud encoder can be easily integrated into various downstream tasks, including not only high-level tasks like 3D detection and segmentation but also low-level tasks like 3D reconstruction and image synthesis. Extensive experiments on various tasks demonstrate the superiority of our approach compared to existing pre-training methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers13
- UniPAD: A Universal Pre-Training Paradigm for Autonomous DrivingHonghui Yang, Sha Zhang, Di Huang, Xiaoyang Wu et al.CVPR 2024 · 31 citations
- Point-GCC: Universal Self-supervised 3D Scene Pre-training via Geometry-Color ContrastGuofan Fan, Zekun Qi, Wenkai Shi, Kaisheng MaACM MM 2024 · 12 citations
- PRED: Pre-training via Semantic Rendering on LiDAR Point CloudsHao Yang, Haiyang Wang, Di Dai, Liwei WangNeurIPS 2023 · 10 citations
- TREND: Unsupervised 3D Representation Learning via Temporal Forecasting for LiDAR PerceptionRunjian Chen, Hyoungseob Park, Bo Zhang, Wenqi Shao et al.NeurIPS 2025 · 4 citations
- DynaRend: Learning 3D Dynamics via Masked Future Rendering for Robotic ManipulationJingyi Tian, Le Wang, Sanping Zhou, Sen Wang et al.NeurIPS 2025 · 3 citations
Builds on34
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 4,089 citations
- NeuS: Learning Neural Implicit Surfaces by Volume Rendering for Multi-view ReconstructionPeng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt et al.NeurIPS 2021 · 2,500 citations
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 2,336 citations
Related papers
- Self-Supervised Pretraining for Large-Scale Point CloudsZaiwei Zhang, Min Bai, Li Erran LiNeurIPS 2022 · 12 citations
- Self-Supervised Learning on 3D Point Clouds by Learning Discrete Generative ModelsBenjamin Eckart, Wentao Yuan, Chao Liu, Jan KautzCVPR 2021
- Harnessing Text-to-Image Diffusion Models for Point Cloud Self-Supervised LearningYiyang Chen, Shanshan Zhao, Lunhao Duan, Changxing Ding et al.ICCV 2025
- ToThePoint: Efficient Contrastive Learning of 3D Point Clouds via RecyclingXinglin Li, Jiajing Chen, Jinhui Ouyang, Hanhui Deng et al.CVPR 2023
- ALSO: Automotive Lidar Self-Supervision by Occupancy EstimationAlexandre Boulch, Corentin Sautier, Björn Michele, Gilles Puy et al.CVPR 2023
