MM-3DScene: 3D Scene Understanding by Customizing Masked Modeling with Informative-Preserved Reconstruction and Self-Distilled Consistency
Mingye Xu, Mutian Xu, Tong He, Wanli Ouyang, Yali Wang, Xiaoguang Han, Yu Qiao
Abstract
mingyexu.github.io/mm3dscene Figure 1 . How to apply masked modeling for large-scale 3D scenes? (a) Conventional random masked modeling on 3D scenes may cause a high risk of uncertainty.In this figure, a chair and a TV are totally masked, which are extremely difficult to be recovered without any context guidance. (b) Our MM-3DScene exploits local statistics to discover and preserve representative structured points, effectively simplifying the pretext task. At each learning step, our method focuses on restoring regional geometry, and enjoys less ambiguity. Moreover, since unmasked areas are underexplored during reconstruction, the model is encouraged to maintain the intrinsic spatial consistency on unmasked points between different masking ratios, which requires the consistent understanding of unmasked areas.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a2c680d4-48d4-4d46-aebb-a70e30e614dcCited by top-tier papers8
- Segment Any Point Cloud Sequences by Distilling Vision Foundation ModelsYouquan Liu, Lingdong Kong, Jun Cen, Runnan Chen et al.NeurIPS 2023 · 169 citations
- A Unified Framework for 3D Scene UnderstandingWei Xu, Chunsheng Shi, Sifan Tu, Xin Zhou et al.NeurIPS 2024 · 25 citations
- Fine-grained Image-to-LiDAR Contrastive Distillation with Visual Foundation ModelsYifan Zhang, Junhui HouNeurIPS 2024 · 9 citations
- Multi-View Representation is What You Need for Point-Cloud Pre-TrainingSiming Yan, Chen Song, Youkang Kong, Qixing HuangICLR 2024 · 6 citations
- PointCSP: Cross-Sample Semantic Propagation and Stability Preservation in Self-Supervised Point Cloud LearningXinxing Yu, Ajian Liu, Sunyuan Qiang, Hui Ma et al.CVPR 2026 · 1 citation
Builds on31
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- KPConv: Flexible and Deformable Convolution for Point CloudsHugues Thomas, Charles R. Qi, Jean-Emmanuel Deschaud, Beatriz Marcotegui et al.ICCV 2019 · 3,193 citations
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
Related papers
- DiffRF: Rendering-Guided 3D Radiance Field DiffusionNorman Müller, Yawar Siddiqui, Lorenzo Porzi, Samuel Rota Bulò et al.CVPR 2023
- 3D Mesh Editing Using Masked LRMsWill Gao, Dilin Wang, Yuchen Fan, Aljaz Bozic et al.ICCV 2025 · 6 citations
- Self-Supervised Pre-Training with Masked Shape Prediction for 3D Scene UnderstandingLi Jiang, Zetong Yang, Shaoshuai Shi, Vladislav Golyanik et al.CVPR 2023
- Clutter Detection and Removal in 3D Scenes with View-Consistent InpaintingFangyin Wei, Thomas A. Funkhouser, Szymon RusinkiewiczICCV 2023 · 10 citations
- CPCM: Contextual Point Cloud Modeling for Weakly-supervised Point Cloud Semantic SegmentationLizhao Liu, Zhuangwei Zhuang, Shangxin Huang, Xunlong Xiao et al.ICCV 2023 · 31 citations
