Masked Scene Modeling: Narrowing the Gap Between Supervised and Self-Supervised Learning in 3D Scene Understanding
Pedro Hermosilla, Christian Stippel, Leon Sick
Abstract
Figure 1 . Self-Supervised Feature Visualization using PCA. We reduce the point features obtained with our self-supervised model to three dimensions using PCA and visualize them as colors. Features learned by our model are semantic-aware, which is visible from the color separation: Similar objects result in similar features, such as the sofas in the first figure or the chairs in the last one, while different objects result in different features, such as the counter and the tables in the second image or the crib and the curtains in the third one.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ac988213-e6b0-44cc-b970-9014a3e0bcdcCited by top-tier papers4
- Towards Foundation Models for 3D Scene Understanding: Instance-Aware Self-Supervised Learning for Point CloudsBin Yang, Mohamed Abdelsamad, Miao Zhang, Alexandru Paul ConduracheCVPR 2026 · 4 citations
- PointCSP: Cross-Sample Semantic Propagation and Stability Preservation in Self-Supervised Point Cloud LearningXinxing Yu, Ajian Liu, Sunyuan Qiang, Hui Ma et al.CVPR 2026 · 1 citation
- DOS: Distilling Observable Softmaps of Zipfian Prototypes for Self-Supervised Point RepresentationMohamed Abdelsamad, Michael Ulrich, Bin Yang, Miao Zhang et al.AAAI 2026 · 1 citation
- Masked Representation Modeling for Domain-Adaptive SegmentationWenlve Zhou, Zhiheng Zhou, Tiantao Xian, Yikui Zhai et al.CVPR 2026
Builds on28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- SimMIM: a Simple Framework for Masked Image ModelingZhenda Xie, Zheng Zhang, Yue Cao, Yutong Lin et al.CVPR 2022 · 1,129 citations
Related papers
- 4D LangSplat: 4D Language Gaussian Splatting via Multimodal Large Language ModelsWanhua Li, Renping Zhou, Jiawei Zhou, Yingwei Song et al.CVPR 2025
- Self-Supervised Learning of Depth Inference for Multi-View StereoJiayu Yang, José M. Álvarez, Miaomiao LiuCVPR 2021
- Augmentation Component Analysis: Modeling Similarity via the Augmentation OverlapsLu Han, Han-Jia Ye, De-Chuan ZhanICLR 2023
- Unsupervised visualization of image datasets using contrastive learningJan Niklas Böhm, Philipp Berens, Dmitry KobakICLR 2023 · 6 citations
- Discovering Universal Geometry in Embeddings with ICAHiroaki Yamagiwa, Momose Oyama, Hidetoshi ShimodairaEMNLP 2023 · 6 citations
