Sonata: Self-Supervised Learning of Reliable Point Representations
Xiaoyang Wu, Daniel DeTone, Duncan P. Frost, Tianwei Shen, Chris Xie, Nan Yang, Jakob J. Engel, Richard A. Newcombe, Hengshuang Zhao, Julian Straub
摘要
We address it through two key strategies: obscuring spatial information and enhancing the reliance on input features, ultimately composing a Sonata of 140k point clouds through self-distillation. Sonata is simple and intuitive, yet its learned representations are strong and reliable: zero-shot visualizations demonstrate semantic grouping, alongside strong spatial reasoning through nearestneighbor relationships. Sonata demonstrates exceptional parameter and data efficiency, tripling linear probing accuracy (from 21.8% to 72.5%) on ScanNet and nearly doubling performance with only 1% of the data compared to previous approaches. Full fine-tuning further advances SOTA across both 3D indoor and outdoor perception tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper30
- SpatialLM: Training Large Language Models for Structured Indoor ModelingYongsen Mao, Junhao Zhong, Chuan Fang, Jia Zheng 等NeurIPS 2025 · 被引用 89 次
- PointWorld: Scaling 3D World Models for In-The-Wild Robotic ManipulationWenlong Huang, Yu-Wei Chao, Arsalan Mousavian, Ming-Yu Liu 等CVPR 2026 · 被引用 87 次
- Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial RepresentationsYujia Zhang, Xiaoyang Wu, Yixing Lao, Chengyao Wang 等NeurIPS 2025 · 被引用 47 次
- LitePT: Lighter Yet Stronger Point TransformerYuanwen Yue, Damien Robert, Jianyuan Wang, Sunghwan Hong 等CVPR 2026 · 被引用 25 次
- Chorus: Multi-Teacher Pretraining for Holistic 3D Gaussian Scene EncodingYue Li, Qi Ma, Runyi Yang, Mengjiao Ma 等CVPR 2026 · 被引用 10 次
它引用的顶会 Paper55
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal 等NeurIPS 2020 · 被引用 5,249 次
相关 Paper
- GaussianCross: Cross-modal Self-supervised 3D Representation Learning via Gaussian SplattingLei Yao, Yi Wang, Yi Zhang, Moyun Liu 等ACM MM 2025 · 被引用 2 次
- Real-time 3D Object Detection with Inference-Aligned LearningChenyu Zhao, Xianwei Zheng, Zimin Xia, Linwei Yue 等AAAI 2026 · 被引用 1 次
- Asymmetric Dual Self-Distillation for 3D Self-Supervised Representation LearningRemco F. Leijenaar, Hamidreza KasaeiNeurIPS 2025
- Self-Supervised Pretraining of 3D Features on any Point-CloudZaiwei Zhang, Rohit Girdhar, Armand Joulin, Ishan MisraICCV 2021 · 被引用 333 次
- DOS: Distilling Observable Softmaps of Zipfian Prototypes for Self-Supervised Point RepresentationMohamed Abdelsamad, Michael Ulrich, Bin Yang, Miao Zhang 等AAAI 2026 · 被引用 1 次
