Utonia: Toward One Encoder for All Point Clouds
Yujia Zhang, Xiaoyang Wu, Yunhan Yang, Xianzhe Fan, Han Li, Yuechen Zhang, Zehao Huang, Naiyan Wang, Hengshuang Zhao
Abstract
We dream of a future where point clouds from all domains can come together to shape a single model that benefits them all. Toward this goal, we present Utonia, a first step toward training a single self-supervised point transformer encoder across heterogeneous domains, spanning remote sensing, outdoor LiDAR, indoor RGB-D sequences, object-centric CAD models, and point clouds lifted from RGB-only videos. Despite their distinct sensing geometries, densities, and priors, Utonia learns a consistent representation space that transfers across domains. This unification improves perception capability while revealing intriguing emergent behaviors that arise only when domains are trained jointly. Beyond perception, we observe that Utonia representations can also benefit embodied and multimodal reasoning: conditioning vision-language-action policies on Utonia features improves robotic manipulation, and integrating them into vision-language models yields gains on spatial reasoning. We hope Utonia can serve as a step toward foundation models for sparse 3D data, and support downstream applications in AR/VR, robotics, and autonomous driving.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4c081134-29bb-4ce4-b2ed-6dab12b9fa66Builds on30
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- Point Transformer V2: Grouped Vector Attention and Partition-based PoolingXiaoyang Wu, Yixing Lao, Li Jiang, Xihui Liu et al.NeurIPS 2022 · 924 citations
- Depth Anything 3: Recovering the Visual Space from Any ViewsHaotong Lin, Sili Chen, Jun Hao Liew, Donny Y. Chen et al.ICLR 2026 · 720 citations
- ScanNet++: A High-Fidelity Dataset of 3D Indoor ScenesChandan Yeshwanth, Yueh-Cheng Liu, Matthias Nießner, Angela DaiICCV 2023 · 659 citations
Related papers
- Towards Universal LiDAR-Based 3D Object Detection by Multi-Domain Knowledge TransferGuile Wu, Tongtong Cao, Bingbing Liu, Xingxin Chen et al.ICCV 2023 · 7 citations
- UniDet3D: Multi-dataset Indoor 3D Object DetectionMaksim Kolodiazhnyi, Anna Vorontsova, Matvey Skripkin, Danila Rukhovich et al.AAAI 2025 · 7 citations
- Uni-Sight: An E2E Vision-Language-Action System Unifying Multi-View Alignment and Multi-Modal FusionDaixun Li, Sibo He, Jiayun Tian, Yusi Zhang et al.ACM MM 2025 · 1 citation
- Multi-Modal Fusion Transformer for End-to-End Autonomous DrivingAditya Prakash, Kashyap Chitta, Andreas GeigerCVPR 2021
- UniTR: A Unified and Efficient Multi-Modal Transformer for Bird's-Eye-View RepresentationHaiyang Wang, Hao Tang, Shaoshuai Shi, Aoxue Li et al.ICCV 2023 · 106 citations
