One-shot 3D Object Canonicalization based on Geometric and Semantic Consistency
Li Jin, Yujie Wang, Wenzheng Chen, Qiyu Dai, Qingzhe Gao, Xueying Qin, Baoquan Chen
Abstract
3D object canonicalization is a fundamental task, essential for various downstream tasks. Existing methods rely on either cumbersome manual processes or priors learned from extensive, per-category training samples. Real-world datasets, however, often exhibit long-tail distributions, challenging existing learning-based methods, especially in categories with limited samples. We address this by introducing the first one-shot category-level object canonicalization framework that operates under arbitrary poses, requiring only a single canonical model as a reference (the "prior model") for each category. To canonicalize any object, our framework first extracts semantic cues with large language models (LLMs) and vision-language models (VLMs) to establish correspondences with the prior model. We introduce a novel joint energy function to enforce geometric and semantic consistency, aligning object orientations precisely despite significant shape variations. Moreover, we adopt a support-plane strategy to reduce search space for initial poses and utilize a semantic relationship map to select the canonical pose from multiple hypotheses. Extensive experiments on multiple datasets demonstrate that our framework achieves state-of-the-art performance and validates key design choices. Using our framework, we create the Canonical Objaverse Dataset (COD), canonicalizing 32K samples in the Objaverse-LVIS dataset, underscoring the effectiveness of our framework on handling large-scale datasets. Project page at https://github.com/JinLi998/CanonObjaverseDataset
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 656cb2c5-e1ba-4e27-a5b4-8868206a9475Cited by top-tier papers2
- Orientation Matters: Making 3D Generative Models Orientation-AlignedYichong Lu, Yuzhuo Tian, Zijin Jiang, Yikun Zhao et al.NeurIPS 2025 · 15 citations
- CoSMo3D: Open-World Promptable 3D Semantic Segmentation through LLM-Guided Canonical Spatial ModelingLi Jin, Weikai Chen, Yujie Wang, Yingda Yin et al.CVPR 2026
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- PointNeXt: Revisiting PointNet++ with Improved Training and Scaling StrategiesGuocheng Qian, Yuchen Li, Houwen Peng, Jinjie Mai et al.NeurIPS 2022 · 1,270 citations
- Learning Shape Templates With Structured Implicit FunctionsKyle Genova, Forrester Cole, Daniel Vlasic, Aaron Sarna et al.ICCV 2019 · 427 citations
- CLAY: A Controllable Large-scale Generative Model for Creating High-quality 3D AssetsLongwen Zhang, Ziyu Wang, Qixuan Zhang, Qiwei Qiu et al.SIGGRAPH 2024 · 148 citations
Related papers
- One-Shot Open Affordance Learning with Foundation ModelsGen Li, Deqing Sun, Laura Sevilla-Lara, Varun JampaniCVPR 2024
- Learning Canonical Shape Space for Category-Level 6D Object Pose and Size EstimationDengsheng Chen, Jun Li, Zheng Wang, Kai XuCVPR 2020
- PointAlign: Feature-Level Alignment Regularization for 3D Vision-Language ModelsYuanhao Su, Shaofeng Zhang, Xiaosong Jia, Qi FanCVPR 2026 · 1 citation
- Zero-Shot Scene Reconstruction from Single Images with Deep Prior AssemblyJunsheng Zhou, Yu-Shen Liu, Zhizhong HanNeurIPS 2024 · 42 citations
- ConceptPose: Training-Free Zero-Shot Object Pose Estimation using Concept VectorsLiming Kuang, Yordanka Velikova, Mahdi Saleh, Jan-Nico Zaech et al.CVPR 2026 · 5 citations
