ShapeCaptioner: Generative Caption Network for 3D Shapes by Learning a Mapping from Parts Detected in Multiple Views to Sentences
Zhizhong Han, Chao Chen, Yu-Shen Liu, Matthias Zwicker
Abstract
3D shape captioning is a challenging application in 3D shape understanding. Captions from recent multi-view based methods reveal that they cannot capture part-level characteristics of 3D shapes. This leads to a lack of detailed part-level description in captions, which human tend to focus on. To resolve this issue, we propose ShapeCaptioner, a generative caption network, to perform 3D shape captioning from semantic parts detected in multiple views. Our novelty lies in learning the knowledge of part detection in multiple views from 3D shape segmentations and transferring this knowledge to facilitate learning the mapping from 3D shapes to sentences. Specifically, ShapeCaptioner aggregates the parts detected in multiple colored views using our novel part class specific aggregation to represent a 3D shape, and then, employs a sequence to sequence model to generate the caption. Our outperforming results show that ShapeCaptioner can learn 3D shape features with more detailed part characteristics to facilitate better 3D shape captioning than previous work.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 94efea07-08ac-47cf-aad3-d42c95fd4030Cited by top-tier papers18
- SnowflakeNet: Point Cloud Completion by Snowflake Point Deconvolution with Skip-TransformerPeng Xiang, Xin Wen, Yu-Shen Liu, Yan-Pei Cao et al.ICCV 2021 · 318 citations
- Neural-Pull: Learning Signed Distance Function from Point clouds by Learning to Pull Space onto SurfaceBaorui Ma, Zhizhong Han, Yu-Shen Liu, Matthias ZwickerICML 2021 · 215 citations
- Surface Reconstruction from Point Clouds by Learning Predictive Context PriorsBaorui Ma, Yu-Shen Liu, Matthias Zwicker, Zhizhong HanCVPR 2022 · 67 citations
- Reconstructing Surfaces for Sparse Point Clouds with On-Surface PriorsBaorui Ma, Yu-Shen Liu, Zhizhong HanCVPR 2022 · 66 citations
- Towards Implicit Text-Guided 3D Shape GenerationZhengzhe Liu, Yi Wang, Xiaojuan Qi, Chi-Wing FuCVPR 2022 · 59 citations
Builds on1
Related papers
- ExCap3d: Expressive 3D Scene Understanding via Object Captioning with Varying DetailChandan Yeshwanth, Dávid Rozenberszki, Angela DaiICCV 2025 · 2 citations
- Scan2Cap: Context-Aware Dense Captioning in RGB-D ScansDave Zhenyu Chen, Ali Gholami, Matthias Nießner, Angel X. ChangCVPR 2021
- Learning Part Generation and Assembly for Structure-Aware Shape SynthesisJun Li, Chengjie Niu, Kai XuAAAI 2020 · 85 citations
- Generative 3D Part Assembly via Part-Whole-Hierarchy Message PassingBi'an Du, Xiang Gao, Wei Hu, Renjie LiaoCVPR 2024
- AutoPartGen: Autoregressive 3D Part Generation and DiscoveryMinghao Chen, Jianyuan Wang, Roman Shapovalov, Tom Monnier et al.NeurIPS 2025 · 29 citations
