Graphics Capsule: Learning Hierarchical 3D Face Representations from 2D Images
Chang Yu, Xiangyu Zhu, Xiaomei Zhang, Zhaoxiang Zhang, Zhen Lei
Abstract
The function of constructing the hierarchy of objects is important to the visual process of the human brain. Previous studies have successfully adopted capsule networks to decompose the digits and faces into parts in an unsupervised manner to investigate the similar perception mechanism of neural networks. However, their descriptions are restricted to the 2D space, limiting their capacities to imitate the intrinsic 3D perception ability of humans. In this paper, we propose an Inverse Graphics Capsule Network (IGC-Net) to learn the hierarchical 3D face representations from large-scale unlabeled images. The core of IGC-Net is a new type of capsule, named graphics capsule, which represents 3D primitives with interpretable parameters in computer graphics (CG), including depth, albedo, and 3D pose. Specifically, IGC-Net first decomposes the objects into a set of semantic-consistent part-level descriptions and then assembles them into object-level descriptions to build the hierarchy. The learned graphics capsules reveal how the neural networks, oriented at visual perception, understand faces as a hierarchy of 3D models. Besides, the discovered parts can be deployed to the unsupervised face segmentation task to evaluate the semantic consistency of our method. Moreover, the part-level descriptions with explicit physical meanings provide insight into the face analysis that originally runs in a black box, such as the importance of shape and texture for face recognition. Experiments on CelebA, BP4D, and Multi-PIE demonstrate the characteristics of our IGC-Net.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0d09bbae-1e44-411f-8815-d3ee6faf241cCited by top-tier papers1
Ask how each one uses itBuilds on7
- Unsupervised Part Discovery from Contrastive ReconstructionSubhabrata Choudhury, Iro Laina, Christian Rupprecht, Andrea VedaldiNeurIPS 2021 · 74 citations
- Unsupervised Part Representation by Flow CapsulesSara Sabour, Andrea Tagliasacchi, Soroosh Yazdani, Geoffrey E. Hinton et al.ICML 2021 · 41 citations
- HP-Capsule: Unsupervised Face Part Discovery by Hierarchical Parsing Capsule NetworkChang Yu, Xiangyu Zhu, Xiaomei Zhang, Zidu Wang et al.CVPR 2022 · 18 citations
- Physically-guided Disentangled Implicit Rendering for 3D Face ModelingZhenyu Zhang, Yanhao Ge, Ying Tai, Weijian Cao et al.CVPR 2022 · 5 citations
- Unsupervised Part Segmentation Through Disentangling Appearance and ShapeShilong Liu, Lei Zhang, Xiao Yang, Hang Su et al.CVPR 2021
Related papers
- RIM-Net: Recursive Implicit Fields for Unsupervised Learning of Hierarchical Shape StructuresChengjie Niu, Manyi Li, Kai Xu, Hao ZhangCVPR 2022 · 18 citations
- 3DP3: 3D Scene Perception via Probabilistic ProgrammingNishad Gothoskar, Marco F. Cusumano-Towner, Ben Zinberg, Matin Ghavamizadeh et al.NeurIPS 2021 · 59 citations
- Generative Scene Graph NetworksFei Deng, Zhuo Zhi, Donghun Lee, Sungjin AhnICLR 2021 · 10 citations
- CSG-Stump: A Learning Friendly CSG-Like Representation for Interpretable Shape ParsingDaxuan Ren, Jianmin Zheng, Jianfei Cai, Jiatong Li et al.ICCV 2021 · 71 citations
- Neural Parts: Learning Expressive 3D Shape Abstractions With Invertible Neural NetworksDespoina Paschalidou, Angelos Katharopoulos, Andreas Geiger, Sanja FidlerCVPR 2021
