Semantify: Simplifying the Control of 3D Morphable Models using CLIP
Omer Gralnik, Guy Gafni, Ariel Shamir
摘要
We present Semantify: a self-supervised method that utilizes the semantic power of CLIP language-vision foundation model [32] to simplify the control of 3D morphable models. Given a parametric model, training data is created by randomly sampling the model’s parameters, creating various shapes and rendering them. The similarity between the output images and a set of word descriptors is calculated in CLIP’s latent space. Our key idea is first to choose a small set of semantically meaningful and disentangled descriptors that characterize the 3DMM, and then learn a non-linear mapping from scores across this set to the parametric coefficients of the given 3DMM. The nonlinear mapping is defined by training a neural network without a human-in-the-loop. We present results on numerous 3DMMs: body shape models, face shape and expression models, as well as animal shapes. We demonstrate how our method defines a simple slider interface for intuitive modeling, and show how the mapping can be used to instantly fit a 3D parametric body shape to in-the-wild images. See our project page at https://omergral.github.io/Semantify/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- UP2You: Fast Reconstruction of Yourself from Unconstrained Photo CollectionsZeyu Cai, Ziyang Li, Xiaoben Li, Boqian Li 等ICLR 2026 · 被引用 11 次
- iPose: Interactive Human Pose Reconstruction from VideoJingyuan Liu, Li-Yi Wei, Ariel Shamir, Takeo IgarashiCHI 2024 · 被引用 8 次
- UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and EditingYiheng Li, Ruibing Hou, Hong Chang, Shiguang Shan 等CVPR 2025
它引用的顶会 Paper15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or 等ICCV 2021 · 被引用 1,437 次
- Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the LoopNikos Kolotouros, Georgios Pavlakos, Michael J. Black, Kostas DaniilidisICCV 2019 · 被引用 1,139 次
- Zero-Shot Text-Guided Object Generation with Dream FieldsAjay Jain, Ben Mildenhall, Jonathan T. Barron, Pieter Abbeel 等CVPR 2022 · 被引用 361 次
相关 Paper
- ClipFace: Text-guided Editing of Textured 3D Morphable ModelsShivangi Aneja, Justus Thies, Angela Dai, Matthias NießnerSIGGRAPH 2023 · 被引用 40 次
- StyleRig: Rigging StyleGAN for 3D Control Over Portrait ImagesAyush Tewari, Mohamed A. Elgharib, Gaurav Bharaj, Florian Bernard 等CVPR 2020
- Common3D: Self-Supervised Learning of 3D Morphable Models for Common Objects in Neural Feature SpaceLeonhard Sommer, Olaf Dünkel, Christian Theobalt, Adam KortylewskiCVPR 2025
- AvatarCLIP: zero-shot text-driven generation and animation of 3D avatarsFangzhou Hong, Mingyuan Zhang, Liang Pan, Zhongang Cai 等SIGGRAPH 2022 · 被引用 213 次
- StyleMorph: Disentangled 3D-Aware Image Synthesis with a 3D Morphable StyleGANEric-Tuan Le, Edward Bartrum, Iasonas KokkinosICLR 2023
