TextDeformer: Geometry Manipulation using Text Guidance
William Gao, Noam Aigerman, Thibault Groueix, Vova Kim, Rana Hanocka
Abstract
Fig. 1. TextDeformer deforms a source shape into various text-specified targets. The mesh colors visualize the smoothness of the mappings.
We present a technique for automatically producing a deformation of an input triangle mesh, guided solely by a text prompt. Our framework is capable of deformations that produce both large, low-frequency shape changes, and small high-frequency details. Our framework relies on differentiable rendering to connect geometry to powerful pre-trained image encoders, such as CLIP and DINO. Notably, updating mesh geometry by taking gradient steps through differentiable rendering is notoriously challenging, commonly resulting in deformed meshes with significant artifacts. These difficulties are amplified by noisy and inconsistent gradients from CLIP. To overcome this limitation, we opt to represent our mesh deformation through Jacobians, which updates deformations in a global, smooth manner (rather than locallysub-optimal steps). Our key observation is that Jacobians are a representation that favors smoother, large deformations, leading to a global relation between vertices and pixels, and avoiding localized noisy gradients. Additionally, to ensure the resulting shape is coherent from all 3D viewpoints, we encourage the deep features computed on the 2D encoding of the rendering to be consistent for a given vertex from all viewpoints. We demonstrate that our method is capable of smoothly-deforming a wide variety of source mesh and target text prompts, achieving both large modifications to, e.g., body proportions of animals, as well as adding fine semantic details, such as shoe laces on an army boot and fine details of a face.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 918082f9-932f-430f-bb4f-5a1504419b50Cited by top-tier papers46
- Magic123: One Image to High-Quality 3D Object Generation Using Both 2D and 3D Diffusion PriorsGuocheng Qian, Jinjie Mai, Abdullah Hamdi, Jian Ren et al.ICLR 2024 · 444 citations
- GaussianEditor: Editing 3D Gaussians Delicately with Text InstructionsJunjie Wang, Jiemin Fang, Xiaopeng Zhang, Lingxi Xie et al.CVPR 2024 · 65 citations
- HumanGaussian: Text-Driven 3D Human Generation with Gaussian SplattingXian Liu, Xiaohang Zhan, Jiaxiang Tang, Ying Shan et al.CVPR 2024 · 42 citations
- Style2Fab: Functionality-Aware Segmentation for Fabricating Personalized 3D Models with Generative AIFaraz Faruqi, Ahmed Katary, Tarik Hasic, Amira Abdel-Rahman et al.UIST 2023 · 39 citations
- Lyra: Generative 3D Scene Reconstruction via Video Diffusion Model Self-DistillationSherwin Bahmani, Tianchang Shen, Jiawei Ren, Jiahui Huang et al.ICLR 2026 · 33 citations
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam et al.ICML 2022 · 4,691 citations
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or et al.ICCV 2021 · 1,437 citations
Related papers
- Text2Mesh: Text-Driven Neural Stylization for MeshesOscar Michel, Roi Bar-On, Richard Liu, Sagie Benaim et al.CVPR 2022
- 3D Highlighter: Localizing Regions on 3D Shapes via Text DescriptionsDale Decatur, Itai Lang, Rana HanockaCVPR 2023
- TANGO: Text-driven Photorealistic and Robust 3D Stylization via Lighting DecompositionYongwei Chen, Rui Chen, Jiabao Lei, Yabin Zhang et al.NeurIPS 2022 · 112 citations
- CLIPVG: Text-Guided Image Manipulation Using Differentiable Vector GraphicsYiren Song, Xuning Shao, Kang Chen, Weidong Zhang et al.AAAI 2023 · 50 citations
- Text2Tex: Text-driven Texture Synthesis via Diffusion ModelsDave Zhenyu Chen, Yawar Siddiqui, Hsin-Ying Lee, Sergey Tulyakov et al.ICCV 2023 · 262 citations
