TexSliders: Diffusion-Based Texture Editing in CLIP Space
Julia Guerrero-Viu, Milos Hasan, Arthur Roullier, Midhun Harikumar, Yiwei Hu, Paul Guerrero, Diego Gutierrez, Belén Masiá, Valentin Deschaintre
摘要
Generative models have enabled intuitive image creation and manipulation using natural language. In particular, diffusion models have recently shown remarkable results for natural image editing. In this work, we propose to apply diffusion techniques to edit textures, a specific class of images that are an essential part of 3D content creation pipelines. We analyze existing editing methods and show that they are not directly applicable to textures, since their common underlying approach, manipulating attention maps, is unsuitable for the texture domain. To address this, we propose a novel approach that instead manipulates CLIP image embeddings to condition the diffusion generation. We define editing directions using simple text prompts (e.g., “aged wood” to “new wood”) and map these to CLIP image embedding space using a texture prior, with a sampling-based approach that gives us identity-preserving directions in CLIP space. To further improve identity preservation, we project these directions to a CLIP subspace that minimizes identity variations resulting from entangled texture attributes. Our editing pipeline facilitates the creation of arbitrary sliders using natural language prompts only, with no ground-truth annotated data necessary.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Magnet: We Never Know How Text-to-Image Diffusion Models Work, Until We Learn How Vision-Language Models FunctionChenyi Zhuang, Ying Hu, Pan GaoNeurIPS 2024 · 被引用 28 次
- PrimitiveAnything: Human-Crafted 3D Primitive Assembly Generation with Auto-Regressive transformerJingwen Ye, Yuze He, Yanning Zhou, Yiqin Zhu 等SIGGRAPH 2025 · 被引用 5 次
- IP-Composer: Semantic Composition of Visual ConceptsSara Dorfman, Dana Cohen-Bar, Rinon Gal, Daniel Cohen-OrSIGGRAPH 2025 · 被引用 4 次
- From Cradle to Cane: A Two-Pass Framework for High-Fidelity Lifespan Face AgingTao Liu, Dafeng Zhang, Gengchen Li, Shizhuo Liu 等NeurIPS 2025 · 被引用 3 次
- Generative detail enhancement for physically based materialsSaeed Hadadan, Benedikt Bitterli, Tizian Zeltner, Jan Novák 等SIGGRAPH 2025 · 被引用 3 次
它引用的顶会 Paper27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam 等ICML 2022 · 被引用 4,691 次
相关 Paper
- TEXTure: Text-Guided Texturing of 3D ShapesElad Richardson, Gal Metzer, Yuval Alaluf, Raja Giryes 等SIGGRAPH 2023 · 被引用 196 次
- Blended Diffusion for Text-driven Editing of Natural ImagesOmri Avrahami, Dani Lischinski, Ohad FriedCVPR 2022 · 被引用 670 次
- Prompt Tuning Inversion for Text-Driven Image Editing Using Diffusion ModelsWenkai Dong, Song Xue, Xiaoyue Duan, Shumin HanICCV 2023 · 被引用 104 次
- DiffEdit: Diffusion-based semantic image editing with mask guidanceGuillaume Couairon, Jakob Verbeek, Holger Schwenk, Matthieu CordICLR 2023 · 被引用 102 次
- Diffusion Handles Enabling 3D Edits for Diffusion Models by Lifting Activations to 3DKarran Pandey, Paul Guerrero, Matheus Gadelha, Yannick Hold-Geoffroy 等CVPR 2024 · 被引用 19 次
