pOps: Photo-Inspired Diffusion Operators
Elad Richardson, Yuval Alaluf, Ali Mahdavi-Amiri, Daniel Cohen-Or
Abstract
Text-guided image generation enables the creation of visual content from textual descriptions. However, certain visual concepts cannot be effectively conveyed through language alone. This has sparked a renewed interest in utilizing the CLIP image embedding space for more visually-oriented tasks through methods such as IP-Adapter. Interestingly, the CLIP image embedding space has been shown to be semantically meaningful, where linear operations within this space yield semantically meaningful results. Yet, the specific meaning of these operations can vary unpredictably across different images. To harness this potential, we introduce pOps, a framework that trains specific semantic operators directly on CLIP image embeddings. Each pOps operator is built upon a pretrained Diffusion Prior model. While the Diffusion Prior model was originally trained to map between text embeddings and image embeddings, we demonstrate that it can be tuned to accommodate new input conditions, resulting in a diffusion operator. Working directly over image embeddings not only improves our ability to learn semantic operations but also allows us to directly use a textual CLIP loss as an additional supervision when needed. We show that pOps can be used to learn a variety of photo-inspired operators with distinct semantic meanings. These operators can then serve as creative tools within a design process, enabling artists to semantically manipulate visual concepts as part of their generative workflow. Finally, we show that pOps can be easily plugged into pretrained image diffusion models alongside existing spatial adapters, offering control over both semantics and structure.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 83b21a11-af3d-4c5f-8f4c-5317c71b43f5Cited by top-tier papers4
- Creative Blends of Visual ConceptsZhida Sun, Zhenyao Zhang, Yue Zhang, Min Lu et al.CHI 2025 · 11 citations
- IP-Composer: Semantic Composition of Visual ConceptsSara Dorfman, Dana Cohen-Bar, Rinon Gal, Daniel Cohen-OrSIGGRAPH 2025 · 4 citations
- Muses: Designing, Composing, Generating Nonexistent Fantasy 3D Creatures without TrainingHexiao Lu, Xiaokun Sun, Zeyu Cai, Hao Guo et al.CVPR 2026
- Alterbute: Editing Intrinsic Attributes of Objects in ImagesTal Reiss, Daniel Winter, Matan Cohen, Alex Rav-Acha et al.ICML 2026
Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- Instruct-CLIP: Improving Instruction-Guided Image Editing with Automated Data Refinement Using Contrastive LearningSherry X. Chen, Misha Sra, Pradeep SenCVPR 2025
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or et al.ICCV 2021 · 1,437 citations
- CLIP-PAE: Projection-Augmentation Embedding to Extract Relevant Features for a Disentangled, Interpretable and Controllable Text-Guided Face ManipulationChenliang Zhou, Fangcheng Zhong, Cengiz ÖztireliSIGGRAPH 2023 · 13 citations
- TexSliders: Diffusion-Based Texture Editing in CLIP SpaceJulia Guerrero-Viu, Milos Hasan, Arthur Roullier, Midhun Harikumar et al.SIGGRAPH 2024 · 18 citations
- Towards Counterfactual Image Manipulation via CLIPYingchen Yu, Fangneng Zhan, Rongliang Wu, Jiahui Zhang et al.ACM MM 2022 · 33 citations
