SpaceEdit: Learning a Unified Editing Space for Open-Domain Image Color Editing
Jing Shi, Ning Xu, Haitian Zheng, Alex Smith, Jiebo Luo, Chenliang Xu
Abstract
Recently, large pretrained models (e.g., BERT, Style-GAN, CLIP) show great knowledge transfer and generalization capability on various downstream tasks within their domains. Inspired by these efforts, in this paper we propose a unified model for open-domain image editing focusing on color and tone adjustment of open-domain images while keeping their original content and structure. Our model learns a unified editing space that is more semantic, intu-itive, and easy to manipulate than the operation space (e.g., contrast, brightness, color curve) used in many existing photo editing softwares. Our model belongs to the image-to-image translation framework which consists of an image encoder and decoder, and is trained on pairs of before-and-after edited images to produce multimodal outputs. We show that by inverting image pairs into latent codes of the learned editing space, our model can be leveraged for vari-ous downstream editing tasks such as language-guided image editing, personalized editing, editing-style clustering, retrieval, etc. We extensively study the unique properties of the editing space in experiments and demonstrate superior performance on the aforementioned tasks <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> Code and supplementary material can be found at the project page https://jshi31.github.io/SpaceEdit.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7e8e3e85-3ee9-4b33-a7cf-26e6cc738840Cited by top-tier papers4
- StyleT2I: Toward Compositional and High-Fidelity Text-to-Image SynthesisZhiheng Li, Martin Renqiang Min, Kai Li, Chenliang XuCVPR 2022 · 38 citations
- Hierarchical Dynamic Image HarmonizationHaoxing Chen, Zhangxuan Gu, Yaohui Li, Jun Lan et al.ACM MM 2023 · 24 citations
- Goal Conditioned Reinforcement Learning for Photo Finishing TuningJiarui Wu, Yujin Wang, Lingen Li, Zhang Fan et al.NeurIPS 2024 · 9 citations
- DocEdit: Language-Guided Document EditingPuneet Mathur, Rajiv Jain, Jiuxiang Gu, Franck Dernoncourt et al.AAAI 2023
Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or et al.ICCV 2021 · 1,437 citations
- Unsupervised Discovery of Interpretable Directions in the GAN Latent SpaceAndrey Voynov, Artem BabenkoICML 2020 · 459 citations
- On the "steerability" of generative adversarial networksAli Jahanian, Lucy Chai, Phillip IsolaICLR 2020 · 421 citations
Related papers
- FlexIT: Towards Flexible Semantic Image TranslationGuillaume Couairon, Asya Grechka, Jakob Verbeek, Holger Schwenk et al.CVPR 2022 · 36 citations
- Unpaired Image-to-Image Translation via Latent Energy TransportYang Zhao, Changyou ChenCVPR 2021
- Hist2Style: Histogram-Guided Stylization with Bilateral GridsDekel Galor, Adam Pikielny, Zhoutong Zhang, Ke Wang et al.CVPR 2026
- DeltaEdit: Exploring Text-free Training for Text-Driven Image ManipulationCVPR 2023
- Prompt Tuning Inversion for Text-Driven Image Editing Using Diffusion ModelsWenkai Dong, Song Xue, Xiaoyue Duan, Shumin HanICCV 2023 · 104 citations
