SpaceEdit: Learning a Unified Editing Space for Open-Domain Image Color Editing
Jing Shi, Ning Xu, Haitian Zheng, Alex Smith, Jiebo Luo, Chenliang Xu
摘要
Recently, large pretrained models (e.g., BERT, Style-GAN, CLIP) show great knowledge transfer and generalization capability on various downstream tasks within their domains. Inspired by these efforts, in this paper we propose a unified model for open-domain image editing focusing on color and tone adjustment of open-domain images while keeping their original content and structure. Our model learns a unified editing space that is more semantic, intu-itive, and easy to manipulate than the operation space (e.g., contrast, brightness, color curve) used in many existing photo editing softwares. Our model belongs to the image-to-image translation framework which consists of an image encoder and decoder, and is trained on pairs of before-and-after edited images to produce multimodal outputs. We show that by inverting image pairs into latent codes of the learned editing space, our model can be leveraged for vari-ous downstream editing tasks such as language-guided image editing, personalized editing, editing-style clustering, retrieval, etc. We extensively study the unique properties of the editing space in experiments and demonstrate superior performance on the aforementioned tasks <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> Code and supplementary material can be found at the project page https://jshi31.github.io/SpaceEdit.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- StyleT2I: Toward Compositional and High-Fidelity Text-to-Image SynthesisZhiheng Li, Martin Renqiang Min, Kai Li, Chenliang XuCVPR 2022 · 被引用 38 次
- Hierarchical Dynamic Image HarmonizationHaoxing Chen, Zhangxuan Gu, Yaohui Li, Jun Lan 等ACM MM 2023 · 被引用 24 次
- Goal Conditioned Reinforcement Learning for Photo Finishing TuningJiarui Wu, Yujin Wang, Lingen Li, Zhang Fan 等NeurIPS 2024 · 被引用 9 次
- DocEdit: Language-Guided Document EditingPuneet Mathur, Rajiv Jain, Jiuxiang Gu, Franck Dernoncourt 等AAAI 2023
它引用的顶会 Paper25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or 等ICCV 2021 · 被引用 1,437 次
- Unsupervised Discovery of Interpretable Directions in the GAN Latent SpaceAndrey Voynov, Artem BabenkoICML 2020 · 被引用 459 次
- On the "steerability" of generative adversarial networksAli Jahanian, Lucy Chai, Phillip IsolaICLR 2020 · 被引用 421 次
相关 Paper
- FlexIT: Towards Flexible Semantic Image TranslationGuillaume Couairon, Asya Grechka, Jakob Verbeek, Holger Schwenk 等CVPR 2022 · 被引用 36 次
- Unpaired Image-to-Image Translation via Latent Energy TransportYang Zhao, Changyou ChenCVPR 2021
- Hist2Style: Histogram-Guided Stylization with Bilateral GridsDekel Galor, Adam Pikielny, Zhoutong Zhang, Ke Wang 等CVPR 2026
- DeltaEdit: Exploring Text-free Training for Text-Driven Image ManipulationCVPR 2023
- Prompt Tuning Inversion for Text-Driven Image Editing Using Diffusion ModelsWenkai Dong, Song Xue, Xiaoyue Duan, Shumin HanICCV 2023 · 被引用 104 次
