LEMaRT: Label-Efficient Masked Region Transform for Image Harmonization
Sheng Liu, Cong Phuoc Huynh, Cong Chen, Maxim Arap, Raffay Hamid
Abstract
We present a simple yet effective self-supervised pretraining method for image harmonization which can leverage large-scale unannotated image datasets. To achieve this goal, we first generate pre-training data online with our Label-Efficient Masked Region Transform (LEMaRT) pipeline. Given an image, LEMaRT generates a foreground mask and then applies a set of transformations to perturb various visual attributes, e.g., defocus blur, contrast, saturation, of the region specified by the generated mask. We then pre-train image harmonization models by recovering the original image from the perturbed image. Secondly, we introduce an image harmonization model, namely SwinIH, by retrofitting the Swin Transformer [27] with a combination of local and global self-attention mechanisms. Pretraining SwinIH with LEMaRT results in a new state of the art for image harmonization, while being label-efficient, i.e., consuming less annotated data for fine-tuning than existing methods. Notably, on iHarmony4 dataset [8], SwinIH outperforms the state of the art, i.e., SCS-Co [16] by a margin of 0.4 dB when it is fine-tuned on only 50% of the training data, and by 1.0 dB when it is trained on the full training dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7de298d9-1e45-488a-9bc2-202541fcc850Cited by top-tier papers6
- Hierarchical Dynamic Image HarmonizationHaoxing Chen, Zhangxuan Gu, Yaohui Li, Jun Lan et al.ACM MM 2023 · 24 citations
- Progressive Painterly Image Harmonization from Low-Level Styles to High-Level StylesLi Niu, Yan Hong, Junyan Cao, Liqing ZhangAAAI 2024 · 3 citations
- Painterly Image Harmonization by Learning from Painterly ObjectsLi Niu, Junyan Cao, Yan Hong, Liqing ZhangAAAI 2024 · 2 citations
- Video Harmonization with Triplet Spatio-Temporal Variation PatternsZonghui Guo, Xinyu Han, Jie Zhang, Shiguang Shan et al.CVPR 2024
- Towards Enhanced Image Inpainting: Mitigating Unwanted Object Insertion and Preserving Color ConsistencyYikai Wang, Chenjie Cao, Junqiu Yu, Ke Fan et al.CVPR 2025
Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
Related papers
- SSH: A Self-Supervised Framework for Image HarmonizationYifan Jiang, He Zhang, Jianming Zhang, Yilin Wang et al.ICCV 2021 · 108 citations
- Image Harmonization with TransformerZonghui Guo, Dongsheng Guo, Haiyong Zheng, Zhaorui Gu et al.ICCV 2021 · 95 citations
- FreePIH: Training-Free Painterly Image Harmonization with Diffusion ModelRuibin Li, Jingcai Guo, Qihua Zhou, Song GuoACM MM 2024 · 2 citations
- PCT-Net: Full Resolution Image Harmonization Using Pixel-Wise Color TransformationsJulian Jorge Andrade Guerreiro, Mitsuru Nakazawa, Björn StengerCVPR 2023
- SinDDM: A Single Image Denoising Diffusion ModelVladimir Kulikov, Shahar Yadin, Matan Kleiner, Tomer MichaeliICML 2023 · 113 citations
