Refine-by-Align: Reference-Guided Artifacts Refinement through Semantic Alignment
Yizhi Song, Liu He, Zhifei Zhang, Soo Ye Kim, He Zhang, Wei Xiong, Zhe Lin, Brian L. Price, Scott Cohen, Jianming Zhang, Daniel G. Aliaga
Abstract
Personalized image generation has emerged from the recent advancements in generative models. However, these generated personalized images often suffer from localized artifacts such as incorrect logos, reducing fidelity and fine-grained identity details of the generated results. Furthermore, there is little prior work tackling this problem. To help improve these identity details in the personalized image generation, we introduce a new task: reference-guided artifacts refinement. We present Refine-by-Align , a first-of-its-kind model that employs a diffusion-based framework to address this challenge. Our model consists of two stages: Alignment Stage and Refinement Stage, which share weights of a unified neural network model. Given a generated image, a masked artifact region, and a reference image, the alignment stage identifies and extracts the corresponding regional features in the reference, which are then used by the refinement stage to fix the artifacts. Our model-agnostic pipeline requires no test-time tuning or optimization. It automatically enhances image fidelity and reference identity in the generated image, generalizing well to existing models on various tasks including but not limited to customization, generative compositing, view synthesis, and virtual tryon. Extensive experiments and comparisons demonstrate that our pipeline greatly 1 INTRODUCTION Generative models (Goodfellow et al., 2014;Karras et al., 2019;He & Aliaga, 2024;2023) for image synthesis (Ho et al., 2020;Rombach et al., 2022;Peebles & Xie, 2023;Podell et al., 2023;Luo et al., 2023b) have made significant advancement. Moreover, the traditional task of referenceguided image generation (Sheng et al., 2022;2023;2024) has been enabled by recent diffusion models (DM), where a text and/or visual prompt is provided and the subject object is generated in a specified novel context. This ability has been widely applied to applications such as subject customization (Ruiz et al., 2023a), object composition (Chen et al., 2023), novel view synthesis (Shi et al., 2023b) and virtual try-on (Choi et al., 2024). While these works seek generation in a single step, in practice undesired blemishes, detail omissions, and blurriness may occur in the generated images. These localized unpleasant anomalies are typically called artifacts as perceived by human eyes (e.g., "artifacts" row in Figure 1 (Zhang et al., 2023b)). The artifacts reduce image fidelity and the overall prompt-alignment quality of the synthesized images. Thus, a localized refinement tool to remove or reduce artifacts is beneficial.
Recently, a few limited approaches to artifact detection and refinement have been presented. PAL (Zhang et al., 2023b) presents an early work in this area that trains an artifact detection model in a supervised end-to-end manner using input images with artifacts and corresponding ground truth images. The detected artifacts can be partially removed using a pre-trained image inpainting tool. As PAL identifies, the artifacts are typically very small and irregular-shaped (Appendix Fig. 8), which further complicates the refinement process. PAL ameliorates artifacts but struggles with diversity of artifact refinement. Lack of diversity is also observed in RealisHuman (Wang et al., 2024a) that focuses on artifacts in human hands and faces, and the SynArtifact (Cao et al., 2024) using visionlanguage model and reinforcement learning for artifact annotation and removal. In general, none of these methods are able to provide a controllable and predictable artifact refinement output with free-form support automatically to precisely preserve the original identity details.
Our main approach is to leverage the identity detail info in the reference-image to guide the refine the artifacts. This provides a locally-controllable output (i.e., we specify the desired refinement), works for arbitrarily-shaped artifacts, preserves the identity and background in the provided generated image, and is applicable to multiple image generation approaches. As shown in Fig. 1, we provide reference-guided artifacts refinement: Given a generated image with marked artifact regions, and a reference image (containing a reference object), our model refines the artifacts by transferring corresponding details from the reference image to the artifact regions in the originally generated image. Our reference-guided approach shares insights with many reference-based image customization models (Song et al., 2023;Chen et al., 2023;Yang et al., 2023), but those models overlooked Figure 2: Comparisons of our region-matching method with keypoint matching. We utilize DIFT (Tang et al., 2023) and DHF (Luo et al., 2023a) to perform keypoint matching from the artifacts region (10 points are sampled along the artifacts contour) to the reference. DIFT and DHF often fail to find the accurate corresponding region; in addition, they have trouble in distinguishing between repeating patterns such as (a)(c). In contrast,
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2acc9a1a-3ff2-4e6c-be1b-4bd784f259deCited by top-tier papers3
- Tuning Timestep-Distilled Diffusion Model Using Pairwise Sample OptimizationZichen Miao, Zhengyuan Yang, Kevin Lin, Ze Wang et al.ICLR 2025
- Large Convolutional Model Tuning via Filter SubspaceWei Chen, Zichen Miao, Qiang QiuICLR 2025
- Coeff-Tuning: A Graph Filter Subspace View for Tuning Attention-Based Large ModelsZichen Miao, Wei Chen, Qiang QiuCVPR 2025
Builds on45
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- OmniPortrait: Fine-Grained Personalized Portrait Synthesis via Pivotal OptimizationDongxu Yue, Bo Lin, Yao Tang, Jiajun Liang et al.ICLR 2026
- PFStorer: Personalized Face Restoration and Super-ResolutionTuomas Varanka, Tapani Toivonen, Soumya Tripathy, Guoying Zhao et al.CVPR 2024 · 13 citations
- Direct Consistency Optimization for Robust Customization of Text-to-Image Diffusion modelsKyungmin Lee, Sangkyung Kwak, Kihyuk Sohn, Jinwoo ShinNeurIPS 2024 · 13 citations
- No Way To Steal My Face: Proactive Defense Against Identity-Preserving Personalized GenerationLizhi Xiong, Jun Li, Ziqiang Li, Weiwei Jiang et al.CVPR 2026 · 1 citation
- Face2Diffusion for Fast and Editable Face PersonalizationKaede Shiohara, Toshihiko YamasakiCVPR 2024 · 13 citations
