FlowFixer: Towards Detail-Preserving Subject-Driven Generation
Jinyoung Jun, Won-Dong Jang, Wenbin Ouyang, Raghudeep Gadde, Jungbeom Lee
摘要
We present FlowFixer, a refinement framework for subject-driven generation (SDG) that restores fine details lost during generation caused by changes in scale and perspective of a subject. FlowFixer proposes direct image-to-image translation from visual references, avoiding ambiguities in language prompts. To enable image-to-image training, we introduce a one-step denoising scheme to generate self-supervised training data, which automatically removes high-frequency details while preserving global structure, effectively simulating real-world SDG errors. We further propose a keypoint matching-based metric to properly assess fidelity in details beyond semantic similarities usually measured by CLIP or DINO. Experimental results demonstrate that FlowFixer outperforms state-of-the-art SDG methods in both qualitative and quantitative evaluations, setting a new benchmark for high-fidelity subject-driven generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper41
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
相关 Paper
- Mind-the-Glitch: Visual Correspondence for Detecting Inconsistencies in Subject-Driven GenerationAbdelrahman Eldesokey, Aleksandar Cvejic, Bernard Ghanem, Peter WonkaNeurIPS 2025 · 被引用 6 次
- Self-Evaluation Unlocks Any-Step Text-to-Image GenerationXin Yu, Xiaojuan Qi, Zhengqi Li, Kai Zhang 等CVPR 2026 · 被引用 3 次
- DiffSim: Taming Diffusion Models for Evaluating Visual SimilarityYiren Song, Xiaokang Liu, Mike Zheng ShouICCV 2025 · 被引用 1 次
- SPatchGAN: A Statistical Feature Based Discriminator for Unsupervised Image-to-Image TranslationXuning Shao, Weidong ZhangICCV 2021 · 被引用 34 次
- Self-Supervised Flow Matching for Scalable Multi-Modal SynthesisHila Chefer, Patrick Esser, Dominik Lorenz, Dustin Podell 等ICML 2026 · 被引用 13 次
