FreeControl: Training-Free Spatial Control of Any Text-to-Image Diffusion Model with Any Condition
Sicheng Mo, Fangzhou Mu, Kuan Heng Lin, Yanli Liu, Bochen Guan, Yin Li, Bolei Zhou
摘要
Recent approaches such as ControlNet [59] offer users fine-grained spatial control over text-to-image (T2I) diffusion models. However, auxiliary modules have to be trained for each spatial condition type, model architecture, and checkpoint, putting them at odds with the diverse intents and preferences a human designer would like to convey to the AI models during the content creation process. In this work, we present FreeControl, a training-free approach for controllable T2I generation that supports multiple conditions, architectures, and checkpoints simultaneously. Free Control enforces structure guidance to facilitate the global alignment with a guidance image, and appearance guidance to collect visual details from images generated without control. Extensive qualitative and quantitative experiments demonstrate the superior performance of Free Control across a variety of pre-trained T2I models. In particular, FreeControl enables convenient training-free control over many different architectures and checkpoints, allows the challenging input conditions on which most of the existing training-free methods fail, and achieves competitive synthesis quality compared to training-based approaches. Project page: https://genforce.github.io/freecontrol/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper32
- Video Diffusion Models are Training-free Motion Interpreter and ControllerZeqi Xiao, Yifan Zhou, Shuai Yang, Xingang PanNeurIPS 2024 · 被引用 71 次
- Attentive Eraser: Unleashing Diffusion Model's Object Removal Potential via Self-Attention Redirection GuidanceWenhao Sun, Xue-Mei Dong, Benlei Cui, Jingqun TangAAAI 2025 · 被引用 50 次
- Understanding and Improving Training-free Loss-based Diffusion GuidanceYifei Shen, Xinyang Jiang, Yifan Yang, Yezhen Wang 等NeurIPS 2024 · 被引用 36 次
- SpotActor: Training-Free Layout-Controlled Consistent Image GenerationJiahao Wang, Caixia Yan, Weizhan Zhang, Haonan Lin 等AAAI 2025 · 被引用 13 次
- SceneDesigner: Controllable Multi-Object Image Generation with 9-DoF Pose ManipulationZhenyuan Qin, Xincheng Shuai, Henghui DingNeurIPS 2025 · 被引用 11 次
它引用的顶会 Paper40
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- Ctrl-X: Controlling Structure and Appearance for Text-To-Image Generation Without GuidanceKuan Heng Lin, Sicheng Mo, Ben Klingher, Fangzhou Mu 等NeurIPS 2024 · 被引用 51 次
- FreeControl: Efficient, Training-Free Structural Control via One-Step Attention ExtractionJiang Lin, Xinyu Chen, Song Wu, Zhiqiu Zhang 等NeurIPS 2025 · 被引用 3 次
- Dual Recursive Feedback on Generation and Appearance Latents for Pose-Robust Text-to-Image DiffusionJiwon Kim, Pu-Reum Kim, Seonhwa Kim, Soobin Park 等ICCV 2025
- Cross-ControlNet: Training-Free Fusion of Multiple Conditions for Text-to-Image GenerationXiang Liu, Junjun Jiang, Wei Han, Kui Jiang 等ICLR 2026
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
