CtrLoRA: An Extensible and Efficient Framework for Controllable Image Generation
Yifeng Xu, Zhenliang He, Shiguang Shan, Xilin Chen
摘要
Recently, large-scale diffusion models have made impressive progress in text-toimage (T2I) generation. To further equip these T2I models with fine-grained spatial control, approaches like ControlNet introduce an extra network that learns to follow a condition image. However, for every single condition type, Con-trolNet requires independent training on millions of data pairs with hundreds of GPU hours, which is quite expensive and makes it challenging for ordinary users to explore and develop new types of conditions. To address this problem, we propose the CtrLoRA framework, which trains a Base ControlNet to learn the common knowledge of image-to-image generation from multiple base conditions, along with condition-specific LoRAs to capture distinct characteristics of each condition. Utilizing our pretrained Base ControlNet, users can easily adapt it to new conditions, requiring as few as 1,000 data pairs and less than one hour of single-GPU training to obtain satisfactory results in most scenarios. Moreover, our CtrLoRA reduces the learnable parameters by 90% compared to ControlNet, significantly lowering the threshold to distribute and deploy the model weights. Extensive experiments on various types of conditions demonstrate the efficiency and effectiveness of our method. Codes and model weights will be released at https://github.com/xyfJASON/ctrlora.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- BideDPO: Conditional Image Generation with Simultaneous Text and Condition AlignmentDewei Zhou, Mingwei Li, Zongxin Yang, Yu Lu 等ICLR 2026 · 被引用 10 次
- DreamRenderer: Taming Multi-Instance Attribute Control in Large-Scale Text-to-Image ModelsDewei Zhou, Mingwei Li, Zongxin Yang, Yi YangICCV 2025 · 被引用 5 次
- FullDiT: Video Generative Foundation Models with Multimodal Control via Full AttentionXuan Ju, Weicai Ye, Quande Liu, Qiulin Wang 等ICCV 2025 · 被引用 5 次
- DivControl: Knowledge Diversion for Controllable Image GenerationYucheng Xie, Fu Feng, Ruixiao Shi, Jing Wang 等AAAI 2026 · 被引用 4 次
- Knowledge Diversion for Efficient Morphology Control and Policy TransferFu Feng, Ruixiao Shi, Yucheng Xie, Jianlu Shen 等ICML 2026 · 被引用 1 次
它引用的顶会 Paper24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- Simplifying Control Mechanism in Text-to-Image Diffusion ModelsZhida Feng, Li Chen, Yuenan Sun, Jiaxiang Liu 等AAAI 2025
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- Ctrl-Adapter: An Efficient and Versatile Framework for Adapting Diverse Controls to Any Diffusion ModelHan Lin, Jaemin Cho, Abhay Zala, Mohit BansalICLR 2025
- ModuleTeam: Open-Set Multi-Conditional Image Generation with Training-Free Latent Mixture of Any Control ModuleYuwei Zhou, Xin Wang, Hong Chen, Yipeng Zhang 等ACM MM 2025
- FreeControl: Training-Free Spatial Control of Any Text-to-Image Diffusion Model with Any ConditionSicheng Mo, Fangzhou Mu, Kuan Heng Lin, Yanli Liu 等CVPR 2024 · 被引用 31 次
