EasyControl: Adding Efficient and Flexible Control for Diffusion Transformer
Yuxuan Zhang, Yirui Yuan, Yiren Song, Haofan Wang, Jiaming Liu
摘要
Recent advancements in Unet-based diffusion models, such as ControlNet and IP-Adapter, have introduced effective spatial and subject control mechanisms. However, the DiT (Diffusion Transformer) architecture still struggles with efficient and flexible control. To tackle this issue, we propose EasyControl, a novel framework designed to unify condition-guided diffusion transformers with high efficiency and flexibility. Our framework is built on three key innovations. First, we introduce a lightweight Condition Injection LoRA Module. This module processes conditional signals in isolation, acting as a plug-and-play solution. It avoids modifying the base model's weights, ensuring compatibility with customized models and enabling the flexible injection of diverse conditions. Notably, this module also supports harmonious and robust zero-shot multi-condition generalization, even when trained only on single-condition data. Second, we propose a Position-Aware Training Paradigm. This approach standardizes input conditions to fixed reso-lutions, allowing the generation of images with arbitrary aspect ratios and flexible resolutions. At the same time, it optimizes computational efficiency, making the framework more practical for real-world applications. Third, we develop a Causal Attention Mechanism combined with the KV Cache technique, adapted for conditional generation tasks. This innovation significantly reduces the latency of image synthesis, improving the overall efficiency of the framework. Through extensive experiments, we demonstrate that EasyControl achieves exceptional performance across various application scenarios. These innovations collectively make our framework highly efficient, flexible, and suitable for a wide range of tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper56
- OmniConsistency: Learning Style-Agnostic Consistency from Paired Stylization DataYiren Song, Cheng Liu, Mike Zheng ShouNeurIPS 2025 · 被引用 33 次
- Follow-Your-Shape: Shape-Aware Image Editing via Trajectory-Guided Region ControlZeqian Long, Mingzhe Zheng, Kunyu Feng, Xinhua Zhang 等ICLR 2026 · 被引用 29 次
- All Patches Matter, More Patches Better: Enhance AI-Generated Image Detection via Panoptic Patch LearningZheng Yang, Ruoxin Chen, Zhiyuan Yan, Ke-Yue Zhang 等ICLR 2026 · 被引用 28 次
- UniTEX: Universal High Fidelity Generative Texturing for 3D ShapesYixun Liang, Kunming Luo, Xiao Chen, Rui Chen 等CVPR 2026 · 被引用 27 次
- LayerCraft: Enhancing Text-to-Image Generation with CoT Reasoning and Layered Object IntegrationYuyao Zhang, Jinghao Li, Yu-Wing TaiNeurIPS 2025 · 被引用 21 次
它引用的顶会 Paper55
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 被引用 11,724 次
相关 Paper
- Ctrl-Adapter: An Efficient and Versatile Framework for Adapting Diverse Controls to Any Diffusion ModelHan Lin, Jaemin Cho, Abhay Zala, Mohit BansalICLR 2025
- Condition-Aware Neural Network for Controlled Image GenerationHan Cai, Muyang Li, Qinsheng Zhang, Ming-Yu Liu 等CVPR 2024
- DivControl: Knowledge Diversion for Controllable Image GenerationYucheng Xie, Fu Feng, Ruixiao Shi, Jing Wang 等AAAI 2026 · 被引用 4 次
- CtrLoRA: An Extensible and Efficient Framework for Controllable Image GenerationYifeng Xu, Zhenliang He, Shiguang Shan, Xilin ChenICLR 2025
- Unicombine: Unified Multi-Conditional Combination with Diffusion TransformerHaoxuan Wang, Jinlong Peng, Qingdong He, Hao Yang 等ICCV 2025 · 被引用 6 次
