Improved Training Technique for Latent Consistency Models
Quan Dao, Khanh Doan, Di Liu, Trung Le, Dimitris N. Metaxas
Abstract
Consistency models are a new family of generative models capable of producing high-quality samples in either a single step or multiple steps. Recently, consistency models have demonstrated impressive performance, achieving results on par with diffusion models in the pixel space. However, the success of scaling consistency training to large-scale datasets, particularly for text-to-image and video generation tasks, is determined by performance in the latent space. In this work, we analyze the statistical differences between pixel and latent spaces, discovering that latent data often contains highly impulsive outliers, which significantly degrade the performance of iCT in the latent space. To address this, we replace Pseudo-Huber losses with Cauchy losses, effectively mitigating the impact of outliers. Additionally, we introduce a diffusion loss at early timesteps and employ optimal transport (OT) coupling to further enhance performance. Lastly, we introduce the adaptive scaling- scheduler to manage the robust training process and adopt Non-scaling LayerNorm in the architecture to better capture the statistics of the features and reduce outlier impact. With these strategies, we successfully train latent consistency models capable of high-quality sampling with one or two steps, significantly narrowing the performance gap between latent consistency and diffusion models. The implementation is released here: https://github.com/quandao10/sLCT/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- SoFlow: Solution Flow Models for One-Step Generative ModelingTianze Luo, Haotian Yuan, Zhuang LiuICLR 2026 · 18 citations
- An Optimal Transport-driven Approach for Cultivating Latent Space in Online Incremental LearningQuyen Tran, Hai Nguyen, Minh Quan Dao, Hoang Phan et al.CVPR 2026
- Multi-Patch Global-to-Local Transformer Architecture For Efficient Flow Matching and Diffusion ModelQuan Dao, Dimitris N. MetaxasCVPR 2026
- LUCAS: Layered Universal Codec AvatarsDi Liu, Teng Deng, Giljoo Nam, Yu Rong et al.CVPR 2025
Builds on35
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 3,959 citations
Related papers
- Improved Techniques for Training Consistency ModelsYang Song, Prafulla DhariwalICLR 2024 · 383 citations
- Truncated Consistency ModelsSangyun Lee, Yilun Xu, Tomas Geffner, Giulia Fanti et al.ICLR 2025
- Phased Consistency ModelsFu-Yun Wang, Zhaoyang Huang, Alexander William Bergman, Dazhong Shen et al.NeurIPS 2024 · 86 citations
- There is No VAE: End-to-End Pixel-Space Generative Modeling via Self-Supervised Pre-TrainingJiachen Lei, Keli Liu, Julius Berner, Y HoiM et al.ICLR 2026 · 24 citations
- Consistency Models Made EasyZhengyang Geng, Ashwini Pokle, Weijian Luo, Justin Lin et al.ICLR 2025
