Improved Training Technique for Latent Consistency Models
Quan Dao, Khanh Doan, Di Liu, Trung Le, Dimitris N. Metaxas
摘要
Consistency models are a new family of generative models capable of producing high-quality samples in either a single step or multiple steps. Recently, consistency models have demonstrated impressive performance, achieving results on par with diffusion models in the pixel space. However, the success of scaling consistency training to large-scale datasets, particularly for text-to-image and video generation tasks, is determined by performance in the latent space. In this work, we analyze the statistical differences between pixel and latent spaces, discovering that latent data often contains highly impulsive outliers, which significantly degrade the performance of iCT in the latent space. To address this, we replace Pseudo-Huber losses with Cauchy losses, effectively mitigating the impact of outliers. Additionally, we introduce a diffusion loss at early timesteps and employ optimal transport (OT) coupling to further enhance performance. Lastly, we introduce the adaptive scaling- scheduler to manage the robust training process and adopt Non-scaling LayerNorm in the architecture to better capture the statistics of the features and reduce outlier impact. With these strategies, we successfully train latent consistency models capable of high-quality sampling with one or two steps, significantly narrowing the performance gap between latent consistency and diffusion models. The implementation is released here: https://github.com/quandao10/sLCT/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- SoFlow: Solution Flow Models for One-Step Generative ModelingTianze Luo, Haotian Yuan, Zhuang LiuICLR 2026 · 被引用 18 次
- An Optimal Transport-driven Approach for Cultivating Latent Space in Online Incremental LearningQuyen Tran, Hai Nguyen, Minh Quan Dao, Hoang Phan 等CVPR 2026
- Multi-Patch Global-to-Local Transformer Architecture For Efficient Flow Matching and Diffusion ModelQuan Dao, Dimitris N. MetaxasCVPR 2026
- LUCAS: Layered Universal Codec AvatarsDi Liu, Teng Deng, Giljoo Nam, Yu Rong 等CVPR 2025
它引用的顶会 Paper35
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 被引用 3,959 次
相关 Paper
- Improved Techniques for Training Consistency ModelsYang Song, Prafulla DhariwalICLR 2024 · 被引用 383 次
- Truncated Consistency ModelsSangyun Lee, Yilun Xu, Tomas Geffner, Giulia Fanti 等ICLR 2025
- Phased Consistency ModelsFu-Yun Wang, Zhaoyang Huang, Alexander William Bergman, Dazhong Shen 等NeurIPS 2024 · 被引用 86 次
- There is No VAE: End-to-End Pixel-Space Generative Modeling via Self-Supervised Pre-TrainingJiachen Lei, Keli Liu, Julius Berner, Y HoiM 等ICLR 2026 · 被引用 24 次
- Consistency Models Made EasyZhengyang Geng, Ashwini Pokle, Weijian Luo, Justin Lin 等ICLR 2025
