VCT: Training Consistency Models with Variational Noise Coupling
Gianluigi Silvestri, Luca Ambrogioni, Chieh-Hsin Lai, Yuhta Takida, Yuki Mitsufuji
Abstract
Consistency Training (CT) has recently emerged as a strong alternative to diffusion models for image generation. However, non-distillation CT often suffers from high variance and instability, motivating ongoing research into its training dynamics. We propose Variational Consistency Training (VCT), a flexible and effective framework compatible with various forward kernels, including those in flow matching. Its key innovation is a learned noise-data coupling scheme inspired by Variational Autoencoders, where a data-dependent encoder models noise emission. This enables VCT to adaptively learn noise-todata pairings, reducing training variance relative to the fixed, unsorted pairings in classical CT. Experiments on multiple image datasets demonstrate significant improvements: our method surpasses baselines, achieves state-of-the-art FID among non-distillation CT approaches on CIFAR-10, and matches SoTA performance on ImageNet 64 × 64 with only two sampling steps. Code is available at https://github.com/sony/vct .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6f4e6967-36be-47ec-8952-0388f9aecd99Cited by top-tier papers2
- CMT: Mid-Training for Efficient Learning of Consistency, Mean Flow, and Flow-Map ModelsZheyuan Hu, Chieh-Hsin Lai, Yuki Mitsufuji, Stefano ErmonICLR 2026 · 21 citations
- Learning Straight Flows: Variational Flow Matching for Efficient GenerationChenrui Ma, Xi Xiao, Tianyang Wang, Xiao Wang et al.CVPR 2026 · 9 citations
Builds on39
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 3,959 citations
Related papers
- Improved Techniques for Training Consistency ModelsYang Song, Prafulla DhariwalICLR 2024 · 383 citations
- Consistency ModelsYang Song, Prafulla Dhariwal, Mark Chen, Ilya SutskeverICML 2023 · 1,720 citations
- Improving Consistency Models with Generator-Augmented FlowsThibaut Issenhuth, Sangchul Lee, Ludovic Dos Santos, Jean-Yves Franceschi et al.ICML 2025
- ACT-Diffusion: Efficient Adversarial Consistency Training for One-Step Diffusion ModelsFei Kong, Jinhao Duan, Lichao Sun, Hao Cheng et al.CVPR 2024
- See Further When Clear: Curriculum Consistency ModelYunpeng Liu, Boxiao Liu, Yi Zhang, Xingzhong Hou et al.CVPR 2025
