CMT: Mid-Training for Efficient Learning of Consistency, Mean Flow, and Flow-Map Models
Zheyuan Hu, Chieh-Hsin Lai, Yuki Mitsufuji, Stefano Ermon
摘要
Flow map models such as Consistency Models (CM) and Mean Flow (MF) enable few-step generation by learning the long jump of the ODE solution of diffusion models, yet training remains unstable, sensitive to hyperparameters, and costly. Initializing from a pre-trained diffusion model helps, but still requires converting infinitesimal steps into a long-jump map, leaving instability unresolved. We introduce mid-training, the first concept and practical method that inserts a lightweight intermediate stage between the (diffusion) pre-training and the final flow map training (i.e., post-training) for vision generation. Concretely, Consistency Mid-Training (CMT) is a compact and principled stage that trains a model to map points along a solver trajectory from a pre-trained model, starting from a prior sample, directly to the solver-generated clean sample. It yields a trajectory-consistent and stable initialization. This initializer outperforms random and diffusion-based baselines and enables fast, robust convergence without heuristics. Initializing post-training with CMT weights further simplifies flow map learning. Empirically, CMT achieves state-of-the-art two-step FIDs of 1.97 (CIFAR-10), 1.32 (ImageNet ), and 1.84 (ImageNet ), using up to % less training data and GPU time than CMs. On ImageNet , it attains 1-step FID 3.34 with % less training than MF from scratch (FID 3.43). On MSCOCO T2I, CMT reaches the best FID with % less training. This establishes CMT as a principled, efficient, and general framework for training flow map models. Code and models are available at https://github.com/sony/cmt.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Improved Mean Flows: On the Challenges of Fastforward Generative ModelsZhengyang Geng, Yiyang Lu, Zongze Wu, Eli Shechtman 等CVPR 2026 · 被引用 116 次
- Learning Straight Flows: Variational Flow Matching for Efficient GenerationChenrui Ma, Xi Xiao, Tianyang Wang, Xiao Wang 等CVPR 2026 · 被引用 9 次
- MeanFlow Transformers with Representation AutoencodersZheyuan Hu, Chieh-Hsin Lai, Ge Wu, Yuki Mitsufuji 等CVPR 2026 · 被引用 6 次
- Understanding, Accelerating, and Improving MeanFlow TrainingJin-Young Kim, Hyojun Go, Lea Bogensperger, Julius Erbach 等CVPR 2026 · 被引用 4 次
- Self-Evaluation Unlocks Any-Step Text-to-Image GenerationXin Yu, Xiaojuan Qi, Zhengqi Li, Kai Zhang 等CVPR 2026 · 被引用 3 次
它引用的顶会 Paper35
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 被引用 3,959 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
相关 Paper
- Flow Map Learning Via Non-Gradient Vector FlowMark Goldstein, Anshuk Uppal, Raghav Singhal, Aahlad Manas Puli 等ICLR 2026
- FACM: Flow-Anchored Consistency ModelsYansong Peng, Kai Zhu, Yu Liu, Pingyu Wu 等ICLR 2026 · 被引用 16 次
- Truncated Consistency ModelsSangyun Lee, Yilun Xu, Tomas Geffner, Giulia Fanti 等ICLR 2025
- Align Your Flow: Scaling Continuous-Time Flow Map DistillationAmirmojtaba Sabour, Sanja Fidler, Karsten KreisNeurIPS 2025 · 被引用 91 次
- Decoupled MeanFlow: Turning Flow Models into Flow Maps for Accelerated SamplingKyungmin Lee, Sihyun Yu, Jinwoo ShinICLR 2026 · 被引用 18 次
