CMT: Mid-Training for Efficient Learning of Consistency, Mean Flow, and Flow-Map Models
Zheyuan Hu, Chieh-Hsin Lai, Yuki Mitsufuji, Stefano Ermon
Abstract
Flow map models such as Consistency Models (CM) and Mean Flow (MF) enable few-step generation by learning the long jump of the ODE solution of diffusion models, yet training remains unstable, sensitive to hyperparameters, and costly. Initializing from a pre-trained diffusion model helps, but still requires converting infinitesimal steps into a long-jump map, leaving instability unresolved. We introduce mid-training, the first concept and practical method that inserts a lightweight intermediate stage between the (diffusion) pre-training and the final flow map training (i.e., post-training) for vision generation. Concretely, Consistency Mid-Training (CMT) is a compact and principled stage that trains a model to map points along a solver trajectory from a pre-trained model, starting from a prior sample, directly to the solver-generated clean sample. It yields a trajectory-consistent and stable initialization. This initializer outperforms random and diffusion-based baselines and enables fast, robust convergence without heuristics. Initializing post-training with CMT weights further simplifies flow map learning. Empirically, CMT achieves state-of-the-art two-step FIDs of 1.97 (CIFAR-10), 1.32 (ImageNet ), and 1.84 (ImageNet ), using up to % less training data and GPU time than CMs. On ImageNet , it attains 1-step FID 3.34 with % less training than MF from scratch (FID 3.43). On MSCOCO T2I, CMT reaches the best FID with % less training. This establishes CMT as a principled, efficient, and general framework for training flow map models. Code and models are available at https://github.com/sony/cmt.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 738a9f74-ce6e-4869-bfc5-c8dd1b7ae3ccCited by top-tier papers6
- Improved Mean Flows: On the Challenges of Fastforward Generative ModelsZhengyang Geng, Yiyang Lu, Zongze Wu, Eli Shechtman et al.CVPR 2026 · 116 citations
- Learning Straight Flows: Variational Flow Matching for Efficient GenerationChenrui Ma, Xi Xiao, Tianyang Wang, Xiao Wang et al.CVPR 2026 · 9 citations
- MeanFlow Transformers with Representation AutoencodersZheyuan Hu, Chieh-Hsin Lai, Ge Wu, Yuki Mitsufuji et al.CVPR 2026 · 6 citations
- Understanding, Accelerating, and Improving MeanFlow TrainingJin-Young Kim, Hyojun Go, Lea Bogensperger, Julius Erbach et al.CVPR 2026 · 4 citations
- Self-Evaluation Unlocks Any-Step Text-to-Image GenerationXin Yu, Xiaojuan Qi, Zhengqi Li, Kai Zhang et al.CVPR 2026 · 3 citations
Builds on35
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 3,959 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
Related papers
- Flow Map Learning Via Non-Gradient Vector FlowMark Goldstein, Anshuk Uppal, Raghav Singhal, Aahlad Manas Puli et al.ICLR 2026
- FACM: Flow-Anchored Consistency ModelsYansong Peng, Kai Zhu, Yu Liu, Pingyu Wu et al.ICLR 2026 · 16 citations
- Truncated Consistency ModelsSangyun Lee, Yilun Xu, Tomas Geffner, Giulia Fanti et al.ICLR 2025
- Align Your Flow: Scaling Continuous-Time Flow Map DistillationAmirmojtaba Sabour, Sanja Fidler, Karsten KreisNeurIPS 2025 · 91 citations
- Decoupled MeanFlow: Turning Flow Models into Flow Maps for Accelerated SamplingKyungmin Lee, Sihyun Yu, Jinwoo ShinICLR 2026 · 18 citations
