One-Step Offline Distillation of Diffusion-based Models via Koopman Modeling
Nimrod Berman, Ilan Naiman, Moshe Eliasof, Hedi Zisling, Omri Azencot
Abstract
Diffusion-based generative models have demonstrated exceptional performance, yet their iterative sampling procedures remain computationally expensive. A prominent strategy to mitigate this cost is distillation, with offline distillation offering particular advantages in terms of efficiency, modularity, and flexibility. In this work, we identify two key observations that motivate a principled distillation framework: (1) while diffusion models have been viewed through the lens of dynamical systems theory, powerful and underexplored tools can be further leveraged; and (2) diffusion models inherently impose structured, semantically coherent trajectories in latent space. Building on these observations, we introduce the Koopman Distillation Model (KDM), a novel offline distillation approach grounded in Koopman theorya classical framework for representing nonlinear dynamics linearly in a transformed space. KDM encodes noisy inputs into an embedded space where a learned linear operator propagates them forward, followed by a decoder that reconstructs clean samples. This enables single-step generation while preserving semantic fidelity. We provide theoretical justification for our approach: (1) under mild assumptions, the learned diffusion dynamics admit a finite-dimensional Koopman representation; and (2) proximity in the Koopman latent space correlates with semantic similarity in the generated outputs, allowing for effective trajectory alignment. KDM achieves highly competitive performance across standard offline distillation benchmarks.
and conditional generation. These results demonstrate the effectiveness of our Koopman-based formulation in bridging the performance gap with online methods, offering a fast and high-fidelity offline diffusion model distillation. Our key contributions are:
- Observing diffusion dynamics. We identify, analyze and formalize two key observations: (i) tools from dynamical systems theory, specifically Koopman operator theory, remain underexplored in diffusion modeling and distillation; (ii) training induces coherent semantic organization in noise space. These insights motivate a structure-promoting dynamics-aware approach to distillation.
We develop a theoretical foundation proving the existence of finite-dimensional Koopman representations and semantic structure preservation under mild assumptions and introduce a simple, scalable encoder-lineardynamics-decoder architecture that preserves the teacher generation dynamics semantic structure. 3. Highly competitive results. We conduct a comprehensive unconditional and conditional evaluation showing that our method outperforms prior offline distillation approaches, achieving FID improvement, while enabling efficient, single-step generation. 2 Related Work Generative Modeling. Deep generative models have made significant progress in recent years, with diffusion models emerging as a leading framework for high-quality image, audio, and molecular generation. Building on principles from nonequilibrium thermodynamics and score matching [72, 76], these models define a forward process that gradually corrupts data with noise and learn to reverse it via a neural network-based denoising process. Variants like DDPMs [30] and score-based generative models [76] have shown competitive performance across a wide range of benchmarks, often surpassing VAEs and GANs [40, 27] in sample fidelity and diversity. Accelerating Diffusion Sampling. Diffusion models generate high-quality samples but are slow due to iterative sampling. One approach improves integration using higher-order solvers and expressive noise schedules [73, 49], achieving high-quality results in significantly fewer steps. Another class of approaches seeks to distill a student model from a pre-trained teacher, reducing the sampling process to a single or few steps. Online distillation techniques [50, 68] supervise the student directly with teacher predictions during training, but require continuous access to the full teacher model, resulting in high memory and compute overhead. Recent methods demonstrate high-quality one-step generation across CIFAR-10, ImageNet, and text-to-image benchmarks [86, 93], using objectives based on distribution matching [86], score identity [93], adversarial training [69, 71], and consistency models [75, 38] among others [13, 52, 62, 94, 6]. Offline distillation methods [24] rely on teacher-generated noise-data pairs and decouple student training from the teacher. Aside from the concurrent work [82], current distillation methods largely unexplored Koopman-based perspectives and their relation to the teacher's underlying dynamics.
Koopman-Based Modeling. Koopman operator theory provides a powerful framework for modeling nonlinear dynamical systems via linear operators acting on observable functions [42,67]. This view has been applied to tasks such as deep learning of dynamical systems [79,53,84,65,58,60], sequential disentanglement [7, 5], and control [29]. A relate
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3c53fa4d-e713-499e-8a0b-4311224c7f2fCited by top-tier papers5
- A Diffusion Model for Regular Time Series Generation from Irregular Data with Completion and MaskingGal Fadlon, Idan Arbiv, Nimrod Berman, Omri AzencotNeurIPS 2025 · 12 citations
- Towards General Modality Translation with Contrastive and Predictive Latent Diffusion BridgeNimrod Berman, Omkar Joglekar, Eitan Kosman, Dotan Di Castro et al.NeurIPS 2025 · 5 citations
- Who Said Neural Networks Aren't Linear?Nimrod Berman, Assaf Hallak, Assaf ShocherICML 2026 · 3 citations
- DiffSDA: Unsupervised Diffusion Sequential Disentanglement Across ModalitiesHedi Zisling, Ilan Naiman, Nimrod Berman, Supasorn Suwajanakorn et al.ICLR 2026 · 2 citations
- Unfolding Generative Flows with Koopman Operators: Trajectory-Preserving LinearizationErkan Turan, Ari Siozopoulos, Louis Martinez, Julien Gaubil et al.ICML 2026
Builds on47
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 5,234 citations
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 3,959 citations
Related papers
- Hierarchical Koopman Diffusion: Fast Generation with Interpretable Diffusion TrajectoryHanru Bai, Weiyang Ding, Difan ZouNeurIPS 2025 · 3 citations
- Data-free Distillation of Diffusion Models with BootstrappingJiatao Gu, Chen Wang, Shuangfei Zhai, Yizhe Zhang et al.ICML 2024 · 5 citations
- Simple Distillation for One-Step Diffusion ModelsHuaisheng Zhu, Teng Xiao, Shijie Zhou, Zhimeng Guo et al.NeurIPS 2025 · 7 citations
- EM Distillation for One-step Diffusion ModelsSirui Xie, Zhisheng Xiao, Diederik P. Kingma, Tingbo Hou et al.NeurIPS 2024 · 69 citations
- DKDM: Data-Free Knowledge Distillation for Diffusion Models with Any ArchitectureQianlong Xiang, Miao Zhang, Yuzhang Shang, Jianlong Wu et al.CVPR 2025
