The Drift Kernel: Why Diffusion Models Change Even When Told Not To
Gokul Srinath Seetha Ram, Rashmi Elavazhagan
Abstract
Even when told to "do nothing," modern diffusion models subtly alter their output relative to the input they are supposed to preserve. We call this effect No-Op Drift. We introduce the Drift Kernel K M (σ) = E[∥I ′ -I 0 ∥ 2 2 | σ]-the expected perceptual deviation induced by running a diffusion model at noise strength σ under a null instructionand measure it across 120,000 baseline samples (30,000 from each of four models: SD15, SD21, SDXL, and In-structPix2Pix) and 9,600 ablation samples (4,800 with null prompts and 4,800 with copy prompts) from four architectures. We derive the quadratic form K M (σ) ≈ k M σ 2 + c M where k M = Tr(J D J ⊤ D ) from first principles via Taylor expansion of the decoder. To demonstrate that this structure is mechanistic rather than dataset-specific, we construct synthetic decoders with analytically controlled Jacobians: linear decoders recover the ideal K(σ) = Tr(AA ⊤ )σ 2 with R 2 = 0.9999, curved decoders retain the quadratic trend with Jacobian-dependent offsets, and edit-biased decoders flatten the relationship, reproducing the two regimes seen in practice. Across SD15, SD21, and SDXL, the drift kernel is well-approximated by K M (σ) ≈ k M σ 2 + c M (aggregate R 2 = 0.97 despite per-model fits R 2 = 0.10-0.26). Instruction-tuned InstructPix2Pix exhibits a distinct edit-driven kernel: flat in mean but highly variant. Drift kernels therefore provide a structural abstraction for why diffusion models cannot preserve identity under "do nothing" prompts. We release NoOp-Bench, a comprehensive benchmark with 10,000 inputs generating 120,000 baseline outputs (30K from each of four models) and 100 inputs generating 9,600 ablation samples, along with code to support reproducible kernel estimation and future analysis. Additional theory, proofs, ablations, LPIPS/CLIP metrics, and extended visualizations are provided in the Supplementary Material.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c814e360-4e78-4b2e-97f2-b491f57e9f8cBuilds on21
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- The Blessing of Randomness: SDE Beats ODE in General Diffusion-based Image EditingShen Nie, Hanzhong Allan Guo, Cheng Lu, Yuhao Zhou et al.ICLR 2024 · 64 citations
- Golden Noise for Diffusion Models: A Learning FrameworkZikai Zhou, Shitong Shao, Lichen Bai, Shufei Zhang et al.ICCV 2025 · 9 citations
- Instruct-CLIP: Improving Instruction-Guided Image Editing with Automated Data Refinement Using Contrastive LearningSherry X. Chen, Misha Sra, Pradeep SenCVPR 2025
- Decouple and Track: Benchmarking and Improving Video Diffusion Transformers for Motion TransferQingyu Shi, Jianzong Wu, Jinbin Bai, Jiangning Zhang et al.ICCV 2025 · 1 citation
- From Inpainting to Editing: Unlocking Robust Mask-Free Visual Dubbing via Generative BootstrappingXu He, Haoxian Zhang, Hejia Chen, Changyuan Zheng et al.ICML 2026
