The Drift Kernel: Why Diffusion Models Change Even When Told Not To
Gokul Srinath Seetha Ram, Rashmi Elavazhagan
摘要
Even when told to "do nothing," modern diffusion models subtly alter their output relative to the input they are supposed to preserve. We call this effect No-Op Drift. We introduce the Drift Kernel K M (σ) = E[∥I ′ -I 0 ∥ 2 2 | σ]-the expected perceptual deviation induced by running a diffusion model at noise strength σ under a null instructionand measure it across 120,000 baseline samples (30,000 from each of four models: SD15, SD21, SDXL, and In-structPix2Pix) and 9,600 ablation samples (4,800 with null prompts and 4,800 with copy prompts) from four architectures. We derive the quadratic form K M (σ) ≈ k M σ 2 + c M where k M = Tr(J D J ⊤ D ) from first principles via Taylor expansion of the decoder. To demonstrate that this structure is mechanistic rather than dataset-specific, we construct synthetic decoders with analytically controlled Jacobians: linear decoders recover the ideal K(σ) = Tr(AA ⊤ )σ 2 with R 2 = 0.9999, curved decoders retain the quadratic trend with Jacobian-dependent offsets, and edit-biased decoders flatten the relationship, reproducing the two regimes seen in practice. Across SD15, SD21, and SDXL, the drift kernel is well-approximated by K M (σ) ≈ k M σ 2 + c M (aggregate R 2 = 0.97 despite per-model fits R 2 = 0.10-0.26). Instruction-tuned InstructPix2Pix exhibits a distinct edit-driven kernel: flat in mean but highly variant. Drift kernels therefore provide a structural abstraction for why diffusion models cannot preserve identity under "do nothing" prompts. We release NoOp-Bench, a comprehensive benchmark with 10,000 inputs generating 120,000 baseline outputs (30K from each of four models) and 100 inputs generating 9,600 ablation samples, along with code to support reproducible kernel estimation and future analysis. Additional theory, proofs, ablations, LPIPS/CLIP metrics, and extended visualizations are provided in the Supplementary Material.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper21
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
相关 Paper
- The Blessing of Randomness: SDE Beats ODE in General Diffusion-based Image EditingShen Nie, Hanzhong Allan Guo, Cheng Lu, Yuhao Zhou 等ICLR 2024 · 被引用 64 次
- Golden Noise for Diffusion Models: A Learning FrameworkZikai Zhou, Shitong Shao, Lichen Bai, Shufei Zhang 等ICCV 2025 · 被引用 9 次
- Instruct-CLIP: Improving Instruction-Guided Image Editing with Automated Data Refinement Using Contrastive LearningSherry X. Chen, Misha Sra, Pradeep SenCVPR 2025
- Decouple and Track: Benchmarking and Improving Video Diffusion Transformers for Motion TransferQingyu Shi, Jianzong Wu, Jinbin Bai, Jiangning Zhang 等ICCV 2025 · 被引用 1 次
- From Inpainting to Editing: Unlocking Robust Mask-Free Visual Dubbing via Generative BootstrappingXu He, Haoxian Zhang, Hejia Chen, Changyuan Zheng 等ICML 2026
