From Softmax to Score: Transformers Can Effectively Implement In-Context Denoising Steps
Paul Rosu, Lawrence Carin, Xiang Cheng
Abstract
Transformers have emerged as powerful meta-learners, with growing evidence that they implement learning algorithms within their forward pass. We study this phenomenon in the context of denoising, presenting a unified framework that shows Transformers can implement (a) manifold denoising via Laplacian flows, (b) score-based denoising from diffusion models, and (c) a generalized form of anisotropic diffusion denoising. Our theory establishes exact equivalence between Transformer attention updates and these algorithms. Empirically, we validate these findings on image denoising tasks, showing that even simple Transformers can perform robust denoising both with and without context. These results illustrate the Transformer's flexibility as a denoising meta-learner. Code available at https://github.com/paulrosu11/Transformers_are_ Diffusion_Denoisers
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 907c996e-d255-4451-bf35-78b62af9bbf1Cited by top-tier papers1
Ask how each one uses itBuilds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Diffusion-LM Improves Controllable Text GenerationXiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang et al.NeurIPS 2022 · 1,546 citations
Related papers
- Transformers Learn the Optimal DDPM Denoiser for Multi-Token GMMsHongkang Li, Hancheng Min, Rene VidalICML 2026 · 1 citation
- Masked Image Training for Generalizable Deep Image DenoisingHaoyu Chen, Jinjin Gu, Yihao Liu, Salma Abdel Magid et al.CVPR 2023
- Towards a Mechanistic Explanation of Diffusion Model GeneralizationMatthew Niedoba, Berend Zwartsenberg, Kevin Patrick Murphy, Frank WoodICML 2025
- On Inductive Biases That Enable Generalization in Diffusion TransformersJie An, De Wang, Pengsheng Guo, Jiebo Luo et al.NeurIPS 2025 · 1 citation
- Normalization-equivariant Diffusion Models: Learning Posterior Samplers From Noisy And Partial MeasurementsBrett Levac, Jon Tamir, Marcelo Pereyra, Julián TachellaICML 2026 · 2 citations
