From Softmax to Score: Transformers Can Effectively Implement In-Context Denoising Steps
Paul Rosu, Lawrence Carin, Xiang Cheng
摘要
Transformers have emerged as powerful meta-learners, with growing evidence that they implement learning algorithms within their forward pass. We study this phenomenon in the context of denoising, presenting a unified framework that shows Transformers can implement (a) manifold denoising via Laplacian flows, (b) score-based denoising from diffusion models, and (c) a generalized form of anisotropic diffusion denoising. Our theory establishes exact equivalence between Transformer attention updates and these algorithms. Empirically, we validate these findings on image denoising tasks, showing that even simple Transformers can perform robust denoising both with and without context. These results illustrate the Transformer's flexibility as a denoising meta-learner. Code available at https://github.com/paulrosu11/Transformers_are_ Diffusion_Denoisers
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 被引用 11,724 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Diffusion-LM Improves Controllable Text GenerationXiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang 等NeurIPS 2022 · 被引用 1,546 次
相关 Paper
- Transformers Learn the Optimal DDPM Denoiser for Multi-Token GMMsHongkang Li, Hancheng Min, Rene VidalICML 2026 · 被引用 1 次
- Masked Image Training for Generalizable Deep Image DenoisingHaoyu Chen, Jinjin Gu, Yihao Liu, Salma Abdel Magid 等CVPR 2023
- Towards a Mechanistic Explanation of Diffusion Model GeneralizationMatthew Niedoba, Berend Zwartsenberg, Kevin Patrick Murphy, Frank WoodICML 2025
- On Inductive Biases That Enable Generalization in Diffusion TransformersJie An, De Wang, Pengsheng Guo, Jiebo Luo 等NeurIPS 2025 · 被引用 1 次
- Normalization-equivariant Diffusion Models: Learning Posterior Samplers From Noisy And Partial MeasurementsBrett Levac, Jon Tamir, Marcelo Pereyra, Julián TachellaICML 2026 · 被引用 2 次
