DiScoFormer: Plug-In Density and Score Estimation with Transformers
Vasily Ilin, Peter Sushko, Ranjay Krishna
摘要
Estimating probability density and its score from samples remains a core problem in generative modeling, Bayesian inference, and kinetic theory. Existing methods are bifurcated: classical kernel density estimators (KDE) generalize across distributions but suffer from the curse of dimensionality, while neural score-matching models achieve high precision but require retraining for every target distribution. We introduce DiScoFormer (Density and Score Transformer), an equivariant Transformer that maps i.i.d. samples to both density values and score vectors. Unlike score matching, which learns a fixed function R d → R d for a single distribution, DiScoFormer learns a sequence-tosequence operator that generalizes across distributions and sample sizes without retraining. Analytically, we prove that self-attention can recover normalized KDE, establishing it as a functional generalization of kernel methods; empirically, individual attention heads learn multi-scale, kernellike behaviors. The model outperforms KDE for density and score estimation, and provides a plugin score oracle for score-debiased KDE, Fisher information computation, and Fokker-Planck-type PDEs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 被引用 3,959 次
- Fourier Neural Operator for Parametric Partial Differential EquationsZongyi Li, Nikola Borislavov Kovachki, Kamyar Azizzadenesheli, Burigede Liu 等ICLR 2021 · 被引用 3,911 次
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 被引用 2,665 次
相关 Paper
- Nonparametric Score EstimatorsYuhao Zhou, Jiaxin Shi, Jun ZhuICML 2020 · 被引用 30 次
- KDEformer: Accelerating Transformers via Kernel Density EstimationAmir Zandieh, Insu Han, Majid Daliri, Amin KarbasiICML 2023 · 被引用 55 次
- All-in-one simulation-based inferenceManuel Glöckler, Michael Deistler, Christian Dietrich Weilbach, Frank Wood 等ICML 2024 · 被引用 74 次
- Rethinking Attention in Spiking Transformers: Overcoming Density Bias with Set SimilarityJinGyo Lim, Seunggyu Jeong, Seong-Eun KimICML 2026
- Particle Denoising Diffusion SamplerAngus Phillips, Hai-Dang Dau, Michael John Hutchinson, Valentin De Bortoli 等ICML 2024 · 被引用 60 次
