A Universal Approximation Theorem of Deep Neural Networks for Expressing Probability Distributions
Yulong Lu, Jianfeng Lu
摘要
This paper studies the universal approximation property of deep neural networks for representing probability distributions. Given a target distribution and a source distribution both defined on , we prove under some assumptions that there exists a deep neural network with ReLU activation such that the push-forward measure of under the map is arbitrarily close to the target measure . The closeness are measured by three classes of integral probability metrics between probability distributions: -Wasserstein distance, maximum mean distance (MMD) and kernelized Stein discrepancy (KSD). We prove upper bounds for the size (width and depth) of the deep neural network in terms of the dimension and the approximation error with respect to the three discrepancies. In particular, the size of neural network can grow exponentially in when -Wasserstein distance is used as the discrepancy, whereas for both MMD and KSD the size of neural network only depends on at most polynomially. Our proof relies on convergence estimates of empirical measures under aforementioned discrepancies and semi-discrete optimal transport.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- Learning Linear Causal Representations from Interventions under General Nonlinear MixingSimon Buchholz, Goutham Rajendran, Elan Rosenfeld, Bryon Aragam 等NeurIPS 2023 · 被引用 113 次
- Identifiability of deep generative models without auxiliary informationBohdan Kivva, Goutham Rajendran, Pradeep Ravikumar, Bryon AragamNeurIPS 2022 · 被引用 87 次
- Improving Diffusion-Based Image Synthesis with Context PredictionLing Yang, Jingwei Liu, Shenda Hong, Zhilong Zhang 等NeurIPS 2023 · 被引用 70 次
- Graph Auto-Encoder via Neighborhood Wasserstein ReconstructionMingyue Tang, Pan Li, Carl YangICLR 2022 · 被引用 68 次
- Normalizing flow neural networks by JKO schemeChen Xu, Xiuyuan Cheng, Yao XieNeurIPS 2023 · 被引用 51 次
它引用的顶会 Paper1
相关 Paper
- Neural Wasserstein Gradient Flows for Discrepancies with Riesz KernelsFabian Altekrüger, Johannes Hertrich, Gabriele SteidlICML 2023 · 被引用 15 次
- Distribution Regression with Sliced Wasserstein KernelsDimitri Meunier, Massimiliano Pontil, Carlo CilibertoICML 2022 · 被引用 24 次
- Constructive Universal High-Dimensional Distribution Generation through Deep ReLU NetworksDmytro Perekrestenko, Stephan Müller, Helmut BölcskeiICML 2020 · 被引用 15 次
- How Many Neurons Does it Take to Approximate the Maximum?Itay Safran, Daniel Reichman, Paul ValiantSODA 2024 · 被引用 3 次
- Quantitative Universal Approximation Bounds for Deep Belief NetworksJulian Sieber, Johann GehringerICML 2023 · 被引用 2 次
