A Universal Approximation Theorem of Deep Neural Networks for Expressing Probability Distributions
Yulong Lu, Jianfeng Lu
Abstract
This paper studies the universal approximation property of deep neural networks for representing probability distributions. Given a target distribution and a source distribution both defined on , we prove under some assumptions that there exists a deep neural network with ReLU activation such that the push-forward measure of under the map is arbitrarily close to the target measure . The closeness are measured by three classes of integral probability metrics between probability distributions: -Wasserstein distance, maximum mean distance (MMD) and kernelized Stein discrepancy (KSD). We prove upper bounds for the size (width and depth) of the deep neural network in terms of the dimension and the approximation error with respect to the three discrepancies. In particular, the size of neural network can grow exponentially in when -Wasserstein distance is used as the discrepancy, whereas for both MMD and KSD the size of neural network only depends on at most polynomially. Our proof relies on convergence estimates of empirical measures under aforementioned discrepancies and semi-discrete optimal transport.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8ed5caad-8104-4e55-b9b8-6afa48abdf67Cited by top-tier papers21
- Learning Linear Causal Representations from Interventions under General Nonlinear MixingSimon Buchholz, Goutham Rajendran, Elan Rosenfeld, Bryon Aragam et al.NeurIPS 2023 · 113 citations
- Identifiability of deep generative models without auxiliary informationBohdan Kivva, Goutham Rajendran, Pradeep Ravikumar, Bryon AragamNeurIPS 2022 · 87 citations
- Improving Diffusion-Based Image Synthesis with Context PredictionLing Yang, Jingwei Liu, Shenda Hong, Zhilong Zhang et al.NeurIPS 2023 · 70 citations
- Graph Auto-Encoder via Neighborhood Wasserstein ReconstructionMingyue Tang, Pan Li, Carl YangICLR 2022 · 68 citations
- Normalizing flow neural networks by JKO schemeChen Xu, Xiuyuan Cheng, Yao XieNeurIPS 2023 · 51 citations
Builds on1
Related papers
- Neural Wasserstein Gradient Flows for Discrepancies with Riesz KernelsFabian Altekrüger, Johannes Hertrich, Gabriele SteidlICML 2023 · 15 citations
- Distribution Regression with Sliced Wasserstein KernelsDimitri Meunier, Massimiliano Pontil, Carlo CilibertoICML 2022 · 24 citations
- Constructive Universal High-Dimensional Distribution Generation through Deep ReLU NetworksDmytro Perekrestenko, Stephan Müller, Helmut BölcskeiICML 2020 · 15 citations
- How Many Neurons Does it Take to Approximate the Maximum?Itay Safran, Daniel Reichman, Paul ValiantSODA 2024 · 3 citations
- Quantitative Universal Approximation Bounds for Deep Belief NetworksJulian Sieber, Johann GehringerICML 2023 · 2 citations
