Lune

NeurIPS2020Top-tier venue

A Universal Approximation Theorem of Deep Neural Networks for Expressing Probability Distributions

Yulong Lu, Jianfeng Lu

2020Year
146Citations
21Top-tier citations

Abstract

This paper studies the universal approximation property of deep neural networks for representing probability distributions. Given a target distribution ππ and a source distribution pzp_z both defined on Rd\mathbb{R}^d, we prove under some assumptions that there exists a deep neural network g:Rd→Rg:\mathbb{R}^d\rightarrow \mathbb{R} with ReLU activation such that the push-forward measure (∇g)#pz(\nabla g)_\# p_z of pzp_z under the map ∇g\nabla g is arbitrarily close to the target measure ππ. The closeness are measured by three classes of integral probability metrics between probability distributions: 11-Wasserstein distance, maximum mean distance (MMD) and kernelized Stein discrepancy (KSD). We prove upper bounds for the size (width and depth) of the deep neural network in terms of the dimension dd and the approximation error ε\varepsilon with respect to the three discrepancies. In particular, the size of neural network can grow exponentially in dd when 11-Wasserstein distance is used as the discrepancy, whereas for both MMD and KSD the size of neural network only depends on dd at most polynomially. Our proof relies on convergence estimates of empirical measures under aforementioned discrepancies and semi-discrete optimal transport.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 8ed5caad-8104-4e55-b9b8-6afa48abdf67

Cited by top-tier papers21

Ask how each one uses it

Builds on1

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines