Pareto Variational Autoencoder
Mincheol Cho, Yedarm Seong, Joong-Ho Won
Abstract
This paper introduces a new class of multivariate power-law distributions-the symmetric Pareto (symPareto) distribution-which can be viewed as an ℓ 1 -norm-based counterpart of the multivariate t distribution, with the motivation of capturing the heavy tail of the target distribution in generative modeling and bringing robustness to noise in downstream tasks such as image denoising. The symPareto distribution possesses many attractive information-geometric properties with respect to the γ-power divergence that is a natural alternative to the Kullback-Leibler divergence, the core of the conventional variational autoencoder (VAE) models, for power families. Leveraging on the joint minimization view of variational inference, this paper proposes the ParetoVAE, a probabilistic autoencoder that minimizes the γ-power divergence between two statistical manifolds. ParetoVAE employs the symPareto distribution for both prior and encoder, with flexible decoder options including multivariate t and symPareto distributions. Empirical evidences demonstrate the effectiveness of ParetoVAE across multiple domains through varying the types of the decoder. The t decoder achieves superior performance in sparse, heavy-tailed data reconstruction and word frequency analysis; the symPareto decoder enables robust high-dimensional denoising. * Equal contribution. THEORETICAL BACKGROUND VARIATIONAL AUTOENCODER (VAE) VAE aims to approximate the true data distribution p data (x) by modeling the marginal likelihood as p θ (x) = p θ (x|z)p Z (z) dz, where z represents a latent variable. Due to the intractability of
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b37827c8-71f3-4622-b389-296be7de1eaaBuilds on7
- Pareto GAN: Extending the Representational Power of GANs to Heavy-Tailed DistributionsTodd Huster, Jeremy E. J. Cohen, Zinan Lin, Kevin Chan et al.ICML 2021 · 37 citations
- Marginal Tail-Adaptive Normalizing FlowsMike Laszkiewicz, Johannes Lederer, Asja FischerICML 2022 · 12 citations
- Coupled Variational AutoencoderXiaoran Hao, Patrick ShaftoICML 2023 · 7 citations
- Cauchy Diffusion: A Heavy-tailed Denoising Diffusion Probabilistic Model for Speech SynthesisQi Lian, Yu Qi, Yueming WangAAAI 2025 · 3 citations
- Flexible Tails for Normalizing FlowsTennessee Hickling, Dennis PrangleICML 2025
Related papers
- -Variational Autoencoder: Learning Heavy-tailed Data with Student's t and Power DivergenceJuno Kim, Jaehyuk Kwon, Mincheol Cho, Hyunjong Lee et al.ICLR 2024 · 11 citations
- Multiobjective distribution matchingXiaoyuan Zhang, Peijie Li, Yingying Yu, Yichi Zhang et al.ICML 2025
- Phase-Type Variational Autoencoders for Heavy-Tailed DataAbdelhakim Ziani, Andras Horvath, Paolo BallariniICML 2026
- Distributional Learning of Variational AutoEncoder: Application to Synthetic Data GenerationSeunghwan An, Jong-June JeonNeurIPS 2023 · 19 citations
- Sparse Autoencoders, Again?Yin Lu, Xuening Zhu, Tong He, David WipfICML 2025
