Multi-Rate VAE: Train Once, Get the Full Rate-Distortion Curve
Juhan Bae, Michael R. Zhang, Michael Ruan, Eric Wang, So Hasegawa, Jimmy Ba, Roger Baker Grosse
Abstract
Variational autoencoders (VAEs) are powerful tools for learning latent representations of data used in a wide range of applications. In practice, VAEs usually require multiple training rounds to choose the amount of information the latent variable should retain. This trade-off between the reconstruction error (distortion) and the KL divergence (rate) is typically parameterized by a hyperparameter . In this paper, we introduce Multi-Rate VAE (MR-VAE), a computationally efficient framework for learning optimal parameters corresponding to various in a single training run. The key idea is to explicitly formulate a response function that maps to the optimal parameters using hypernetworks. MR-VAEs construct a compact response hypernetwork where the pre-activations are conditionally gated based on . We justify the proposed architecture by analyzing linear VAEs and showing that it can represent response functions exactly for linear VAEs. With the learned hypernetwork, MR-VAEs can construct the rate-distortion curve without additional training and can be deployed with significantly less hyperparameter tuning. Empirically, our approach is competitive and often exceeds the performance of multiple -VAEs training with minimal computation and memory overheads.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Probabilistic Inference in Language Models via Twisted Sequential Monte CarloStephen Zhao, Rob Brekelmans, Alireza Makhzani, Roger Baker GrosseICML 2024 · 61 citations
- Tree Variational AutoencodersLaura Manduchi, Moritz Vandenhirtz, Alain Ryser, Julia E. VogtNeurIPS 2023 · 17 citations
- Magnitude Invariant Parametrizations Improve Hypernetwork LearningJose Javier Gonzalez Ortiz, John V. Guttag, Adrian V. DalcaICLR 2024 · 13 citations
- Generalization in VAE and Diffusion Models: A Unified Information-Theoretic AnalysisQi Chen, Jierui Zhu, Florian ShkurtiICLR 2025
Builds on14
- NVAE: A Deep Hierarchical Variational AutoencoderArash Vahdat, Jan KautzNeurIPS 2020 · 1,141 citations
- Skew-Fit: State-Covering Self-Supervised Reinforcement LearningVitchyr Pong, Murtaza Dalal, Steven Lin, Ashvin Nair et al.ICML 2020 · 303 citations
- Learning the Pareto Front with HypernetworksAviv Navon, Aviv Shamsian, Ethan Fetaya, Gal ChechikICLR 2021 · 189 citations
- Improved Conditional VRNNs for Video PredictionLluís Castrejón, Nicolas Ballas, Aaron C. CourvilleICCV 2019 · 177 citations
- Multiplicative Interactions and Where to Find ThemSiddhant M. Jayakumar, Wojciech M. Czarnecki, Jacob Menick, Jonathan Schwarz et al.ICLR 2020 · 152 citations
Related papers
- Trading Information between Latents in Hierarchical Variational AutoencodersTim Z. Xiao, Robert BamlerICLR 2023
- Simple and Effective VAE Training with Calibrated DecodersOleh Rybkin, Kostas Daniilidis, Sergey LevineICML 2021 · 119 citations
- Learning to Quantize for Training Vector-Quantized NetworksPeijia Qin, Jianguo ZhangICML 2025
- You Only Train Once: Loss-Conditional Training of Deep NetworksAlexey Dosovitskiy, Josip DjolongaICLR 2020 · 96 citations
- Variable Rate Deep Image Compression With a Conditional AutoencoderYoojin Choi, Mostafa El-Khamy, Jungwon LeeICCV 2019 · 265 citations
