Representational aspects of depth and conditioning in normalizing flows
Frederic Koehler, Viraj Mehta, Andrej Risteski
Abstract
Normalizing flows are among the most popular paradigms in generative modeling, especially for images, primarily because we can efficiently evaluate the likelihood of a data point. Normalizing flows also come with difficulties: models which produce good samples typically need to be extremely deep -- which comes with accompanying vanishing/exploding gradient problems. Relatedly, they are often poorly conditioned since typical training data like images intuitively are lower-dimensional, and the learned maps often have Jacobians that are close to being singular. In our paper, we tackle representational aspects around depth and conditioning of normalizing flows -- both for general invertible architectures, and for a particular common architecture -- affine couplings. For general invertible architectures, we prove that invertibility comes at a cost in terms of depth: we show examples where a much deeper normalizing flow model may need to be used to match the performance of a non-invertible generator. For affine couplings, we first show that the choice of partitions isn't a likely bottleneck for depth: we show that any invertible linear map (and hence a permutation) can be simulated by a constant number of affine coupling layers, using a fixed partition. This shows that the extra flexibility conferred by 1x1 convolution layers, as in GLOW, can in principle be simulated by increasing the size by a constant factor. Next, in terms of conditioning, we show that affine couplings are universal approximators -- provided the Jacobian of the model is allowed to be close to singular. We furthermore empirically explore the benefit of different kinds of padding -- a common strategy for improving conditioning -- on both synthetic and real-life datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers11
- On the Universality of Volume-Preserving and Coupling-Based Normalizing FlowsFelix Draxler, Stefan Wahl, Christoph Schnörr, Ullrich KötheICML 2024 · 19 citations
- Low Complexity Homeomorphic Projection to Ensure Neural-Network Solution Feasibility for Optimization over (Non-)Convex SetEnming Liang, Minghua Chen, Steven H. LowICML 2023 · 18 citations
- Universal Approximation Using Well-Conditioned Normalizing FlowsHolden Lee, Chirag Pabbaraju, Anish Prasad Sevekari, Andrej RisteskiNeurIPS 2021 · 15 citations
- MGF: Mixed Gaussian Flow for Diverse Trajectory PredictionJiahe Chen, Jinkun Cao, Dahua Lin, Kris Kitani et al.NeurIPS 2024 · 11 citations
- Whitening Convergence Rate of Coupling-based Normalizing FlowsFelix Draxler, Christoph Schnörr, Ullrich KötheNeurIPS 2022 · 7 citations
Builds on2
- Coupling-based Invertible Neural Networks Are Universal Diffeomorphism ApproximatorsTakeshi Teshima, Isao Ishikawa, Koichi Tojo, Kenta Oono et al.NeurIPS 2020 · 129 citations
- Approximation Capabilities of Neural ODEs and Invertible Residual NetworksHan Zhang, Xi Gao, Jacob Unterman, Tom ArodzICML 2020 · 114 citations
Related papers
- Densely connected normalizing flowsMatej Grcic, Ivan Grubisic, Sinisa SegvicNeurIPS 2021 · 67 citations
- Self Normalizing FlowsT. Anderson Keller, Jorn W. T. Peters, Priyank Jaini, Emiel Hoogeboom et al.ICML 2021 · 14 citations
- On the Robustness of Normalizing Flows for Inverse Problems in ImagingSeongmin Hong, Inbum Park, Se Young ChunICCV 2023 · 9 citations
- Generative Flows with Matrix ExponentialChangyi Xiao, Ligang LiuICML 2020 · 10 citations
- Flowification: Everything is a normalizing flowBálint Máté, Samuel Klein, Tobias Golling, François FleuretNeurIPS 2022 · 9 citations
