Relative gradient optimization of the Jacobian term in unsupervised deep learning
Luigi Gresele, Giancarlo Fissore, Adrián Javaloy, Bernhard Schölkopf, Aapo Hyvärinen
摘要
Learning expressive probabilistic models correctly describing the data is a ubiquitous problem in machine learning. A popular approach for solving it is mapping the observations into a representation space with a simple joint distribution, which can typically be written as a product of its marginals -- thus drawing a connection with the field of nonlinear independent component analysis. Deep density models have been widely used for this task, but their likelihood-based training requires estimating the log-determinant of the Jacobian and is computationally expensive, thus imposing a trade-off between computation and expressive power. In this work, we propose a new approach for exact likelihood-based training of such neural networks. Based on relative gradients, we exploit the matrix structure of neural network parameters to compute updates efficiently even in high-dimensional spaces; the computational cost of the training is quadratic in the input size, in contrast with the cubic scaling of the naive approaches. This allows fast training with objective functions involving the log-determinant of the Jacobian without imposing constraints on its structure, in stark contrast to normalizing flows. An implementation of our method can be found at this https URL
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Towards Nonlinear Disentanglement in Natural Data with Temporal Sparse CodingDavid A. Klindt, Lukas Schott, Yash Sharma, Ivan Ustyuzhaninov 等ICLR 2021 · 被引用 156 次
- Independent mechanism analysis, a new concept?Luigi Gresele, Julius von Kügelgen, Vincent Stimper, Bernhard Schölkopf 等NeurIPS 2021 · 被引用 133 次
- Nonparametric Identifiability of Causal Representations from Unknown InterventionsJulius von Kügelgen, Michel Besserve, Wendong Liang, Luigi Gresele 等NeurIPS 2023 · 被引用 127 次
- Disentangling Identifiable Features from Noisy Data with Structured Nonlinear ICAHermanni Hälvä, Sylvain Le Corff, Luc Lehéricy, Jonathan So 等NeurIPS 2021 · 被引用 87 次
- Causal Component AnalysisWendong Liang, Armin Kekic, Julius von Kügelgen, Simon Buchholz 等NeurIPS 2023 · 被引用 65 次
相关 Paper
- Self Normalizing FlowsT. Anderson Keller, Jorn W. T. Peters, Priyank Jaini, Emiel Hoogeboom 等ICML 2021 · 被引用 14 次
- OT-Flow: Fast and Accurate Continuous Normalizing Flows via Optimal TransportDerek Onken, Samy Wu Fung, Xingjian Li, Lars RuthottoAAAI 2021 · 被引用 210 次
- Convex Potential Flows: Universal Probability Distributions with Optimal Transport and Convex OptimizationChin-Wei Huang, Ricky T. Q. Chen, Christos Tsirigotis, Aaron C. CourvilleICLR 2021 · 被引用 107 次
- NanoFlow: Scalable Normalizing Flows with Sublinear Parameter ComplexitySang-gil Lee, Sungwon Kim, Sungroh YoonNeurIPS 2020 · 被引用 20 次
- Neural Inverse Transform SamplerHenry Li, Yuval KlugerICML 2022 · 被引用 4 次
