Non-Euclidean Universal Approximation
Anastasis Kratsios, Ievgen Bilokopytov
摘要
Modifications to a neural network's input and output layers are often required to accommodate the specificities of most practical learning tasks. However, the impact of such changes on architecture's approximation capabilities is largely not understood. We present general conditions describing feature and readout maps that preserve an architecture's ability to approximate any continuous functions uniformly on compacts. As an application, we show that if an architecture is capable of universal approximation, then modifying its final layer to produce binary values creates a new architecture capable of deterministically approximating any classifier. In particular, we obtain guarantees for deep CNNs and deep feed-forward networks. Our results also have consequences within the scope of geometric deep learning. Specifically, when the input and output spaces are Cartan-Hadamard manifolds, we obtain geometrically meaningful feature and readout maps satisfying our criteria. Consequently, commonly used non-Euclidean regression models between spaces of symmetric positive definite matrices are extended to universal DNNs. The same result allows us to show that the hyperbolic feed-forward networks, used for hierarchical learning, are universal. Our result is also used to show that the common practice of randomizing all but the last two layers of a DNN produces a universal family of functions with probability one. We also provide conditions on a DNN's first (resp. last) few layer's connections and activation function which guarantee that these layer's can have a width equal to the input (resp. output) space's dimension while not negatively effecting the architecture's approximation capabilities.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Universal Approximation Under Constraints is Possible with TransformersAnastasis Kratsios, Behnoosh Zamanlooy, Tianlin Liu, Ivan DokmanicICLR 2022 · 被引用 38 次
- Globally injective and bijective neural operatorsTakashi Furuya, Michael Puthawala, Matti Lassas, Maarten V. de HoopNeurIPS 2023 · 被引用 18 次
- Shape-Informed Clustering of Multi-Dimensional Functional Data via Deep Functional AutoencodersSamuel V. Singh, Shirley Coyle, Mimi ZhangNeurIPS 2025 · 被引用 7 次
- Can neural operators always be continuously discretized?Takashi Furuya, Michael Puthawala, Matti Lassas, Maarten V. de HoopNeurIPS 2024 · 被引用 5 次
- Riemannian Neural Optimal TransportAlessandro Micheli, Yueqi Cao, Anthea Monod, Samir BhattICML 2026 · 被引用 1 次
它引用的顶会 Paper2
相关 Paper
- Universal approximation power of deep residual neural networks via nonlinear control theoryPaulo Tabuada, Bahman GharesifardICLR 2021 · 被引用 31 次
- Batch normalization is sufficient for universal function approximation in CNNsRebekka BurkholzICLR 2024 · 被引用 8 次
- On Universality of Deep Equivariant NetworksMarco Pacini, Mircea Petrache, Bruno Lepri, Shubhendu Trivedi 等ICLR 2026 · 被引用 4 次
- Effects of Data Geometry in Early Deep LearningSaket Tiwari, George KonidarisNeurIPS 2022 · 被引用 14 次
- Asymptotics of representation learning in finite Bayesian neural networksJacob A. Zavatone-Veth, Abdulkadir Canatar, Benjamin S. Ruben, Cengiz PehlevanNeurIPS 2021 · 被引用 45 次
