Non-Euclidean Universal Approximation
Anastasis Kratsios, Ievgen Bilokopytov
Abstract
Modifications to a neural network's input and output layers are often required to accommodate the specificities of most practical learning tasks. However, the impact of such changes on architecture's approximation capabilities is largely not understood. We present general conditions describing feature and readout maps that preserve an architecture's ability to approximate any continuous functions uniformly on compacts. As an application, we show that if an architecture is capable of universal approximation, then modifying its final layer to produce binary values creates a new architecture capable of deterministically approximating any classifier. In particular, we obtain guarantees for deep CNNs and deep feed-forward networks. Our results also have consequences within the scope of geometric deep learning. Specifically, when the input and output spaces are Cartan-Hadamard manifolds, we obtain geometrically meaningful feature and readout maps satisfying our criteria. Consequently, commonly used non-Euclidean regression models between spaces of symmetric positive definite matrices are extended to universal DNNs. The same result allows us to show that the hyperbolic feed-forward networks, used for hierarchical learning, are universal. Our result is also used to show that the common practice of randomizing all but the last two layers of a DNN produces a universal family of functions with probability one. We also provide conditions on a DNN's first (resp. last) few layer's connections and activation function which guarantee that these layer's can have a width equal to the input (resp. output) space's dimension while not negatively effecting the architecture's approximation capabilities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Universal Approximation Under Constraints is Possible with TransformersAnastasis Kratsios, Behnoosh Zamanlooy, Tianlin Liu, Ivan DokmanicICLR 2022 · 38 citations
- Globally injective and bijective neural operatorsTakashi Furuya, Michael Puthawala, Matti Lassas, Maarten V. de HoopNeurIPS 2023 · 18 citations
- Shape-Informed Clustering of Multi-Dimensional Functional Data via Deep Functional AutoencodersSamuel V. Singh, Shirley Coyle, Mimi ZhangNeurIPS 2025 · 7 citations
- Can neural operators always be continuously discretized?Takashi Furuya, Michael Puthawala, Matti Lassas, Maarten V. de HoopNeurIPS 2024 · 5 citations
- Riemannian Neural Optimal TransportAlessandro Micheli, Yueqi Cao, Anthea Monod, Samir BhattICML 2026 · 1 citation
Builds on2
Related papers
- Universal approximation power of deep residual neural networks via nonlinear control theoryPaulo Tabuada, Bahman GharesifardICLR 2021 · 31 citations
- Batch normalization is sufficient for universal function approximation in CNNsRebekka BurkholzICLR 2024 · 8 citations
- On Universality of Deep Equivariant NetworksMarco Pacini, Mircea Petrache, Bruno Lepri, Shubhendu Trivedi et al.ICLR 2026 · 4 citations
- Effects of Data Geometry in Early Deep LearningSaket Tiwari, George KonidarisNeurIPS 2022 · 14 citations
- Asymptotics of representation learning in finite Bayesian neural networksJacob A. Zavatone-Veth, Abdulkadir Canatar, Benjamin S. Ruben, Cengiz PehlevanNeurIPS 2021 · 45 citations
