Universal Approximation Under Constraints is Possible with Transformers
Anastasis Kratsios, Behnoosh Zamanlooy, Tianlin Liu, Ivan Dokmanic
摘要
Many practical problems need the output of a machine learning model to satisfy a set of constraints, . Nevertheless, there is no known guarantee that classical neural network architectures can exactly encode constraints while simultaneously achieving universality. We provide a quantitative constrained universal approximation theorem which guarantees that for any non-convex compact set and any continuous function , there is a probabilistic transformer whose randomized outputs all lie in and whose expected output uniformly approximates . Our second main result is a"deep neural version"of Berge's Maximum Theorem (1963). The result guarantees that given an objective function , a constraint set , and a family of soft constraint sets, there is a probabilistic transformer that approximately minimizes and whose outputs belong to ; moreover, approximately satisfies the soft constraints. Our results imply the first universal approximation theorem for classical transformers with exact convex constraint satisfaction. They also yield that a chart-free universal approximation theorem for Riemannian manifold-valued functions subject to suitable geodesically convex constraints.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Pretrained Language Models as Visual Planners for Human AssistanceDhruvesh Patel, Hamid Eghbalzadeh, Nitin Kamra, Michael Louis Iuzzolino 等ICCV 2023 · 被引用 41 次
- Approximation Rate of the Transformer Architecture for Sequence ModelingHaotian Jiang, Qianxiao LiNeurIPS 2024 · 被引用 32 次
- Are Transformers with One Layer Self-Attention Using Low-Rank Weight Matrices Universal Approximators?Tokio Kajitsuka, Issei SatoICLR 2024 · 被引用 31 次
- Pinet: Optimizing hard-constrained neural networks with orthogonal projection layersPanagiotis D. Grontas, Antonio Terpin, Efe C. Balta, Raffaello D'Andrea 等ICLR 2026 · 被引用 22 次
- Low Complexity Homeomorphic Projection to Ensure Neural-Network Solution Feasibility for Optimization over (Non-)Convex SetEnming Liang, Minghua Chen, Steven H. LowICML 2023 · 被引用 18 次
它引用的顶会 Paper11
- Are Transformers universal approximators of sequence-to-sequence functions?Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank J. Reddi 等ICLR 2020 · 被引用 481 次
- What is Local Optimality in Nonconvex-Nonconcave Minimax Optimization?Chi Jin, Praneeth Netrapalli, Michael I. JordanICML 2020 · 被引用 381 次
- Model-Based Domain GeneralizationAlexander Robey, George J. Pappas, Hamed HassaniNeurIPS 2021 · 被引用 167 次
- The phase diagram of approximation rates for deep neural networksDmitry Yarotsky, Anton ZhevnerchukNeurIPS 2020 · 被引用 156 次
- Minimum Width for Universal ApproximationSejun Park, Chulhee Yun, Jaeho Lee, Jinwoo ShinICLR 2021 · 被引用 148 次
相关 Paper
- Non-Euclidean Universal ApproximationAnastasis Kratsios, Ievgen BilokopytovNeurIPS 2020 · 被引用 64 次
- Deep Ridgelet Transform and Unified Universality Theorem for Deep and Shallow Joint-Group-Equivariant MachinesSho Sonoda, Yuka Hashimoto, Isao Ishikawa, Masahiro IkedaICML 2025
- A closer look at the approximation capabilities of neural networksKai Fong Ernest ChongICLR 2020 · 被引用 18 次
- On Solution Functions of Optimization: Universal Approximation and Covering Number BoundsMing Jin, Vanshaj Khattar, Harshal Kaushik, Bilgehan Sel 等AAAI 2023 · 被引用 15 次
- CAffNet: Hard Constraint-Affine Neural NetworksYang Zhao, Jungeun Lee, Jeong hwan Jeon, Sze Zheng YongICML 2026 · 被引用 1 次
