On the complexity of nonsmooth automatic differentiation
Jérôme Bolte, Ryan Boustany, Edouard Pauwels, Béatrice Pesquet-Popescu
摘要
Using the notion of conservative gradient, we provide a simple model to estimate the computational costs of the backward and forward modes of algorithmic differentiation for a wide class of nonsmooth programs. The overhead complexity of the backward mode turns out to be independent of the dimension when using programs with locally Lipschitz semi-algebraic or definable elementary functions. This considerably extends Baur-Strassen's smooth cheap gradient principle. We illustrate our results by establishing fast backpropagation results of conservative gradients through feedforward neural networks with standard activation and loss functions. Nonsmooth backpropagation's cheapness contrasts with concurrent forward approaches, which have, to this day, dimensional-dependent worst-case overhead estimates. We provide further results suggesting the superiority of backward propagation of conservative gradients. Indeed, we relate the complexity of computing a large number of directional derivatives to that of matrix multiplication, and we show that finding two subgradients in the Clarke subdifferential of a function is an NP-hard problem.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- On the Correctness of Automatic Differentiation for Neural Networks with Machine-Representable ParametersWonyeol Lee, Sejun Park, Alex AikenICML 2023 · 被引用 6 次
- Testing Approximate Stationarity Concepts for Piecewise Affine FunctionsLai Tian, Anthony Man-Cho SoSODA 2025 · 被引用 1 次
它引用的顶会 Paper7
- Efficient and Modular Implicit DifferentiationMathieu Blondel, Quentin Berthet, Marco Cuturi, Roy Frostig 等NeurIPS 2022 · 被引用 386 次
- A Refined Laser Method and Faster Matrix MultiplicationJosh Alman, Virginia Vassilevska WilliamsSODA 2021 · 被引用 275 次
- Monotone operator equilibrium networksEzra Winston, J. Zico KolterNeurIPS 2020 · 被引用 177 次
- Nonsmooth Implicit Differentiation for Machine-Learning and OptimizationJérôme Bolte, Tam Le, Edouard Pauwels, Antonio Silveti-FallsNeurIPS 2021 · 被引用 85 次
- A mathematical model for automatic differentiation in machine learningJérôme Bolte, Edouard PauwelsNeurIPS 2020 · 被引用 84 次
相关 Paper
- What does automatic differentiation compute for neural networks?Sejun Park, Sanghyuk Chun, Wonyeol LeeICLR 2024
- Automatic differentiation of nonsmooth iterative algorithmsJérôme Bolte, Edouard Pauwels, Samuel VaiterNeurIPS 2022 · 被引用 33 次
- One-step differentiation of iterative algorithmsJérôme Bolte, Edouard Pauwels, Samuel VaiterNeurIPS 2023 · 被引用 36 次
- δ is for DialecticaMarie Morgane Kerjean, Pierre-Marie PédrotLICS 2024 · 被引用 1 次
- Exactly Computing the Local Lipschitz Constant of ReLU NetworksMatt Jordan, Alexandros G. DimakisNeurIPS 2020 · 被引用 156 次
