What does automatic differentiation compute for neural networks?
Sejun Park, Sanghyuk Chun, Wonyeol Lee
摘要
Forward-or reverse-mode automatic differentiation (AD) is a popular algorithm for computing the derivative of a function expressed by a program. AD always outputs the correct derivative if a program does not use any non-differentiable functions and control flows; however, it may return an arbitrary value otherwise. In this work, we investigate what AD computes for neural networks that may contain non-differentiable functions such as ReLU and maxpools. We first prove that AD always returns a generalized derivative called a Clarke subderivative for networks with pointwise activation functions, if the minibatch size is one and all non-differentiable neurons have distinct bias parameters. We show that the same conclusion does not hold otherwise, but does hold under some mild sufficient conditions. We also prove similar results for more general networks that can use maxpools and bias parameters shared across different neurons. We empirically check our sufficient conditions over popular network architectures and observe that AD almost always computes a Clarke subderivative in practical learning setups. * Equal contribution Emails: sejun.park000, sanghyuk.chun, wonyeol.lee.cs@gmail.com * AD uses either D -ρ(x) or D + ρ(x) as the proxy derivative of a pointwise activation function ρ at x, where D -ρ and D + ρ denote the left-hand and right-hand derivatives of ρ. † AD uses an element of the Clarke subdifferential as the proxy derivative of ρ. ‡ AD uses λD -ρ(x) + (1 -λ)D + ρ(x) as the proxy derivative of ρ at x; λ ∈ [0, 1] is shared in a layer. ¶ A sufficient condition for AD to be correct at the current parameter/input values is given in Theorem 5. § A sufficient condition for AD to be correct at the current parameter/input values is given in Theorem 8.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le 等ICCV 2019 · 被引用 9,163 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- Gradient flow dynamics of shallow ReLU networks for square loss and orthogonal inputsEtienne Boursier, Loucas Pillaud-Vivien, Nicolas FlammarionNeurIPS 2022 · 被引用 92 次
- A mathematical model for automatic differentiation in machine learningJérôme Bolte, Edouard PauwelsNeurIPS 2020 · 被引用 84 次
- On Correctness of Automatic Differentiation for Non-Differentiable FunctionsWonyeol Lee, Hangyeol Yu, Xavier Rival, Hongseok YangNeurIPS 2020 · 被引用 50 次
相关 Paper
- On the Correctness of Automatic Differentiation for Neural Networks with Machine-Representable ParametersWonyeol Lee, Sejun Park, Alex AikenICML 2023 · 被引用 6 次
- Floating-Point Networks with Automatic Differentiation Can Represent Almost All Floating-Point Functions and Their GradientsSejun Park, Yeachan Park, Geonho HwangICML 2026
- Automatic differentiation in PCFDamiano Mazza, Michele PaganiPOPL 2021 · 被引用 47 次
- Complexity of Finding Stationary Points of Nonconvex Nonsmooth FunctionsJingzhao Zhang, Hongzhou Lin, Stefanie Jegelka, Suvrit Sra 等ICML 2020 · 被引用 98 次
- On the complexity of nonsmooth automatic differentiationJérôme Bolte, Ryan Boustany, Edouard Pauwels, Béatrice Pesquet-PopescuICLR 2023
