Lune

ICML2026Top-tier venue

Floating-Point Networks with Automatic Differentiation Can Represent Almost All Floating-Point Functions and Their Gradients

Sejun Park, Yeachan Park, Geonho Hwang

2026Year

Abstract

Theoretical studies show that for any differentiable function on a compact domain, there exists a neural network that approximates both the function values and gradients. However, such a result cannot be used in practice since it assumes real parameters and exact internal operations. In contrast, real implementations only use a finite subset of reals and machine operations with round-off errors. In this work, we investigate whether a similar result holds for neural networks under floating-point arithmetic, when the gradient with respect to the input is computed by the automatic differentiation algorithm DADD^\mathtt{AD}. We first show that given a floating-point function ϕ\phi (e.g., a loss function), arbitrary function values and gradients can be represented by a floating-point network ff and DAD(ϕ∘f)D^\mathtt{AD}(\phi\circ f), respectively. We further extend this result: given ϕ1,…,ϕn\phi_1,\dots,\phi_n, DAD(ϕi∘f)D^\mathtt{AD}(\phi_i\circ f) can simultaneously represent arbitrary gradients while ff represents the target values, under mild conditions. Our results hold for practical activation functions, e.g., ReLU, ELU, GELU, Swish, Sigmoid, and tanh.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext cea6407a-715e-4665-8ba2-79d6f9fabd4b

Builds on4

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines