Breaking the Simplification Bottleneck in Amortized Neural Symbolic Regression
Paul Saegert, Ullrich Koethe
Abstract
Symbolic regression (SR) aims to discover interpretable analytical expressions that accurately describe observed data. Amortized SR promises to be much more efficient than the predominant genetic programming SR methods, but currently struggles to scale to realistic scientific complexity. We find that a key obstacle is the lack of a fast reduction of equivalent expressions to a concise normalized form. Amortized SR has addressed this with general-purpose Computer Algebra Systems (CAS) like SymPy, but the high computational cost severely limits training and inference speed. We propose SimpliPy , a rule-based simplification engine achieving a 100-fold speed-up over SymPy at comparable quality. This enables substantial improvements in amortized SR, including scalability to much larger training sets, more efficient use of the per-expression token budget, and systematic training set decontamination with respect to equivalent test expressions. We demonstrate these advantages in our Flash-ANSR framework, which achieves much better accuracy than amortized baselines (NeSymReS, E2E) on the FastSRB benchmark. Moreover, it performs on par with state-of-the-art direct optimization (PySR) while recovering more concise rather than more complex expressions with increasing inference budget.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4d618e70-c913-441e-a9c5-ed7d8daeb695Builds on14
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng et al.ICML 2020 · 1,388 citations
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang et al.ICLR 2020 · 1,108 citations
- Deduplicating Training Data Makes Language Models BetterKatherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang et al.ACL 2022 · 844 citations
- PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph CompilationJason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein et al.ASPLOS 2024 · 693 citations
Related papers
- Deep Generative Symbolic RegressionSamuel Holt, Zhaozhi Qian, Mihaela van der SchaarICLR 2023 · 4 citations
- End-to-end Symbolic Regression with TransformersPierre-Alexandre Kamienny, Stéphane d'Ascoli, Guillaume Lample, François ChartonNeurIPS 2022 · 320 citations
- A Unified Framework for Deep Symbolic RegressionMikel Landajuela, Chak Shing Lee, Jiachen Yang, Ruben Glatt et al.NeurIPS 2022 · 160 citations
- Neural–Evolutionary Symbolic Regression with Global Constraints: Constraint-Aware Decoding and Reward ShapingXiangdong Wu, wenjun wu, Ziyu Wei, Bingrun Chen et al.ICML 2026
- ParFam - (Neural Guided) Symbolic Regression via Continuous Global OptimizationPhilipp Scholl, Katharina Bieker, Hillary Hauger, Gitta KutyniokICLR 2025 · 1 citation
