Ab Initio Nonparametric Variable Selection for Scalable Symbolic Regression with Large p
Shengbin Ye, Meng Li
摘要
Symbolic regression (SR) is a powerful technique for discovering symbolic expressions that characterize nonlinear relationships in data, gaining increasing attention for its interpretability, compactness, and robustness. However, existing SR methods do not scale to datasets with a large number of input variables (referred to as extreme-scale SR), which is common in modern scientific applications. This "large p" setting, often accompanied by measurement error, leads to slow performance of SR methods and overly complex expressions that are difficult to interpret. To address this scalability challenge, we propose a method called PAN+SR, which combines a key idea of ab initio nonparametric variable selection with SR to efficiently pre-screen large input spaces and reduce search complexity while maintaining accuracy. The use of nonparametric methods eliminates model misspecification, supporting a strategy called parametric-assisted nonparametric (PAN). We also extend SRBench, an open-source benchmarking platform, by incorporating highdimensional regression problems with various signal-to-noise ratios. Our results demonstrate that PAN+SR consistently enhances the performance of 19 contemporary SR methods, enabling several to achieve state-of-the-art performance on these challenging datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper6
- End-to-end Symbolic Regression with TransformersPierre-Alexandre Kamienny, Stéphane d'Ascoli, Guillaume Lample, François ChartonNeurIPS 2022 · 被引用 320 次
- AI Feynman 2.0: Pareto-optimal symbolic regression exploiting graph modularitySilviu-Marian Udrescu, Andrew K. Tan, Jiahai Feng, Orisvaldo Neto 等NeurIPS 2020 · 被引用 267 次
- A Unified Framework for Deep Symbolic RegressionMikel Landajuela, Chak Shing Lee, Jiachen Yang, Ruben Glatt 等NeurIPS 2022 · 被引用 160 次
- Transformer-based Planning for Symbolic RegressionParshin Shojaee, Kazem Meidani, Amir Barati Farimani, Chandan K. ReddyNeurIPS 2023 · 被引用 116 次
- Deep Generative Symbolic Regression with Monte-Carlo-Tree-SearchPierre-Alexandre Kamienny, Guillaume Lample, Sylvain Lamprier, Marco VirgolinICML 2023 · 被引用 50 次
相关 Paper
- Deep Generative Symbolic RegressionSamuel Holt, Zhaozhi Qian, Mihaela van der SchaarICLR 2023 · 被引用 4 次
- A Neural-Guided Dynamic Symbolic Network for Exploring Mathematical Expressions from DataWenqiang Li, Weijun Li, Lina Yu, Min Wu 等ICML 2024 · 被引用 16 次
- Pareto-Optimal Fronts for Benchmarking Symbolic Regression AlgorithmsKei Sen Fong, Mehul MotaniICML 2025
- Syntax-Aware Retrieval Augmentation for Neural Symbolic RegressionCanmiao Zhou, Han HuangEMNLP 2025
- Breaking the Simplification Bottleneck in Amortized Neural Symbolic RegressionPaul Saegert, Ullrich KoetheICML 2026
