Symbolic regression via MDLformer-guided search: from minimizing prediction error to minimizing description length
Zihan Yu, Jingtao Ding, Yong Li, Depeng Jin
摘要
Symbolic regression, a task discovering the formula best fitting the given data, is typically based on the heuristical search. These methods usually update candidate formulas to obtain new ones with lower prediction errors iteratively. However, since formulas with similar function shapes may have completely different symbolic forms, the prediction error does not decrease monotonously as the search approaches the target formula, causing the low recovery rate of existing methods. To solve this problem, we propose a novel search objective based on the minimum description length, which reflects the distance from the target and decreases monotonically as the search approaches the correct form of the target formula. To estimate the minimum description length of any input data, we design a neural network, MDLformer, which enables robust and scalable estimation through large-scale training. With the MDLformer's output as the search objective, we implement a symbolic regression method, SR4MDL, that can effectively recover the correct mathematical form of the formula. Extensive experiments illustrate its excellent performance in recovering formulas from data. Our method successfully recovers around 50 formulas across two benchmark datasets comprising 133 problems, outperforming state-of-the-art methods by 43.92%. Experiments on 122 unseen black-box problems further demonstrate its generalization performance. We release our code at https://github.com/ tsinghua-fib-lab/SR4MDL .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- GenSR: Symbolic regression based on equation generative spaceQian Li, Yuxiao Hu, Juncheng Liu, Yuntian ChenICLR 2026 · 被引用 7 次
- Beyond Accuracy and Complexity: The Effective Information Criterion for Structurally Stable Symbolic RegressionZihan Yu, Guanren Wang, Jingtao Ding, Huandong Wang 等ICML 2026
- Discovering Ordinary Differential Equations with LLM-Based Qualitative and Quantitative EvaluationSum Kyun Song, Bong Gyun Shin, JaeYong LeeICML 2026
- Breaking the Simplification Bottleneck in Amortized Neural Symbolic RegressionPaul Saegert, Ullrich KoetheICML 2026
- MOD-SR: Unifying Multimodal Learning and Direct Optimization with Gradient-Guided Diffusion Model for Symbolic RegressionChuyang Xiang, Yichen Wei, Junchi YanICML 2026
它引用的顶会 Paper9
- Deep symbolic regression: Recovering mathematical expressions from data via risk-seeking policy gradientsBrenden K. Petersen, Mikel Landajuela, T. Nathan Mundhenk, Cláudio Prata Santiago 等ICLR 2021 · 被引用 444 次
- End-to-end Symbolic Regression with TransformersPierre-Alexandre Kamienny, Stéphane d'Ascoli, Guillaume Lample, François ChartonNeurIPS 2022 · 被引用 320 次
- AI Feynman 2.0: Pareto-optimal symbolic regression exploiting graph modularitySilviu-Marian Udrescu, Andrew K. Tan, Jiahai Feng, Orisvaldo Neto 等NeurIPS 2020 · 被引用 267 次
- A Unified Framework for Deep Symbolic RegressionMikel Landajuela, Chak Shing Lee, Jiachen Yang, Ruben Glatt 等NeurIPS 2022 · 被引用 160 次
- A-NeSI: A Scalable Approximate Method for Probabilistic Neurosymbolic InferenceEmile van Krieken, Thiviyan Thanapalasingam, Jakub M. Tomczak, Frank van Harmelen 等NeurIPS 2023 · 被引用 62 次
相关 Paper
- A Neural-Guided Dynamic Symbolic Network for Exploring Mathematical Expressions from DataWenqiang Li, Weijun Li, Lina Yu, Min Wu 等ICML 2024 · 被引用 16 次
- Symbolic Regression via Deep Reinforcement Learning Enhanced Genetic Programming SeedingT. Nathan Mundhenk, Mikel Landajuela, Ruben Glatt, Cláudio P. Santiago 等NeurIPS 2021 · 被引用 95 次
- Deep Generative Symbolic RegressionSamuel Holt, Zhaozhi Qian, Mihaela van der SchaarICLR 2023 · 被引用 4 次
- MetaSymNet: A Tree-like Symbol Network with Adaptive Architecture and Activation FunctionsYanjie Li, Weijun Li, Lina Yu, Min Wu 等AAAI 2025 · 被引用 1 次
- EGG-SR: Embedding Symbolic Equivalence into Symbolic Regression via Equality GraphNan Jiang, Ziyi Wang, Yexiang XueICLR 2026 · 被引用 3 次
