Bridging the Empirical-Theoretical Gap in Neural Network Formal Language Learning Using Minimum Description Length
Nur Geffen Lan, Emmanuel Chemla, Roni Katzir
Abstract
Neural networks offer good approximation to many tasks but consistently fail to reach perfect generalization, even when theoretical work shows that such perfect solutions can be expressed by certain architectures. Using the task of formal language learning, we focus on one simple formal language and show that the theoretically correct solution is in fact not an optimum of commonly used objectiveseven with regularization techniques that according to common wisdom should lead to simple weights and good generalization (L1, L2) or other meta-heuristics (early-stopping, dropout). On the other hand, replacing standard targets with the Minimum Description Length objective (MDL) results in the correct solution being an optimum.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cbcaf4e4-8918-4414-a059-eb9cb4bd4edeCited by top-tier papers1
Ask how each one uses itBuilds on2
Related papers
- Bridging Kolmogorov Complexity and Deep Learning: Asymptotically Optimal Description Length Objectives for TransformersPeter Shaw, James Cohan, Jacob Eisenstein, Kristina ToutanovaICLR 2026 · 7 citations
- Training Neural Networks as Recognizers of Formal LanguagesAlexandra Butoi, Ghazal Khalighinejad, Anej Svete, Josef Valvoda et al.ICLR 2025
- Symbolic regression via MDLformer-guided search: from minimizing prediction error to minimizing description lengthZihan Yu, Jingtao Ding, Yong Li, Depeng JinICLR 2025
- Neural Networks and the Chomsky HierarchyGrégoire Delétang, Anian Ruoss, Jordi Grau-Moya, Tim Genewein et al.ICLR 2023 · 45 citations
- Generalization error in high-dimensional perceptrons: Approaching Bayes error with convex optimizationBenjamin Aubin, Florent Krzakala, Yue M. Lu, Lenka ZdeborováNeurIPS 2020 · 67 citations
