Statistical Schema Learning with Occam's Razor
Justin Talbot, Daniel Ting
Abstract
A judiciously normalized database schema can increase data interpretability, reduce data size, and improve data integrity. However, real world data sets are often stored or shared in a denormalized state. We examine the problem of automatically creating a good schema for a denormalized table, approaching it as an unsupervised machine learning problem which must learn an optimal schema from the data. This differs from past rule-based approaches that focus on normalization into a canonical form. We define a principled schema optimization criterion, based on Occam's razor, that is robust to noise and extensible---allowing users to easily specify desirable properties of the resulting schema. We develop an efficient learning algorithm for this criterion and empirically demonstrate that it is 3 to 100 times faster than previous work and produces higher quality schemas with 1/5th the errors.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- In-Context Learning and Occam's RazorEric Elmoznino, Tom Marty, Tejas Kasetty, Léo Gagnon et al.ICML 2025
- Bridging Kolmogorov Complexity and Deep Learning: Asymptotically Optimal Description Length Objectives for TransformersPeter Shaw, James Cohan, Jacob Eisenstein, Kristina ToutanovaICLR 2026 · 7 citations
- TabNet: Attentive Interpretable Tabular LearningSercan Ö. Arik, Tomas PfisterAAAI 2021 · 2,148 citations
- PAC-Bayes Compression Bounds So Tight That They Can Explain GeneralizationSanae Lotfi, Marc Finzi, Sanyam Kapoor, Andres Potapczynski et al.NeurIPS 2022 · 98 citations
- Composite Object Normal Forms: Parameterizing Boyce-Codd Normal Form by the Number of Minimal KeysZhuoxing Zhang, Wu Chen, Sebastian LinkSIGMOD 2023 · 9 citations
