Empirical Study of the Benefits of Overparameterization in Learning Latent Variable Models
Rares-Darius Buhai, Yoni Halpern, Yoon Kim, Andrej Risteski, David A. Sontag
Abstract
One of the most surprising and exciting discoveries in supervised learning was the benefit of overparameterization (i.e. training a very large model) to improving the optimization landscape of a problem, with minimal effect on statistical performance (i.e. generalization). In contrast, unsupervised settings have been under-explored, despite the fact that it was observed that overparameterization can be helpful as early as Dasgupta & Schulman (2007) . We perform an empirical study of different aspects of overparameterization in unsupervised learning of latent variable models via synthetic and semi-synthetic experiments. We discuss benefits to different metrics of success (recovering the parameters of the ground-truth model, held-out log-likelihood), sensitivity to variations of the training algorithm, and behavior as the amount of overparameterization increases. We find that across a variety of models (noisy-OR networks, sparse coding, probabilistic context-free grammars) and training algorithms (variational inference, alternating minimization, expectation-maximization), overparameterization can significantly increase the number of ground truth latent variables recovered. The code and supporting files for the experiments are located at https: //github.com/clinicalml/overparam .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2693da7e-13e1-4c3a-a161-ac2b043182e8Cited by top-tier papers7
- Sequence-to-Sequence Learning with Latent Neural GrammarsYoon KimNeurIPS 2021 · 44 citations
- Can Looped Transformers Learn to Implement Multi-step Gradient Descent for In-context Learning?Khashayar Gatmiry, Nikunj Saunshi, Sashank J. Reddi, Stefanie Jegelka et al.ICML 2024 · 43 citations
- Compressible Dynamics in Deep Overparameterized Low-Rank Learning & AdaptationCan Yaras, Peng Wang, Laura Balzano, Qing QuICML 2024 · 29 citations
- The Lazy Neuron Phenomenon: On Emergence of Activation Sparsity in TransformersZonglin Li, Chong You, Srinadh Bhojanapalli, Daliang Li et al.ICLR 2023 · 10 citations
- Learning Noisy OR Bayesian Networks with Max-Product Belief PropagationAntoine Dedieu, Guangyao Zhou, Dileep George, Miguel Lázaro-GredillaICML 2023 · 2 citations
Related papers
- Hiding Data Helps: On the Benefits of Masking for Sparse CodingMuthu Chidambaram, Chenwei Wu, Yu Cheng, Rong GeICML 2023
- On the Provable Advantage of Unsupervised PretrainingJiawei Ge, Shange Tang, Jianqing Fan, Chi JinICLR 2024 · 23 citations
- Amortised Learning by Wake-SleepLi K. Wenliang, Theodore H. Moskovitz, Heishiro Kanagawa, Maneesh SahaniICML 2020 · 7 citations
- Probability Distribution Collapse: A Critical Bottleneck to Compact Unsupervised Neural Grammar InductionJinwook Park, Kangil KimEMNLP 2025
- Cross-Entropy Is All You Need To Invert the Data Generating ProcessPatrik Reizinger, Alice Bizeul, Attila Juhos, Julia E. Vogt et al.ICLR 2025
