Empirical Study of the Benefits of Overparameterization in Learning Latent Variable Models
Rares-Darius Buhai, Yoni Halpern, Yoon Kim, Andrej Risteski, David A. Sontag
摘要
One of the most surprising and exciting discoveries in supervised learning was the benefit of overparameterization (i.e. training a very large model) to improving the optimization landscape of a problem, with minimal effect on statistical performance (i.e. generalization). In contrast, unsupervised settings have been under-explored, despite the fact that it was observed that overparameterization can be helpful as early as Dasgupta & Schulman (2007) . We perform an empirical study of different aspects of overparameterization in unsupervised learning of latent variable models via synthetic and semi-synthetic experiments. We discuss benefits to different metrics of success (recovering the parameters of the ground-truth model, held-out log-likelihood), sensitivity to variations of the training algorithm, and behavior as the amount of overparameterization increases. We find that across a variety of models (noisy-OR networks, sparse coding, probabilistic context-free grammars) and training algorithms (variational inference, alternating minimization, expectation-maximization), overparameterization can significantly increase the number of ground truth latent variables recovered. The code and supporting files for the experiments are located at https: //github.com/clinicalml/overparam .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Sequence-to-Sequence Learning with Latent Neural GrammarsYoon KimNeurIPS 2021 · 被引用 44 次
- Can Looped Transformers Learn to Implement Multi-step Gradient Descent for In-context Learning?Khashayar Gatmiry, Nikunj Saunshi, Sashank J. Reddi, Stefanie Jegelka 等ICML 2024 · 被引用 43 次
- Compressible Dynamics in Deep Overparameterized Low-Rank Learning & AdaptationCan Yaras, Peng Wang, Laura Balzano, Qing QuICML 2024 · 被引用 29 次
- The Lazy Neuron Phenomenon: On Emergence of Activation Sparsity in TransformersZonglin Li, Chong You, Srinadh Bhojanapalli, Daliang Li 等ICLR 2023 · 被引用 10 次
- Learning Noisy OR Bayesian Networks with Max-Product Belief PropagationAntoine Dedieu, Guangyao Zhou, Dileep George, Miguel Lázaro-GredillaICML 2023 · 被引用 2 次
相关 Paper
- Hiding Data Helps: On the Benefits of Masking for Sparse CodingMuthu Chidambaram, Chenwei Wu, Yu Cheng, Rong GeICML 2023
- On the Provable Advantage of Unsupervised PretrainingJiawei Ge, Shange Tang, Jianqing Fan, Chi JinICLR 2024 · 被引用 23 次
- Amortised Learning by Wake-SleepLi K. Wenliang, Theodore H. Moskovitz, Heishiro Kanagawa, Maneesh SahaniICML 2020 · 被引用 7 次
- Probability Distribution Collapse: A Critical Bottleneck to Compact Unsupervised Neural Grammar InductionJinwook Park, Kangil KimEMNLP 2025
- Cross-Entropy Is All You Need To Invert the Data Generating ProcessPatrik Reizinger, Alice Bizeul, Attila Juhos, Julia E. Vogt 等ICLR 2025
