Analyzing and Improving the Optimization Landscape of Noise-Contrastive Estimation
Bingbin Liu, Elan Rosenfeld, Pradeep Kumar Ravikumar, Andrej Risteski
摘要
Noise-contrastive estimation (NCE) is a statistically consistent method for learning unnormalized probabilistic models. It has been empirically observed that the choice of the noise distribution is crucial for NCE's performance. However, such observations have never been made formal or quantitative. In fact, it is not even clear whether the difficulties arising from a poorly chosen noise distribution are statistical or algorithmic in nature. In this work, we formally pinpoint reasons for NCE's poor performance when an inappropriate noise distribution is used. Namely, we prove these challenges arise due to an ill-behaved (more precisely, flat) loss landscape. To address this, we introduce a variant of NCE called"eNCE"which uses an exponential loss and for which normalized gradient descent addresses the landscape issues provably when the target and noise distributions are in a given exponential family.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- The Mechanism of Prediction Head in Non-contrastive Self-supervised LearningZixin Wen, Yuanzhi LiNeurIPS 2022 · 被引用 44 次
- Provable benefits of annealing for estimating normalizing constants: Importance Sampling, Noise-Contrastive Estimation, and beyondOmar Chehab, Aapo Hyvärinen, Andrej RisteskiNeurIPS 2023 · 被引用 18 次
- ``Noisier'’ Noise Contrastive Estimation is (Almost) Maximum LikelihoodPeiyu Yu, Dinghuai Zhang, Hengzhi He, Xiaojian Ma 等ICLR 2026 · 被引用 11 次
- BoltzNCE: Learning likelihoods for Boltzmann Generation with Stochastic Interpolants and Noise Contrastive EstimationRishal Aggarwal, Jacky Chen, Nicholas M. Boffi, David KoesNeurIPS 2025 · 被引用 10 次
- Learning Unnormalized Statistical Models via Compositional OptimizationWei Jiang, Jiayu Qin, Lingyu Wu, Changyou Chen 等ICML 2023 · 被引用 8 次
它引用的顶会 Paper7
- Data-Efficient Image Recognition with Contrastive Predictive CodingOlivier J. HénaffICML 2020 · 被引用 1,553 次
- Your classifier is secretly an energy based model and you should treat it like oneWill Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud 等ICLR 2020 · 被引用 643 次
- A Mutual Information Maximization Perspective of Language Representation LearningLingpeng Kong, Cyprien de Masson d'Autume, Lei Yu, Wang Ling 等ICLR 2020 · 被引用 179 次
- Telescoping Density-Ratio EstimationBenjamin Rhodes, Kai Xu, Michael U. GutmannNeurIPS 2020 · 被引用 148 次
- Training Deep Energy-Based Models with f-Divergence MinimizationLantao Yu, Yang Song, Jiaming Song, Stefano ErmonICML 2020 · 被引用 50 次
相关 Paper
- A Unified View on Learning Unnormalized Distributions via Noise-Contrastive EstimationJongha Jon Ryu, Abhin Shah, Gregory W. WornellICML 2025
- Pitfalls of Gaussians as a noise distribution in NCEHolden Lee, Chirag Pabbaraju, Anish Prasad Sevekari, Andrej RisteskiICLR 2023
- Autoencoding Under Normalization ConstraintsSangwoong Yoon, Yung-Kyun Noh, Frank Chongwoo ParkICML 2021 · 被引用 44 次
- Adaptive Multi-stage Density Ratio Estimation for Learning Latent Space Energy-based ModelZhisheng Xiao, Tian HanNeurIPS 2022 · 被引用 24 次
- Generalized Jensen-Shannon Divergence Loss for Learning with Noisy LabelsErik Englesson, Hossein AzizpourNeurIPS 2021 · 被引用 170 次
