Analyzing and Improving the Optimization Landscape of Noise-Contrastive Estimation
Bingbin Liu, Elan Rosenfeld, Pradeep Kumar Ravikumar, Andrej Risteski
Abstract
Noise-contrastive estimation (NCE) is a statistically consistent method for learning unnormalized probabilistic models. It has been empirically observed that the choice of the noise distribution is crucial for NCE's performance. However, such observations have never been made formal or quantitative. In fact, it is not even clear whether the difficulties arising from a poorly chosen noise distribution are statistical or algorithmic in nature. In this work, we formally pinpoint reasons for NCE's poor performance when an inappropriate noise distribution is used. Namely, we prove these challenges arise due to an ill-behaved (more precisely, flat) loss landscape. To address this, we introduce a variant of NCE called"eNCE"which uses an exponential loss and for which normalized gradient descent addresses the landscape issues provably when the target and noise distributions are in a given exponential family.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- The Mechanism of Prediction Head in Non-contrastive Self-supervised LearningZixin Wen, Yuanzhi LiNeurIPS 2022 · 44 citations
- Provable benefits of annealing for estimating normalizing constants: Importance Sampling, Noise-Contrastive Estimation, and beyondOmar Chehab, Aapo Hyvärinen, Andrej RisteskiNeurIPS 2023 · 18 citations
- ``Noisier'’ Noise Contrastive Estimation is (Almost) Maximum LikelihoodPeiyu Yu, Dinghuai Zhang, Hengzhi He, Xiaojian Ma et al.ICLR 2026 · 11 citations
- BoltzNCE: Learning likelihoods for Boltzmann Generation with Stochastic Interpolants and Noise Contrastive EstimationRishal Aggarwal, Jacky Chen, Nicholas M. Boffi, David KoesNeurIPS 2025 · 10 citations
- Learning Unnormalized Statistical Models via Compositional OptimizationWei Jiang, Jiayu Qin, Lingyu Wu, Changyou Chen et al.ICML 2023 · 8 citations
Builds on7
- Data-Efficient Image Recognition with Contrastive Predictive CodingOlivier J. HénaffICML 2020 · 1,553 citations
- Your classifier is secretly an energy based model and you should treat it like oneWill Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud et al.ICLR 2020 · 643 citations
- A Mutual Information Maximization Perspective of Language Representation LearningLingpeng Kong, Cyprien de Masson d'Autume, Lei Yu, Wang Ling et al.ICLR 2020 · 179 citations
- Telescoping Density-Ratio EstimationBenjamin Rhodes, Kai Xu, Michael U. GutmannNeurIPS 2020 · 148 citations
- Training Deep Energy-Based Models with f-Divergence MinimizationLantao Yu, Yang Song, Jiaming Song, Stefano ErmonICML 2020 · 50 citations
Related papers
- A Unified View on Learning Unnormalized Distributions via Noise-Contrastive EstimationJongha Jon Ryu, Abhin Shah, Gregory W. WornellICML 2025
- Pitfalls of Gaussians as a noise distribution in NCEHolden Lee, Chirag Pabbaraju, Anish Prasad Sevekari, Andrej RisteskiICLR 2023
- Autoencoding Under Normalization ConstraintsSangwoong Yoon, Yung-Kyun Noh, Frank Chongwoo ParkICML 2021 · 44 citations
- Adaptive Multi-stage Density Ratio Estimation for Learning Latent Space Energy-based ModelZhisheng Xiao, Tian HanNeurIPS 2022 · 24 citations
- Generalized Jensen-Shannon Divergence Loss for Learning with Noisy LabelsErik Englesson, Hossein AzizpourNeurIPS 2021 · 170 citations
