JEM++: Improved Techniques for Training JEM
Xiulong Yang, Shihao Ji
Abstract
Joint Energy-based Model (JEM) [17] is a recently proposed hybrid model that retains strong discriminative power of modern CNN classifiers, while generating samples rivaling the quality of GAN-based approaches. In this paper, we propose a variety of new training procedures and architecture features to improve JEM's accuracy, training stability, and speed altogether. 1) We propose a proximal SGLD to generate samples in the proximity of samples from previous step, which improves the stability. 2) We further treat the approximate maximum likelihood learning of EBM as a multi-step differential game, and extend the YOPO framework [59] to cut out redundant calculations during backpropagation, which accelerates the training substantially. 3) Rather than initializing SGLD chain from random noise, we introduce a new informative initialization that samples from a distribution estimated from training data. 4) This informative initialization allows us to enable batch normalization in JEM, which further releases the power of modern CNN architectures for hybrid modeling. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 37433914-5229-418c-bcef-6e67e06d0893Cited by top-tier papers11
- Energy-Based Contrastive Learning of Visual RepresentationsBeomsu Kim, Jong Chul YeNeurIPS 2022 · 23 citations
- TEA: Test-Time Energy AdaptationYige Yuan, Bingbing Xu, Liang Hou, Fei Sun et al.CVPR 2024 · 8 citations
- Guiding Energy-based Models via Contrastive Latent VariablesHankook Lee, Jongheon Jeong, Sejun Park, Jinwoo ShinICLR 2023 · 4 citations
- Exploring Structured Semantic Priors Underlying Diffusion Score for Test-time AdaptationMingjia Li, Shuang Li, Tongrui Su, Longhui Yuan et al.NeurIPS 2024 · 2 citations
- Scalable Energy-Based Models via Adversarial Training: Unifying Discrimination and GenerationXuwang Yin, Claire Zhang, Julie Steele, Nir Shavit et al.ICLR 2026 · 1 citation
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- NVAE: A Deep Hierarchical Variational AutoencoderArash Vahdat, Jan KautzNeurIPS 2020 · 1,141 citations
- Your classifier is secretly an energy based model and you should treat it like oneWill Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud et al.ICLR 2020 · 643 citations
- Certified Defenses for Adversarial PatchesPing-yeh Chiang, Renkun Ni, Ahmed Abdelkader, Chen Zhu et al.ICLR 2020 · 194 citations
Related papers
- Towards Bridging the Performance Gaps of Joint Energy-Based ModelsXiulong Yang, Qing Su, Shihao JiCVPR 2023
- Your Classifier Can Do More: Towards Balancing the Gaps in Classification, Robustness, and GenerationKaichao Jiang, He Wang, Xiaoshuai Hao, Xiulong Yang et al.CVPR 2026 · 1 citation
- No MCMC for me: Amortized sampling for fast and stable training of energy-based modelsWill Sussman Grathwohl, Jacob Jin Kelly, Milad Hashemi, Mohammad Norouzi et al.ICLR 2021 · 75 citations
- Learning Energy-based Model via Dual-MCMC TeachingJiali Cui, Tian HanNeurIPS 2023 · 14 citations
- Learning Energy-Based Models by Cooperative Diffusion Recovery LikelihoodYaxuan Zhu, Jianwen Xie, Ying Nian Wu, Ruiqi GaoICLR 2024 · 18 citations
