JEM++: Improved Techniques for Training JEM
Xiulong Yang, Shihao Ji
摘要
Joint Energy-based Model (JEM) [17] is a recently proposed hybrid model that retains strong discriminative power of modern CNN classifiers, while generating samples rivaling the quality of GAN-based approaches. In this paper, we propose a variety of new training procedures and architecture features to improve JEM's accuracy, training stability, and speed altogether. 1) We propose a proximal SGLD to generate samples in the proximity of samples from previous step, which improves the stability. 2) We further treat the approximate maximum likelihood learning of EBM as a multi-step differential game, and extend the YOPO framework [59] to cut out redundant calculations during backpropagation, which accelerates the training substantially. 3) Rather than initializing SGLD chain from random noise, we introduce a new informative initialization that samples from a distribution estimated from training data. 4) This informative initialization allows us to enable batch normalization in JEM, which further releases the power of modern CNN architectures for hybrid modeling. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Energy-Based Contrastive Learning of Visual RepresentationsBeomsu Kim, Jong Chul YeNeurIPS 2022 · 被引用 23 次
- TEA: Test-Time Energy AdaptationYige Yuan, Bingbing Xu, Liang Hou, Fei Sun 等CVPR 2024 · 被引用 8 次
- Guiding Energy-based Models via Contrastive Latent VariablesHankook Lee, Jongheon Jeong, Sejun Park, Jinwoo ShinICLR 2023 · 被引用 4 次
- Exploring Structured Semantic Priors Underlying Diffusion Score for Test-time AdaptationMingjia Li, Shuang Li, Tongrui Su, Longhui Yuan 等NeurIPS 2024 · 被引用 2 次
- Scalable Energy-Based Models via Adversarial Training: Unifying Discrimination and GenerationXuwang Yin, Claire Zhang, Julie Steele, Nir Shavit 等ICLR 2026 · 被引用 1 次
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- NVAE: A Deep Hierarchical Variational AutoencoderArash Vahdat, Jan KautzNeurIPS 2020 · 被引用 1,141 次
- Your classifier is secretly an energy based model and you should treat it like oneWill Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud 等ICLR 2020 · 被引用 643 次
- Certified Defenses for Adversarial PatchesPing-yeh Chiang, Renkun Ni, Ahmed Abdelkader, Chen Zhu 等ICLR 2020 · 被引用 194 次
相关 Paper
- Towards Bridging the Performance Gaps of Joint Energy-Based ModelsXiulong Yang, Qing Su, Shihao JiCVPR 2023
- Your Classifier Can Do More: Towards Balancing the Gaps in Classification, Robustness, and GenerationKaichao Jiang, He Wang, Xiaoshuai Hao, Xiulong Yang 等CVPR 2026 · 被引用 1 次
- No MCMC for me: Amortized sampling for fast and stable training of energy-based modelsWill Sussman Grathwohl, Jacob Jin Kelly, Milad Hashemi, Mohammad Norouzi 等ICLR 2021 · 被引用 75 次
- Learning Energy-based Model via Dual-MCMC TeachingJiali Cui, Tian HanNeurIPS 2023 · 被引用 14 次
- Learning Energy-Based Models by Cooperative Diffusion Recovery LikelihoodYaxuan Zhu, Jianwen Xie, Ying Nian Wu, Ruiqi GaoICLR 2024 · 被引用 18 次
