Scalable Energy-Based Models via Adversarial Training: Unifying Discrimination and Generation
Xuwang Yin, Claire Zhang, Julie Steele, Nir Shavit, Tony T. Wang
Abstract
Simultaneously achieving robust classification and high-fidelity generative modeling within a single framework presents a significant challenge. Hybrid approaches, such as Joint Energy-Based Models (JEM), interpret classifiers as EBMs but are often limited by the instability and poor sample quality inherent in training based on Stochastic Gradient Langevin Dynamics (SGLD). We address these limitations by proposing a novel training framework that integrates adversarial training (AT) principles for both discriminative robustness and stable generative learning. The proposed method introduces three key innovations: (1) the replacement of SGLD-based JEM learning with a stable, AT-based approach that optimizes the energy function through a Binary Cross-Entropy (BCE) loss that discriminates between real data and contrastive samples generated via Projected Gradient Descent (PGD); (2) adversarial training for the discriminative component that enhances classification robustness while implicitly providing the gradient regularization needed for stable EBM training; and (3) a two-stage training strategy that addresses normalization-related instabilities and enables leveraging pretrained robust classifiers, generalizing effectively across architectures. Experiments on CIFAR-10/100 and ImageNet demonstrate that our approach: (1) is the first EBM-based hybrid to scale to high-resolution datasets with high training stability, simultaneously achieving state-of-the-art discriminative and generative performance on ImageNet 256256; (2) uniquely combines generative quality with adversarial robustness, enabling faithful counterfactual explanations; and (3) functions as a competitive standalone generative model, matching state-of-the-art autoregressive models (VAR-d16) and surpassing strong diffusion baselines (ADM-G, LDM-4-G), while additionally supporting diverse image synthesis tasks and compositional generation within a single model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 13cc0e34-2883-4656-be06-553bc35189e9Cited by top-tier papers1
Ask how each one uses itBuilds on24
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- Your Classifier Can Do More: Towards Balancing the Gaps in Classification, Robustness, and GenerationKaichao Jiang, He Wang, Xiaoshuai Hao, Xiulong Yang et al.CVPR 2026 · 1 citation
- Your classifier is secretly an energy based model and you should treat it like oneWill Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud et al.ICLR 2020 · 643 citations
- Towards Understanding the Generative Capability of Adversarially Robust ClassifiersYao Zhu, Jiacheng Ma, Jiacheng Sun, Zewei Chen et al.ICCV 2021 · 30 citations
- A Unified Contrastive Energy-based Model for Understanding the Generative Ability of Adversarial TrainingYifei Wang, Yisen Wang, Jiansheng Yang, Zhouchen LinICLR 2022 · 19 citations
- Towards Bridging the Performance Gaps of Joint Energy-Based ModelsXiulong Yang, Qing Su, Shihao JiCVPR 2023
