Intriguing Properties of Generative Classifiers
Priyank Jaini, Kevin Clark, Robert Geirhos
Abstract
What is the best paradigm to recognize objects -- discriminative inference (fast but potentially prone to shortcut learning) or using a generative model (slow but potentially more robust)? We build on recent advances in generative modeling that turn text-to-image models into classifiers. This allows us to study their behavior and to compare them against discriminative models and human psychophysical data. We report four intriguing emergent properties of generative classifiers: they show a record-breaking human-like shape bias (99% for Imagen), near human-level out-of-distribution accuracy, state-of-the-art alignment with human classification errors, and they understand certain perceptual illusions. Our results indicate that while the current dominant paradigm for modeling human object recognition is discriminative inference, zero-shot generative models approximate human object recognition data surprisingly well.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a9d1cbda-3f92-496b-abcb-faf40886a4d2Cited by top-tier papers21
- Adversarial Robustness Limits via Scaling-Law and Human-Alignment StudiesBrian R. Bartoldson, James Diffenderfer, Konstantinos Parasyris, Bhavya KailkhuraICML 2024 · 45 citations
- Visual Anagrams Reveal Hidden Differences in Holistic Shape Processing Across Vision ModelsFenil R. Doshi, Thomas Fel, Talia Konkle, George A. AlvarezNeurIPS 2025 · 5 citations
- Your VAR Model is Secretly an Efficient and Explainable Generative ClassifierYi-Chung Chen, David I. Inouye, Jing GaoICLR 2026 · 2 citations
- Exploring Structured Semantic Priors Underlying Diffusion Score for Test-time AdaptationMingjia Li, Shuang Li, Tongrui Su, Longhui Yuan et al.NeurIPS 2024 · 2 citations
- Leveraging Prior Knowledge of Diffusion Model for Person SearchGiyeol Kim, Sooyoung Yang, Jihyong Oh, Myungjoo Kang et al.ICCV 2025 · 2 citations
Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- Text-to-Image Diffusion Models are Zero Shot ClassifiersKevin Clark, Priyank JainiNeurIPS 2023 · 192 citations
- A Universal Discriminator for Zero-Shot GeneralizationHaike Xu, Zongyu Lin, Jing Zhou, Yanan Zheng et al.ACL 2023 · 6 citations
- Compositional Scene Understanding through Inverse Generative ModelingYanbo Wang, Justin Dauwels, Yilun DuICML 2025
- Generative Multi-modal Models are Good Class-Incremental LearnersXusheng Cao, Haori Lu, Linlan Huang, Xialei Liu et al.CVPR 2024
- Is Synthetic Data from Generative Models Ready for Image Recognition?Ruifei He, Shuyang Sun, Xin Yu, Chuhui Xue et al.ICLR 2023 · 56 citations
