Advancing Image Classification with Discrete Diffusion Classification Modeling
Omer Belhasin, Shelly Golan, Ran El-Yaniv, Michael Elad
Abstract
Image classification is a well-studied task in computer vision, and yet it remains challenging under high-uncertainty conditions, such as when input images are corrupted or training data are limited. Conventional classification approaches typically train models to directly predict class labels from input images, but this might lead to suboptimal performance in such scenarios. To address this issue, we propose Discrete Diffusion Classification Modeling (DiDiCM), a novel framework that leverages a diffusionbased procedure to model the posterior distribution of class labels conditioned on the input image. DiDiCM supports diffusion-based predictions either on class probabilities or on discrete class labels, providing flexibility in computation and memory trade-offs. We conduct a comprehensive empirical study demonstrating the superior performance of DiDiCM over standard classifiers, showing that a few diffusion iterations achieve higher classification accuracy on the ImageNet dataset compared to baselines, with accuracy gains increasing as the task becomes more challenging. We release our code at https://github.com/ omerb01/didicm.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 362a0453-fa43-4758-914c-6baea7e1d0c0Cited by top-tier papers1
Ask how each one uses itBuilds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
Related papers
- An Efficient Framework for Enhancing Discriminative Models via Diffusion TechniquesChunxiao Li, Xiaoxiao Wang, Boming Miao, Chuanlong Xie et al.AAAI 2025 · 2 citations
- DisCo-Diff: Enhancing Continuous Diffusion Models with Discrete LatentsYilun Xu, Gabriele Corso, Tommi S. Jaakkola, Arash Vahdat et al.ICML 2024 · 22 citations
- CARD: Classification and Regression Diffusion ModelsXizewen Han, Huangjie Zheng, Mingyuan ZhouNeurIPS 2022 · 185 citations
- Joint Enhancement and Classification using Coupled Diffusion Models of Signals and LogitsGilad Nurko, Roi Benita, Yehoshua Dissen, Tomohiro Nakatani et al.ICML 2026
- Efficient Image-to-Image Diffusion Classifier for Adversarial RobustnessHefei Mei, Minjing Dong, Chang XuAAAI 2025 · 2 citations
