SODA: Bottleneck Diffusion Models for Representation Learning
Drew A. Hudson, Daniel Zoran, Mateusz Malinowski, Andrew K. Lampinen, Andrew Jaegle, James L. McClelland, Loic Matthey, Felix Hill, Alexander Lerchner
Abstract
We introduce SODA, a self-supervised diffusion model, designed for representation learning. The model incorpo-rates an image encoder, which distills a source view into a compact representation, that, in turn, guides the generation of related novel views. We show that by imposing a tight bottleneck between the encoder and a denoising decoder, and leveraging novel view synthesis as a self-supervised ob-jective, we can turn diffusion models into strong represen-tation learners, capable of capturing visual semantics in an unsupervised manner. To the best of our knowledge, SODA is the first diffusion model to succeed at ImageNet linear-probe classification, and, at the same time, it accomplishes reconstruction, editing and synthesis tasks across a wide range of datasets. Further investigation reveals the disentangled nature of its emergent latent space, that serves as an effective interface to control and manipulate the produced images. All in all, we aim to shed light on the exciting and promising potential of diffusion models, not only for image generation, but also for learning rich and robust represen-tations. See our website at soda-diffusion.github.io.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers42
- What matters for Representation Alignment: Global Information or Spatial Structure?Jaskirat Singh, Xingjian Leng, Zongze Wu, Liang Zheng et al.ICLR 2026 · 84 citations
- Schrodinger Bridge Flow for Unpaired Data TranslationValentin De Bortoli, Iryna Korshunova, Andriy Mnih, Arnaud DoucetNeurIPS 2024 · 55 citations
- Diffusion Model with Cross Attention as an Inductive Bias for DisentanglementTao Yang, Cuiling Lan, Yan Lu, Nanning ZhengNeurIPS 2024 · 41 citations
- Learning an Actionable Discrete Diffusion Policy via Large-Scale Actionless Video Pre-TrainingHaoran He, Chenjia Bai, Ling Pan, Weinan Zhang et al.NeurIPS 2024 · 38 citations
- Unleashing the Potential of the Diffusion Model in Few-shot Semantic SegmentationMuzhi Zhu, Yang Liu, Zekai Luo, Chenchen Jing et al.NeurIPS 2024 · 31 citations
Builds on49
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
Related papers
- Denoising Diffusion Autoencoders are Unified Self-supervised LearnersWeilai Xiang, Hongyu Yang, Di Huang, Yunhong WangICCV 2023 · 145 citations
- Boosting Generative Image Modeling via Joint Image-Feature SynthesisTheodoros Kouzelis, Efstathios Karypidis, Ioannis Kakogeorgiou, Spyridon Gidaris et al.NeurIPS 2025 · 47 citations
- Diffusion Based Representation LearningSarthak Mittal, Korbinian Abstreiter, Stefan Bauer, Bernhard Schölkopf et al.ICML 2023 · 71 citations
- DreamTeacher: Pretraining Image Backbones with Deep Generative ModelsDaiqing Li, Huan Ling, Amlan Kar, David Acuna et al.ICCV 2023 · 37 citations
- ProgDiffusion: Progressively Self-encoding Diffusion ModelsZhangkai Wu, Xuhui Fan, Longbing CaoKDD 2025 · 3 citations
