Contrastive Diffusion Alignment: Learning Structured Latents for Controllable Generation
Ruchi Sandilya, Sumaira Perez, Charles Lynch, Lindsay Victoria, Benjamin Zebley, Derrick Buchanan, Mahendra Bhati, Nolan Williams, Timothy Spellman, FAITH GUNNING, Conor Liston, Logan Grosenick
Abstract
Diffusion models excel at generation, but their latent spaces are high dimensional and not explicitly organized for interpretation or control. We introduce ConDA (Contrastive Diffusion Alignment), a plug-and-play geometry layer that applies contrastive learning to pretrained diffusion latents using auxiliary variables (e.g., time, stimulation parameters, facial action units). ConDA learns a low-dimensional embedding whose directions align with underlying dynamical factors, consistent with recent contrastive learning results on structured and disentangled representations. In this embedding, simple nonlinear trajectories support smooth interpolation, extrapolation, and counterfactual editing while rendering remains in the original diffusion space. ConDA separates editing and rendering by lifting embedding trajectories back to diffusion latents with a neighborhood-preserving kNN decoder and is robust across inversion solvers. Across fluid dynamics, neural calcium imaging, therapeutic neurostimulation, facial expression dynamics, and monkey motor cortex activity, ConDA yields more interpretable and controllable latent structure than linear traversals and conditioning-based baselines, indicating that diffusion latents encode dynamics-relevant structure that can be exploited by an explicit contrastive geometry layer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 022e9605-cce3-4b8c-bca6-c932f5146299Cited by top-tier papers1
Ask how each one uses itBuilds on22
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- InfoDiffusion: Representation Learning Using Information Maximizing Diffusion ModelsYingheng Wang, Yair Schiff, Aaron Gokaslan, Weishen Pan et al.ICML 2023 · 64 citations
- Causal-Adapter: Taming Text-to-Image Diffusion for Faithful Counterfactual GenerationLei Tong, Zhihua Liu, Chaochao Lu, Dino Oglic et al.ICML 2026 · 2 citations
- Disentangled Hierarchical VAE for 3D Human-Human Interaction GenerationZichen Geng, Zeeshan Hayder, Bo Miao, Jian Liu et al.ICLR 2026 · 3 citations
- Latent Diffusion for Neural Spiking DataJaivardhan Kapoor, Auguste Schulz, Julius Vetter, Felix Pei et al.NeurIPS 2024 · 24 citations
- Isometric Representation Learning for Disentangled Latent Space of Diffusion ModelsJaehoon Hahm, Junho Lee, Sunghyun Kim, Joonseok LeeICML 2024 · 21 citations
