CaO2: Rectifying Inconsistencies in Diffusion-Based Dataset Distillation
Haoxuan Wang, Zhenghao Zhao, Junyi Wu, Yuzhang Shang, Gaowen Liu, Yan Yan
Abstract
The recent introduction of diffusion models in dataset distillation has shown promising potential in creating compact surrogate datasets for large, high-resolution target datasets, offering improved efficiency and performance over traditional bi-level/uni-level optimization methods. However, current diffusion-based dataset distillation approaches overlook the evaluation process and exhibit two critical inconsistencies in the distillation process: (1) Objective Inconsistency, where the distillation process diverges from the evaluation objective, and (2) Condition Inconsistency, leading to mismatches between generated images and their corresponding conditions. To resolve these issues, we introduce Condition-aware Optimization with Objective-guided Sampling (), a two-stage diffusion-based framework that aligns the distillation process with the evaluation objective. The first stage employs a probability-informed sample selection pipeline, while the second stage refines the corresponding latent representations to improve conditional likelihood. CaO2 achieves state-of-the-art performance on ImageNet and its subsets, surpassing the best-performing baselines by an average of 2.3% accuracy. Code is available at https://github.com/hatchetProject/CaO2.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1c01f0d1-3a20-4104-8a57-c1928891f65eCited by top-tier papers12
- Fixed Anchors Are Not Enough: Dynamic Retrieval and Persistent Homology for Dataset DistillationMuquan Li, Hang Gou, Yingyi Ma, Rongzheng Wang et al.CVPR 2026 · 11 citations
- HierAmp: Coarse-to-Fine Autoregressive Amplification for Generative Dataset DistillationLin Zhao, Xinru Jiang, Xi Xiao, Qihui Fan et al.CVPR 2026 · 9 citations
- Efficient Multimodal Dataset Distillation via Generative ModelsZhenghao Zhao, Haoxuan Wang, Junyi Wu, Yuzhang Shang et al.NeurIPS 2025 · 7 citations
- Multimodal Distribution Matching for Vision-Language Dataset DistillationJongoh Jeong, Hoyong Kwon, Minseok Kim, Kuk-Jin YoonCVPR 2026 · 3 citations
- Diffusion Models as Dataset Distillation PriorsDuo Su, Huyu Wu, Huanran Chen, Yiming Shi et al.ICLR 2026 · 3 citations
Builds on26
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 806 citations
- Autoregressive Image Generation without Vector QuantizationTianhong Li, Yonglong Tian, He Li, Mingyang Deng et al.NeurIPS 2024 · 758 citations
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 684 citations
Related papers
- Noise-Optimized Distribution Distillation for Dataset CondensationTongfei Liu, Yufan Liu, Bing Li, Weiming Hu et al.ACM MM 2025
- ProtoVAR: Efficient Dataset Distillation via Prototype-Guided Visual Autoregressive ModelingMingyu Wang, Wei JiangICML 2026
- MGD3 : Mode-Guided Dataset Distillation using Diffusion ModelsJeffrey A. Chan-Santiago, Praveen Tirupattur, Gaurav Kumar Nayak, Gaowen Liu et al.ICML 2025
- Influence-Guided Diffusion for Dataset DistillationMingyang Chen, Jiawei Du, Bo Huang, Yi Wang et al.ICLR 2025
- An Adaptive Sampling Framework for Diffusion-based Dataset Distillation with High Fidelity and DiversitySunbeom Jeong, Sehwan Kim, Hyeonggeun Han, Hyungjun Joo et al.AAAI 2026
