CaO2: Rectifying Inconsistencies in Diffusion-Based Dataset Distillation
Haoxuan Wang, Zhenghao Zhao, Junyi Wu, Yuzhang Shang, Gaowen Liu, Yan Yan
摘要
The recent introduction of diffusion models in dataset distillation has shown promising potential in creating compact surrogate datasets for large, high-resolution target datasets, offering improved efficiency and performance over traditional bi-level/uni-level optimization methods. However, current diffusion-based dataset distillation approaches overlook the evaluation process and exhibit two critical inconsistencies in the distillation process: (1) Objective Inconsistency, where the distillation process diverges from the evaluation objective, and (2) Condition Inconsistency, leading to mismatches between generated images and their corresponding conditions. To resolve these issues, we introduce Condition-aware Optimization with Objective-guided Sampling (), a two-stage diffusion-based framework that aligns the distillation process with the evaluation objective. The first stage employs a probability-informed sample selection pipeline, while the second stage refines the corresponding latent representations to improve conditional likelihood. CaO2 achieves state-of-the-art performance on ImageNet and its subsets, surpassing the best-performing baselines by an average of 2.3% accuracy. Code is available at https://github.com/hatchetProject/CaO2.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Fixed Anchors Are Not Enough: Dynamic Retrieval and Persistent Homology for Dataset DistillationMuquan Li, Hang Gou, Yingyi Ma, Rongzheng Wang 等CVPR 2026 · 被引用 11 次
- HierAmp: Coarse-to-Fine Autoregressive Amplification for Generative Dataset DistillationLin Zhao, Xinru Jiang, Xi Xiao, Qihui Fan 等CVPR 2026 · 被引用 9 次
- Efficient Multimodal Dataset Distillation via Generative ModelsZhenghao Zhao, Haoxuan Wang, Junyi Wu, Yuzhang Shang 等NeurIPS 2025 · 被引用 7 次
- Multimodal Distribution Matching for Vision-Language Dataset DistillationJongoh Jeong, Hoyong Kwon, Minseok Kim, Kuk-Jin YoonCVPR 2026 · 被引用 3 次
- Diffusion Models as Dataset Distillation PriorsDuo Su, Huyu Wu, Huanran Chen, Yiming Shi 等ICLR 2026 · 被引用 3 次
它引用的顶会 Paper26
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 被引用 806 次
- Autoregressive Image Generation without Vector QuantizationTianhong Li, Yonglong Tian, He Li, Mingyang Deng 等NeurIPS 2024 · 被引用 758 次
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 被引用 684 次
相关 Paper
- Noise-Optimized Distribution Distillation for Dataset CondensationTongfei Liu, Yufan Liu, Bing Li, Weiming Hu 等ACM MM 2025
- ProtoVAR: Efficient Dataset Distillation via Prototype-Guided Visual Autoregressive ModelingMingyu Wang, Wei JiangICML 2026
- MGD3 : Mode-Guided Dataset Distillation using Diffusion ModelsJeffrey A. Chan-Santiago, Praveen Tirupattur, Gaurav Kumar Nayak, Gaowen Liu 等ICML 2025
- Influence-Guided Diffusion for Dataset DistillationMingyang Chen, Jiawei Du, Bo Huang, Yi Wang 等ICLR 2025
- An Adaptive Sampling Framework for Diffusion-based Dataset Distillation with High Fidelity and DiversitySunbeom Jeong, Sehwan Kim, Hyeonggeun Han, Hyungjun Joo 等AAAI 2026
