Lune

ACM MM2025Top-tier venue

Noise-Optimized Distribution Distillation for Dataset Condensation

Tongfei Liu, Yufan Liu, Bing Li, Weiming Hu, Yuming Li, Chenguang Ma

2025Year

Abstract

Dataset condensation distills a large dataset into a small synthetic surrogate dataset with similar training efficacy on downstream tasks. Of the existing condensation methods, diffusion-based methods that synthesize surrogate datasets with diffusion models have successfully distilled high-resolution datasets with high training efficacy and satisfactory cross-architectural transferability. However, these methods exhibit a random sampling bias that impairs their performance in dataset condensation settings. We propose a novel dataset condensation method called Noise-Optimized Distribution Distillation (NODD) that mitigates this sampling bias to improve the training performance of synthetic datasets generated with diffusion models. NODD can integrate with existing diffusion-based methods to produce synthetic datasets with enhanced training performance.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get f64a3a0d-d180-4829-9a4d-3739ac281a1f

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines