Dataset Distillation by Influence Matching
Haoru Tan, Wang Wang, Sitong Wu, Xiuzhe Wu, Yang-Tian Sun, Chirui Chang, Shaofeng Zhang, Xiaojuan Qi
Abstract
We revisit dataset distillation from an outcome-centric perspective. Rather than aligning process surrogates (per-step gradients or training trajectories), Influence Matching (Inf-Match) aligns the final outcome of training: it learns a compact synthetic set whose effect on the converged parameters matches that of the full dataset. Concretely, we introduce a fully differentiable, sample-level influence estimator that quantifies parameter shifts from adding or removing data, without time-consuming inverse-Hessian products or convexity assumptions. The estimator runs in linear time by unrolling the optimization dynamics and applying a first-order Taylor approximation. We then learn the synthetic set by minimizing the mismatch between its influence and that of the real dataset, yielding outcome alignment rather than heuristic process imitation. Inf-Match delivers the best accuracy across standard classification benchmarks. For instance, on Tiny-ImageNet (IPC=10), Inf-Match attains 31.5%, a +4.7% improvement over NCFM. Beyond classification, Inf-Match scales to vision-language distillation on Flickr30K, outperforming strong process-matching baselines. For instance, with 200 to 1000 synthetic samples, our method achieved a leading impressive average on image/text retrieval tasks, higher than NCFM by 2.5%. The code will be released via https://github.com/hrtan/infmatch.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on39
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 784 citations
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 684 citations
- High-Performance Large-Scale Image Recognition Without NormalizationAndy Brock, Soham De, Samuel L. Smith, Karen SimonyanICML 2021 · 613 citations
- Dataset Condensation with Differentiable Siamese AugmentationBo Zhao, Hakan BilenICML 2021 · 390 citations
- Dataset Meta-Learning from Kernel Ridge-RegressionTimothy Nguyen, Zhourong Chen, Jaehoon LeeICLR 2021 · 307 citations
Related papers
- CovMatch: Cross-Covariance Guided Multimodal Dataset Distillation with Trainable Text EncoderYongmin Lee, Hye Won ChungNeurIPS 2025 · 2 citations
- SelMatch: Effectively Scaling Up Dataset Distillation via Selection-Based Initialization and Partial Updates by Trajectory MatchingYongmin Lee, Hye Won ChungICML 2024 · 26 citations
- Towards Lossless Dataset Distillation via Difficulty-Aligned Trajectory MatchingZiyao Guo, Kai Wang, George Cazenavette, Hui Li et al.ICLR 2024 · 142 citations
- Minimizing the Accumulated Trajectory Error to Improve Dataset DistillationJiawei Du, Yidi Jiang, Vincent Y. F. Tan, Joey Tianyi Zhou et al.CVPR 2023
- DataDAM: Efficient Dataset Distillation with Attention MatchingAhmad Sajedi, Samir Khaki, Ehsan Amjadian, Lucy Z. Liu et al.ICCV 2023 · 106 citations
