Diffusion Reconstruction-based Data Likelihood Estimation for Core-Set Selection
Mingyang Chen, Jiawei Du, Bo Huang, Yi Wang, Xiaobo Zhang, Wei Wang
摘要
Existing core-set selection methods predominantly rely on heuristic scoring signals such as training dynamics or model uncertainty, lacking explicit modeling of data likelihood. This omission may hinder the constructed subset from capturing subtle yet critical distributional structures that underpin effective model training. In this work, we propose a novel, theoretically grounded approach that leverages diffusion models to estimate data likelihood via reconstruction deviation induced by partial reverse denoising. Specifically, we establish a formal connection between reconstruction error and data likelihood, grounded in the Evidence Lower Bound (ELBO) of Markovian diffusion processes, thereby enabling a principled, distribution-aware scoring criterion for data selection. Complementarily, we introduce an efficient information-theoretic method to identify the optimal reconstruction timestep, ensuring that the deviation provides a reliable signal indicative of underlying data likelihood. Extensive experiments on Ima-geNet demonstrate that reconstruction deviation offers an effective scoring criterion, consistently outperforming existing baselines across selection ratios, and closely matching fulldata training using only 50% of the data. Further analysis shows that the likelihood-informed nature of our score reveals informative insights in data selection, shedding light on the interplay between data distributional characteristics and model learning preferences. The code is available at https://github.com/mchen725/DRD .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 被引用 784 次
- What Neural Networks Memorize and Why: Discovering the Long Tail via Influence EstimationVitaly Feldman, Chiyuan ZhangNeurIPS 2020 · 被引用 674 次
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 被引用 494 次
相关 Paper
- On Density Estimation with Diffusion ModelsDiederik P. Kingma, Tim Salimans, Ben Poole, Jonathan HoNeurIPS 2021 · 被引用 56 次
- Information-Theoretic Discrete DiffusionMoongyu Jeon, Sangwoo Shin, Dongjae Jeon, Albert NoNeurIPS 2025 · 被引用 4 次
- An Expectation-Maximization Algorithm for Training Clean Diffusion Models from Corrupted ObservationsWeimin Bai, Yifei Wang, Wenzheng Chen, He SunNeurIPS 2024 · 被引用 21 次
- Fast and Scalable Analytical DiffusionXinyi Shang, Peng Sun, Jingyu Lin, Zhiqiang ShenICML 2026 · 被引用 1,092 次
- Maximum Likelihood Training of Score-Based Diffusion ModelsYang Song, Conor Durkan, Iain Murray, Stefano ErmonNeurIPS 2021 · 被引用 958 次
