An Efficient Framework for Crediting Data Contributors of Diffusion Models
Mingyu Lu, Chris Lin, Chanwoo Kim, Su-In Lee
摘要
As diffusion models are deployed in real-world settings, and their performance is driven by training data, appraising the contribution of data contributors is crucial to creating incentives for sharing quality data and to implementing policies for data compensation. Depending on the use case, model performance corresponds to various global properties of the distribution learned by a diffusion model (e.g., overall aesthetic quality). Hence, here we address the problem of attributing global properties of diffusion models to data contributors. The Shapley value provides a principled approach to valuation by uniquely satisfying game-theoretic axioms of fairness. However, estimating Shapley values for diffusion models is computationally impractical because it requires retraining on many training data subsets corresponding to different contributors and rerunning inference. We introduce a method to efficiently retrain and rerun inference for Shapley value estimation, by leveraging model pruning and fine-tuning. We evaluate the utility of our method with three use cases: (i) image quality for a DDPM trained on a CIFAR dataset, (ii) demographic diversity for an LDM trained on CelebA-HQ, and (iii) aesthetic quality for a Stable Diffusion model LoRA-finetuned on Post-Impressionist artworks. Our results empirically demonstrate that our framework can identify important data contributors across models' global properties, outperforming existing attribution methods for diffusion models 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Region-Level Data Attribution for Text-To-Image Generative ModelsTrong Bang Nguyen, Phi Le Nguyen, Simon Lucey, Minh HoaiICCV 2025 · 被引用 1 次
- SurrogateSHAP: Training-Free Contributor Attribution for Text-to-Image (T2I) ModelsMingYu Lu, Soham Gadgil, Chris Lin, Chanwoo Kim 等ICML 2026
- GUDA: Counterfactual Group-wise Training Data Attribution for Diffusion Models via UnlearningNaoki Murata, Yuhta Takida, Chieh-Hsin Lai, Toshimitsu Uesaka 等ICML 2026
- On the Fragility of Data Attribution When Learning Is DistributedXian Gao, Bo Hui, MIN-TE SUN, Wei-Shinn KuICML 2026
它引用的顶会 Paper23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
相关 Paper
- Intriguing Properties of Data Attribution on Diffusion ModelsXiaosen Zheng, Tianyu Pang, Chao Du, Jing Jiang 等ICLR 2024 · 被引用 41 次
- Data Shapley in One Training RunJiachen T. Wang, Prateek Mittal, Dawn Song, Ruoxi JiaICLR 2025
- Ripple Shapley: Data Influence Attribution in One Federated Training RunDewen Zeng, Wenlong Tian, Haozhao Wang, Jianfeng Lu 等AAAI 2026 · 被引用 1 次
- Collaborative Causal Inference with Fair IncentivesRui Qiao, Xinyi Xu, Bryan Kian Hsiang LowICML 2023 · 被引用 8 次
- Diffusion Attribution Score: Evaluating Training Data Influence in Diffusion ModelsJinxu Lin, Linwei Tao, Minjing Dong, Chang XuICLR 2025
