On the Fragility of Data Attribution When Learning Is Distributed
Xian Gao, Bo Hui, MIN-TE SUN, Wei-Shinn Ku
Abstract
Data attribution has become an important component of pricing, auditing, and governance in machine learning pipelines, yet most attribution methods implicitly assume that attribution values faithfully reflect participants' contributions. We show that this assumption can fail: a single participant in a standard distributed training workflow can substantially inflate its measured attribution value while preserving global utility. Our attribution-first attack uses latent optimization to inject small synthetic batches that preserve utility while exploiting non-IID label coverage and evaluator sensitivities. Across datasets, models, and multiple marginal-utility evaluators, the attack consistently increases the adversary’s attribution value and reshapes the relative attribution structure among benign clients without degrading accuracy or triggering geometry-based defenses. These results show that attribution itself forms a new attack surface and motivate the development of attribution-robust and incentive-compatible scoring mechanisms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dc3117b2-35d6-4744-aa54-97b85b4e4031Builds on41
- Addressing Class Imbalance in Federated LearningLixu Wang, Shichao Xu, Xiao Wang, Qi ZhuAAAI 2021 · 314 citations
- Intriguing Properties of Data Attribution on Diffusion ModelsXiaosen Zheng, Tianyu Pang, Chao Du, Jing Jiang et al.ICLR 2024 · 41 citations
- Training Data Attribution via Approximate UnrollingJuhan Bae, Wu Lin, Jonathan Lorraine, Roger B. GrosseNeurIPS 2024 · 41 citations
- Incentives in Federated Learning: Equilibria, Dynamics, and Mechanisms for Welfare MaximizationAniket Murhekar, Zhuowen Yuan, Bhaskar Ray Chaudhury, Bo Li et al.NeurIPS 2023 · 37 citations
- Fair and Efficient Contribution Valuation for Vertical Federated LearningZhenan Fan, Huang Fang, Xinglu Wang, Zirui Zhou et al.ICLR 2024 · 33 citations
Related papers
- Adversarial Attacks on Data AttributionXinhe Wang, Pingbang Hu, Junwei Deng, Jiaqi W. MaICLR 2025
- Faithful Group Shapley ValueKiljae Lee, Ziqi Liu, Weijing Tang, Yuan ZhangNeurIPS 2025 · 4 citations
- ACE: A Model Poisoning Attack on Contribution Evaluation Methods in Federated LearningZhangchen Xu, Fengqing Jiang, Luyao Niu, Jinyuan Jia et al.USENIX Security 2024 · 11 citations
- Rescaled Influence Functions: Accurate Data Attribution in High DimensionIttai Rubinstein, Samuel B. HopkinsNeurIPS 2025 · 3 citations
- Local Model Poisoning Attacks to Byzantine-Robust Federated LearningMinghong Fang, Xiaoyu Cao, Jinyuan Jia, Neil Zhenqiang GongUSENIX Security 2020
