Lune

ICLR2026顶会

Ghost in the Cloud: Your Geo-Distributed Large Language Models Training is Easily Manipulated

Zichen Tang, Zhenheng Tang, Gaoning Pan, Buhua Liu, Xin He, Kunfeng Lai, Xiaowen Chu, Bo Li

出版方
2026年份
5顶会引用

摘要

Geo-distributed training and Federated Learning (FL) provide viable solutions to address the substantial data and computational resource needs associated with training large language models (LLMs). However, we empirically demonstrate that a single attacker can significantly compromise the safety alignment of LLMs through malicious training, and existing defenses like robust aggregation or trust-based frameworks fail under this setting due to data heterogeneity. We identify two existing server-side defense strategies that effectively counter naive jailbreak attacks: Task Performance Check (TPC), which filters out model updates with low downstream performance, and Malicious Output Scrutiny (MOS), which detects harmful outputs by prompting uploaded models with malicious queries. To evade both defenses, we design a trigger-based jailbreak variant that preserves downstream performance using a novel regularization method to limit the excessive model updates on jailbreak datasets. We further conceal malicious triggers by mixing the malicious dataset with pseudo-contrastive safety-aligned answers to maintain the original safety alignment. Experiments on several widely used safety-aligned LLMs show that CloudGhost can consistently implant triggers into the global model without degrading downstream performance, achieving 74–93% attack success rate (ASR) and below 5% detection true rate (DTR).

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper5

问问它们各自怎么用它

它引用的顶会 Paper30

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖