Ghost in the Cloud: Your Geo-Distributed Large Language Models Training is Easily Manipulated
Zichen Tang, Zhenheng Tang, Gaoning Pan, Buhua Liu, Xin He, Kunfeng Lai, Xiaowen Chu, Bo Li
Abstract
Geo-distributed training and Federated Learning (FL) provide viable solutions to address the substantial data and computational resource needs associated with training large language models (LLMs). However, we empirically demonstrate that a single attacker can significantly compromise the safety alignment of LLMs through malicious training, and existing defenses like robust aggregation or trust-based frameworks fail under this setting due to data heterogeneity. We identify two existing server-side defense strategies that effectively counter naive jailbreak attacks: Task Performance Check (TPC), which filters out model updates with low downstream performance, and Malicious Output Scrutiny (MOS), which detects harmful outputs by prompting uploaded models with malicious queries. To evade both defenses, we design a trigger-based jailbreak variant that preserves downstream performance using a novel regularization method to limit the excessive model updates on jailbreak datasets. We further conceal malicious triggers by mixing the malicious dataset with pseudo-contrastive safety-aligned answers to maintain the original safety alignment. Experiments on several widely used safety-aligned LLMs show that CloudGhost can consistently implant triggers into the global model without degrading downstream performance, achieving 74–93% attack success rate (ASR) and below 5% detection true rate (DTR).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a592176a-1afb-4b1f-b509-c1b435acb8e5Cited by top-tier papers5
- ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM InferenceXiang Liu, Zhenheng Tang, Peijie Dong, Zeyu Li et al.NeurIPS 2025 · 71 citations
- Landscape of Thoughts: Visualizing the Reasoning Process of Large Language ModelsZhanke Zhou, Zhaocheng Zhu, Xuan Li, Mikhail Galkin et al.ICLR 2026 · 28 citations
- Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache CompressionXiang Liu, Zhenheng Tang, Hong Chen, Peijie Dong et al.ICML 2026 · 16 citations
- DiffAdapt: Difficulty-Adaptive Reasoning for Token-Efficient LLM InferenceXiang Liu, Xuming Hu, Xiaowen Chu, Eunsol ChoiICLR 2026 · 15 citations
- Identifying and Mitigating Errors in Gradient Aggregation of Distributed Data Parallel TrainingZhenheng Tang, Junlin Huang, Zichen TANG, Xueze Kang et al.ICML 2026
Builds on30
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen et al.ICLR 2024 · 1,104 citations
- An Explanation of In-context Learning as Implicit Bayesian InferenceSang Michael Xie, Aditi Raghunathan, Percy Liang, Tengyu MaICLR 2022 · 1,030 citations
- Merging Models with Fisher-Weighted AveragingMichael Matena, Colin RaffelNeurIPS 2022 · 741 citations
Related papers
- Towards Robust Multimodal Large Language Models Against Jailbreak AttacksZiyi Yin, Yuanpu Cao, Han Liu, Ting Wang et al.CVPR 2026 · 5 citations
- One Model Transfer to All: On Robust Jailbreak Prompts Generation against LLMsLinbao Li, Yannan Liu, Daojing He, Yu LiICLR 2025
- SelfDefend: LLMs Can Defend Themselves against Jailbreaking in a Practical MannerXunguang Wang, Daoyuan Wu, Zhenlan Ji, Zongjie Li et al.USENIX Security 2025
- Patcher: Post-Hoc Patching of Backdoored Large Language ModelsAnjun Gao, Yueyang Quan, Yufei Xia, Zhuqing Liu et al.USENIX Security 2026
- Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time AlignmentSoumya Suvra Ghosal, Souradip Chakraborty, Vaibhav Singh, Tianrui Guan et al.CVPR 2025
