Exploiting Leaderboards for Large-Scale Distribution of Malicious Models
Anshuman Suri, Harsh Chaudhari, Yuefeng Peng, Ali Naseh, Alina Oprea, Amir Houmansadr
摘要
While poisoning attacks on machine learning models have been extensively studied, the mechanisms by which adversaries can distribute poisoned models at scale remain largely unexplored. In this paper, we shed light on how model leaderboards -- ranked platforms for model discovery and evaluation -- can serve as a powerful channel for adversaries for stealthy large-scale distribution of poisoned models. We present TrojanClimb, a general framework that enables injection of malicious behaviors while maintaining competitive leaderboard performance. We demonstrate its effectiveness across four diverse modalities: text-embedding, text-generation, text-to-speech and text-to-image, showing that adversaries can successfully achieve high leaderboard rankings while embedding arbitrary harmful functionalities, from backdoors to bias injection. Our findings reveal a significant vulnerability in the machine learning ecosystem, highlighting the urgent need to redesign leaderboard evaluation mechanisms to detect and filter malicious (e.g., poisoned) models, while exposing broader security implications for the machine learning community regarding the risks of adopting models from unverified sources.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Your Compiler is Backdooring Your Model: Understanding and Exploiting Compilation Inconsistency Vulnerabilities in Deep Learning CompilersSimin Chen, Jinjun Peng, Yixin He, Junfeng Yang 等S&P 2026 · 被引用 11 次
- When Anonymity Breaks: Identifying Models Behind Text-to-Image LeaderboardsAli Naseh, Anshuman Suri, Yuefeng Peng, Harsh Chaudhari 等CVPR 2026
它引用的顶会 Paper34
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos 等ICML 2024 · 被引用 1,212 次
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz 等ICML 2023 · 被引用 854 次
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun 等ICLR 2024 · 被引用 716 次
- Why Do Adversarial Attacks Transfer? Explaining Transferability of Evasion and Poisoning AttacksAmbra Demontis, Marco Melis, Maura Pintor, Matthew Jagielski 等USENIX Security 2019 · 被引用 466 次
相关 Paper
- Imperceptible Content Poisoning in LLM-Powered ApplicationsQuan Zhang, Chijin Zhou, Gwihwan Go, Binqi Zeng 等ASE 2024 · 被引用 3 次
- Customization under Fire: Plugin Poisoning in Text-to-Image EcosystemJiahao Chen, Xing He, Yong Yang, Xinfeng Li 等CCS 2026 · 被引用 2 次
- When LoRA Betrays: Backdooring Text-to-Image Models by Masquerading as Benign AdaptersLiangwei Lyu, Jiaqi Xu, Jianwei Ding, Qiyao DengCVPR 2026 · 被引用 5 次
- Exploring and Mitigating Adversarial Manipulation of Voting-Based LeaderboardsYangsibo Huang, Milad Nasr, Anastasios Nikolas Angelopoulos, Nicholas Carlini 等ICML 2025
- Are Your LLM-based Text-to-SQL Models Secure? Exploring SQL Injection via Backdoor AttacksMeiyu Lin, Haichuan Zhang, Jiale Lao, Renyuan Li 等SIGMOD 2026 · 被引用 4 次
