Function Calling as a Flexible LLM Defense Add-On: Capability and Application Exploration
Zhenlan Ji, Daoyuan Wu, Wenxuan Wang, Pingchuan Ma, Shuai Wang, Lei Ma, Juergen Rahmel
摘要
Large language models (LLMs) exhibit impressive capabilities but are susceptible to adversarial attacks that induce harmful outputs. Although various defenses have been proposed, their practicality is restricted by substantial runtime overhead or degraded model helpfulness. Moreover, LLM applications typically have diverse and evolving security requirements that cannot be fully anticipated during the design of static defenses. These limitations call for a flexible, low-overhead defense mechanism that can be easily customized to meet task-specific needs. In this paper, we explore function calling (FC)—a built-in mechanism in modern LLMs for invoking custom tools—as a lightweight and adaptable defense add-on. We show that by defining functions representing malicious actions, LLMs equipped with FC can intercept harmful prompts by triggering these function calls instead of generating unsafe content. Extensive experiments across mainstream LLMs demonstrate that FC substantially improves defense effectiveness with minimal impact on model helpfulness. To further assess FC's practical utility, we also introduce DSPEC, a new dataset reflecting real-world LLM applications with specific defense requirements. Our evaluations on DSPEC show that FC substantially outperforms existing defenses in this realistic setting. Besides, we also explore the practical applications of FC in various scenarios, including universal defense frameworks and multi-agent systems, further demonstrating its versatility and effectiveness in enhancing LLM security.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Disentangling Adversarial Prompts: A Semantic-Graph Defense for Robust LLM SecurityXiang Fang, Wanlong FangAAAI 2026 · 被引用 4 次
- Gandalf the Red: Adaptive Security for LLMsNiklas Pfister, Václav Volhejn, Manuel Knott, Santiago Arias 等ICML 2025
- Root Defense Strategies: Ensuring Safety of LLM at the Decoding LevelXinyi Zeng, Yuying Shang, Jiawei Chen, Jingyuan Zhang 等ACL 2025 · 被引用 6 次
- MetaDefense: Defending Fine-tuning based Jailbreak Attack Before and During GenerationWeisen Jiang, Sinno Jialin PanNeurIPS 2025 · 被引用 10 次
- Defenses Against Prompt Attacks Learn Surface HeuristicsShawn Li, Chenxiao Yu, Zhiyu Ni, Hao Li 等ACL 2026 · 被引用 8 次
