DualSentinel: A Lightweight Framework for Detecting Targeted Attacks in Black-box LLM via Dual Entropy Lull Pattern
Xiaoyi Pang, Xuanyi Hao, Pengyu Liu, Qi Luo, Song Guo, Zhibo Wang
摘要
Recent intelligent systems integrate powerful Large Language Models (LLMs) through APIs, but their trustworthiness may be critically undermined by sequence-coercing targeted attacks that force the model to emit an attacker-chosen payload, including malicious URLs, system commands, and specific misinformation. Existing defensive approaches for such threats typically rely on high access rights, impose prohibitive costs, and hinder normal inference, rendering them impractical for real-world scenarios. To solve these limitations, we introduce DualSentinel, a lightweight defense framework that can accurately and promptly detect the activation of targeted attacks alongside the LLM generation process. We first identify a characteristic of compromised LLMs, termed Entropy Lull: when a sequence-coercing targeted attack successfully hijacks the generation process, the LLM exhibits a distinct period of abnormally low and stable token probability entropy, indicating it is following a fixed path rather than making creative choices. DualSentinel leverages this pattern by developing an innovative dual-check approach. It first employs a magnitude and trend-aware monitoring method to proactively and sensitively flag an entropy lull pattern at runtime. Upon such flagging, it triggers a lightweight yet powerful secondary verification based on task-flipping. An attack is confirmed only if the entropy lull pattern persists across both the original and the flipped task, proving that the LLM's output is coercively controlled. Extensive evaluations show that Du-alSentinel is both highly effective (superior detection accuracy with the lowest false positives) and remarkably efficient (negligible additional cost), offering a truly practical path toward securing deployed LLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper26
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Adversarial Neuron Pruning Purifies Backdoored Deep ModelsDongxian Wu, Yisen WangNeurIPS 2021 · 被引用 441 次
- Poisoning Language Models During Instruction TuningAlexander Wan, Eric Wallace, Sheng Shen, Dan KleinICML 2023 · 被引用 319 次
- Weight Poisoning Attacks on Pretrained ModelsKeita Kurita, Paul Michel, Graham NeubigACL 2020 · 被引用 312 次
- Formalizing and Benchmarking Prompt Injection Attacks and DefensesYupei Liu, Yuqi Jia, Runpeng Geng, Jinyuan Jia 等USENIX Security 2024 · 被引用 308 次
相关 Paper
- ConfGuard: A Simple and Effective Backdoor Detection for Large Language ModelsZihan Wang, Rui Zhang, Hongwei Li, Wenshu Fan 等AAAI 2026 · 被引用 5 次
- Quantifying Large Language Model Attacks Through the Lens of Model CognitionXiuming Liu, Chaoxiang He, Xuanran Yu, Jichen Chai 等USENIX Security 2026
- Defending Against Social Engineering Attacks in the Age of LLMsLin Ai, Tharindu Kumarage, Amrita Bhattacharjee, Zizhou Liu 等EMNLP 2024 · 被引用 12 次
- RepGuard: Adaptive Feature Decoupling for Robust Backdoor Defense in Large Language ModelsChenxu Niu, Jie M. Zhang, Yanbing Liu, Yunpeng Li 等NeurIPS 2025 · 被引用 1 次
- LoopLLM: Transferable Energy-Latency Attacks in LLMs via Repetitive GenerationXingyu Li, Xiaolei Liu, Cheng Liu, Yixiao Xu 等AAAI 2026 · 被引用 5 次
