PlugGuard: A Streaming Safeguard for Large Models via Latent Dynamics-Guided Risk Detection
Xiaodan Li, Mengjie Wu, Yao Zhu, Yunna Lv, YueFeng Chen, Cen Chen, Jianmei Guo, Hui Xue'
Abstract
Large models (LMs) are powerful content generators, yet their open‑ended nature can also introduce potential risks, such as generating harmful or biased content. Existing guardrails mostly perform post-hoc detection that may expose unsafe content before it is caught, and the latency constraints further push them toward lightweight models, limiting detection accuracy. In this work, we propose PlugGuard, a novel plug-in framework that enables streaming risk detection within the LM generation pipeline. PlugGuard leverages intermediate LM hidden states through a Streaming Latent Dynamics Head (SLD), which models the temporal evolution of risk across the generated sequence for more accurate real-time risk detection. To achieve reliable streaming moderation in real applications, we introduce an Anchored Temporal Consistency (ATC) loss, ensuring that risk assessments remain consistent with a strict stop-if-harmful policy. Besides, for a rigorous evaluation of streaming guardrails, we also present StreamGuardBench—a model-grounded benchmark featuring on-the-fly responses from each protected model, reflecting real-world streaming scenarios in both text and vision–language tasks. Across diverse models and datasets, PlugGuard consistently outperforms state-of-the-art streaming guardrails (achieving a 22.80% F1 score gain), while using only 20M parameters and adding less than 0.5 ms of per-token latency. The code and StreamGuardBench are released at PlugGuard to facilitate research on streaming guardrails.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on8
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual PromptsYichen Gong, Delong Ran, Jinyuan Liu, Conglei Wang et al.AAAI 2025 · 350 citations
- Adversarial Attacks against Closed-Source MLLMs via Feature Optimal AlignmentXiaojun Jia, Sensen Gao, Simeng Qin, Tianyu Pang et al.NeurIPS 2025 · 49 citations
- Res-Tuning: A Flexible and Efficient Tuning Paradigm via Unbinding Tuner from BackboneZeyinzi Jiang, Chaojie Mao, Ziyuan Huang, Ao Ma et al.NeurIPS 2023 · 31 citations
- From Judgment to Interference: Early Stopping LLM Harmful Outputs via Streaming Content MonitoringYang Li, Qiang Sheng, Yehan Yang, Xueyao Zhang et al.NeurIPS 2025 · 21 citations
Related papers
- NExT-Guard: Training-Free Streaming Safeguard without Token-Level LabelsJunfeng Fang, Nachuan Chen, Houcheng Jiang, Dan Zhang et al.ICML 2026 · 4 citations
- FlexGuard: Continuous Risk Scoring for Strictness-Adaptive LLM Content ModerationZhihao Ding, Jinming Li, Ze Lu, Jieming ShiACL 2026 · 2 citations
- C-SafeGen: Certified Safe LLM Generation with Claim-Based Streaming GuardrailsMintong Kang, Zhaorun Chen, Bo LiNeurIPS 2025 · 4 citations
- RST-Guarder: Enhancing Long-Context Robustness for Safeguards via RST Parsing and Probabilistic InferenceXu Zhang, Xiaojun WanACL 2026
- LLM Safety From Within: Detecting Harmful Content with Internal RepresentationsDifan Jiao, Yilun Liu, Ye Yuan, Zhenwei Tang et al.ACL 2026 · 3 citations
