PlugGuard: A Streaming Safeguard for Large Models via Latent Dynamics-Guided Risk Detection
Xiaodan Li, Mengjie Wu, Yao Zhu, Yunna Lv, YueFeng Chen, Cen Chen, Jianmei Guo, Hui Xue'
摘要
Large models (LMs) are powerful content generators, yet their open‑ended nature can also introduce potential risks, such as generating harmful or biased content. Existing guardrails mostly perform post-hoc detection that may expose unsafe content before it is caught, and the latency constraints further push them toward lightweight models, limiting detection accuracy. In this work, we propose PlugGuard, a novel plug-in framework that enables streaming risk detection within the LM generation pipeline. PlugGuard leverages intermediate LM hidden states through a Streaming Latent Dynamics Head (SLD), which models the temporal evolution of risk across the generated sequence for more accurate real-time risk detection. To achieve reliable streaming moderation in real applications, we introduce an Anchored Temporal Consistency (ATC) loss, ensuring that risk assessments remain consistent with a strict stop-if-harmful policy. Besides, for a rigorous evaluation of streaming guardrails, we also present StreamGuardBench—a model-grounded benchmark featuring on-the-fly responses from each protected model, reflecting real-world streaming scenarios in both text and vision–language tasks. Across diverse models and datasets, PlugGuard consistently outperforms state-of-the-art streaming guardrails (achieving a 22.80% F1 score gain), while using only 20M parameters and adding less than 0.5 ms of per-token latency. The code and StreamGuardBench are released at PlugGuard to facilitate research on streaming guardrails.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
- FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual PromptsYichen Gong, Delong Ran, Jinyuan Liu, Conglei Wang 等AAAI 2025 · 被引用 350 次
- Adversarial Attacks against Closed-Source MLLMs via Feature Optimal AlignmentXiaojun Jia, Sensen Gao, Simeng Qin, Tianyu Pang 等NeurIPS 2025 · 被引用 49 次
- Res-Tuning: A Flexible and Efficient Tuning Paradigm via Unbinding Tuner from BackboneZeyinzi Jiang, Chaojie Mao, Ziyuan Huang, Ao Ma 等NeurIPS 2023 · 被引用 31 次
- From Judgment to Interference: Early Stopping LLM Harmful Outputs via Streaming Content MonitoringYang Li, Qiang Sheng, Yehan Yang, Xueyao Zhang 等NeurIPS 2025 · 被引用 21 次
相关 Paper
- NExT-Guard: Training-Free Streaming Safeguard without Token-Level LabelsJunfeng Fang, Nachuan Chen, Houcheng Jiang, Dan Zhang 等ICML 2026 · 被引用 4 次
- FlexGuard: Continuous Risk Scoring for Strictness-Adaptive LLM Content ModerationZhihao Ding, Jinming Li, Ze Lu, Jieming ShiACL 2026 · 被引用 2 次
- C-SafeGen: Certified Safe LLM Generation with Claim-Based Streaming GuardrailsMintong Kang, Zhaorun Chen, Bo LiNeurIPS 2025 · 被引用 4 次
- RST-Guarder: Enhancing Long-Context Robustness for Safeguards via RST Parsing and Probabilistic InferenceXu Zhang, Xiaojun WanACL 2026
- LLM Safety From Within: Detecting Harmful Content with Internal RepresentationsDifan Jiao, Yilun Liu, Ye Yuan, Zhenwei Tang 等ACL 2026 · 被引用 3 次
