NExT-Guard: Training-Free Streaming Safeguard without Token-Level Labels
Junfeng Fang, Nachuan Chen, Houcheng Jiang, Dan Zhang, Xiangnan He, Tat-Seng Chua, Xiang Wang
Abstract
Large language models are increasingly deployed in streaming scenarios, rendering conventional post-hoc safeguards ineffective as they fail to interdict unsafe content in real-time. While streaming safeguards based on token-level supervised training could address this, they necessitate expensive annotations and suffer from severe overfitting. In this work, we challenge the paradigm that streaming safety must rely on token-level supervised training. Instead, it is an inherent capability of well-trained post-hoc safeguards, as they already encode token-level risk signals in hidden representations. Hence, we introduce NEXT-GUARD, a training-free framework that achieves streaming safeguards by monitoring interpretable latent features from Sparse Autoencoders (SAEs). It uses pretrained SAEs from publicly available base LLMs, enabling flexible, low-cost deployment without token-level supervision. Experimental results show that NEXT-GUARD outperforms both post-hoc and streaming safeguards based on supervised training, with superior robustness across models, SAE variants, and risk scenarios. These results make NEXT-GUARD a universal and scalable paradigm for real-time safety, accelerating the practical deployment of streaming safeguards. Code is available at https://github. com/NashChennc/NExTGuard .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1455ca7c-3ac7-4f08-8e37-a6f0237dd65cCited by top-tier papers1
Ask how each one uses itBuilds on6
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic ToolsJunde Wu, Jiayuan Zhu, Yuyuan Liu, Min Xu et al.ACL 2025 · 88 citations
- From Judgment to Interference: Early Stopping LLM Harmful Outputs via Streaming Content MonitoringYang Li, Qiang Sheng, Yehan Yang, Xueyao Zhang et al.NeurIPS 2025 · 21 citations
- Scaling and evaluating sparse autoencodersLeo Gao, Tom Dupré la Tour, Henk Tillman, Gabriel Goh et al.ICLR 2025 · 10 citations
- Self-Regularization with Sparse Autoencoders for Controllable LLM-based ClassificationXuansheng Wu, Wenhao Yu, Xiaoming Zhai, Ninghao LiuKDD 2025
Related papers
- PlugGuard: A Streaming Safeguard for Large Models via Latent Dynamics-Guided Risk DetectionXiaodan Li, Mengjie Wu, Yao Zhu, Yunna Lv et al.ICML 2026 · 5 citations
- C-SafeGen: Certified Safe LLM Generation with Claim-Based Streaming GuardrailsMintong Kang, Zhaorun Chen, Bo LiNeurIPS 2025 · 4 citations
- SCOPE: Streaming Covariance-Orthogonal Post-Hoc Editing for Continual LLM Safety GovernanceYizhe Yang, Xuanming Jiang, Jisheng Dang, Aoying Wang et al.KDD 2026 · 1 citation
- Gamma-Guard: Lightweight Residual Adapters for Robust Guardrails in Large Language ModelsLijia Lv, Yuanshu Zhao, Guan Wang, Xuehai Tang et al.EMNLP 2025
- LLM Safety From Within: Detecting Harmful Content with Internal RepresentationsDifan Jiao, Yilun Liu, Ye Yuan, Zhenwei Tang et al.ACL 2026 · 3 citations
