SCOPE: Streaming Covariance-Orthogonal Post-Hoc Editing for Continual LLM Safety Governance
Yizhe Yang, Xuanming Jiang, Jisheng Dang, Aoying Wang, Baoyi An, Hao Wu, Bimei Wang, Hong Peng, Guoshuai Zhao, Bin Hu, Zhongyu Yang
Abstract
Deployed large language models (LLMs) require continual post-hoc safety updates to counter evolving threats, yet current methods often entangle refusal behaviors with capability-critical representations, causing cumulative utility drift in sequential edits. We propose SCOPE, a streaming post-hoc editing framework for capability-invariant LLM safety governance. SCOPE derives layer-wise capability subspaces via online activation covariance and constrains safety updates to their orthogonal spaces, decoupling capability preservation from safety-specific edits. Within the derived null space, it extracts low-rank refusal directions from contrastive harmful activations to steer refusal behaviors without interfering with core capabilities. Its online covariance maintenance enables efficient, adaptive expansion of the protected capability subspace against open-ended threat evolution. Our experiments validate that SCOPE delivers significant safety robustness gains by substantially reducing attack success rates, while keeping general capabilities within natural run-to-run variance under repeated edits, supporting long-term safety maintenance for deployed LLMs.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get d998ce53-98ad-49e4-a5ec-e1c39bc24c72Related papers
- The Geometry of Refusal in Large Language Models: Concept Cones and Representational IndependenceTom Wollschläger, Jannes Elstner, Simon Geisler, Vincent Cohen-Addad et al.ICML 2025
- SAME: Safety-Aware Model Editing Guided by Safety TransformationJiayi Wang, Shipeng Wang, Ji Wu, Jian SunACL 2026
- Safety Subspaces are Not Linearly Distinct: A Fine-Tuning Case StudyKaustubh Ponkshe, Shaan Shah, Raghav Singhal, Praneeth VepakommaICLR 2026 · 9 citations
- SCOPE: Intrinsic Semantic Space Control for Mitigating Copyright Infringement in LLMsZhenliang Zhang, Xinyu Hu, Xiaojun WanAAAI 2026 · 1 citation
- NExT-Guard: Training-Free Streaming Safeguard without Token-Level LabelsJunfeng Fang, Nachuan Chen, Houcheng Jiang, Dan Zhang et al.ICML 2026 · 4 citations
