TraceRouter: Robust Safety for Large Foundation Models via Path-Level Intervention
Chuancheng Shi, shangze li, Wenjun Lu, Wenhua Wu, Fei Shen, Cong Wang, Zifeng Cheng, Tat-Seng Chua
摘要
Despite their capabilities, large foundation models (LFMs) remain susceptible to adversarial manipulation. Current defenses predominantly rely on the ``locality hypothesis", suppressing isolated neurons or features. However, harmful semantics act as distributed, cross-layer circuits, rendering such localized interventions brittle and detrimental to utility. To bridge this gap, we propose TraceRouter, a path-level framework that traces and disconnects the causal propagation circuits of illicit semantics. TraceRouter operates in three stages: (1) it pinpoints a sensitive onset layer by analyzing attention divergence; (2) it leverages sparse autoencoders (SAEs) and differential activation analysis to disentangle and isolate malicious features; and (3) it maps these features to downstream causal pathways via feature influence scores (FIS) derived from zero-out interventions. By selectively suppressing these causal chains, TraceRouter physically severs the flow of harmful information while leaving orthogonal computation routes intact. Extensive experiments demonstrate that TraceRouter significantly outperforms state-of-the-art baselines, achieving a superior trade-off between adversarial robustness and general utility. Our code will be publicly released. WARNING: This paper contains unsafe model responses.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper23
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
相关 Paper
- Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMsYukun Jiang, Hai Huang, Mingjie Li, Yage Zhang 等ICML 2026 · 被引用 9 次
- Breaking Bad Tokens: Detoxification of LLMs Using Sparse AutoencodersAgam Goyal, Vedant Rathi, William Yeh, Yian Wang 等EMNLP 2025 · 被引用 1 次
- Language Model Circuits Are Sparse in the Neuron BasisAryaman Arora, Zhengxuan Wu, Jacob Steinhardt, Sarah SchwettmannICML 2026
- Model Unlearning via Sparse Autoencoder Subspace Guided ProjectionsXu Wang, Zihao Li, Benyou Wang, Yan Hu 等EMNLP 2025 · 被引用 9 次
- SEAL: Reinforcing Global Safety in Mixture-of-Experts through Shared Expert ALignmentQingyu Meng, Yiwei Zha, Jiahuan Pei, Koen Hindriks 等CCS 2026
