A Causal Explainable Guardrails for Large Language Models
Zhixuan Chu, Yan Wang, Longfei Li, Zhibo Wang, Zhan Qin, Kui Ren
Abstract
Large Language Models (LLMs) have shown impressive performance in natural language tasks, but their outputs can exhibit undesirable attributes or biases. Existing methods for steering LLMs toward desired attributes often assume unbiased representations and rely solely on steering prompts. However, the representations learned from pre-training can introduce semantic biases that influence the steering process, leading to suboptimal results. We propose LLMGuardrail, a novel framework that incorporates causal analysis and adversarial learning to obtain unbiased steering representations in LLMs. LLMGuardrail systematically identifies and blocks the confounding effects of biases, enabling the extraction of unbiased steering representations. Additionally, it includes an explainable component that provides insights into the alignment between the generated output and the desired direction. Experiments demonstrate LLMGuardrail's effectiveness in steering LLMs toward desired attributes while mitigating biases. Our work contributes to the development of safe and reliable LLMs that align with desired attributes. CCS CONCEPTS • Computing methodologies → Natural language generation; Causal reasoning and diagnostics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4b786ce0-b523-44d5-8978-43a62a57a3eeCited by top-tier papers6
- Large Language Model Unlearning via Embedding-Corrupted PromptsChris Yuhao Liu, Yaxuan Wang, Jeffrey Flanigan, Yang LiuNeurIPS 2024 · 138 citations
- Sok: Evaluating Jailbreak Guardrails for Large Language ModelsXunguang Wang, Zhenlan Ji, Wenxuan Wang, Zongjie Li et al.S&P 2026 · 27 citations
- Understanding and Mitigating Overrefusal in LLMs from an Unveiling Perspective of Safety Decision BoundaryLicheng Pan, Yongqi Tong, Xin Zhang, Xiaolu Zhang et al.EMNLP 2025 · 3 citations
- Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM SafetySeongmin Lee, Aeree Cho, Grace C. Kim, Shengyun Peng et al.EMNLP 2025 · 1 citation
- Explainable Token-level Noise Filtering for LLM Fine-tuning DatasetsYuchen Yang, Wenze Lin, Enhao Huang, Zhixuan Chu et al.ICLR 2026 · 1 citation
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
Related papers
- A Causal Perspective for Enhancing Jailbreak Attack and DefenseLicheng Pan, Yunsheng Lu, Jiexi Liu, Jialing Tao et al.NDSS 2026
- Causal-Guided Active Learning for Debiasing Large Language ModelsZhouhao Sun, Li Du, Xiao Ding, Yixuan Ma et al.ACL 2024
- A Lightweight Explainable Guardrail for Prompt SafetyMd. Asiful Islam, Mihai SurdeanuACL 2026
- MrGuard: A Multilingual Reasoning Guardrail for Universal LLM SafetyYahan Yang, Soham Dan, Shuo Li, Dan Roth et al.EMNLP 2025
- ExpGuard: LLM Content Moderation in Specialized DomainsMinseok Choi, Dongjin Kim, Seungbin Yang, Subin Kim et al.ICLR 2026 · 3 citations
