Think Twice Before You Act: Enhancing Agent Behavioral Safety with Thought Correction
Changyue Jiang, Wenqi Zhang, Xudong Pan, Geng Hong, Min Yang
摘要
LLM-based agents solve complex tasks through iterative reasoning, tool use, and environment interaction, where each intermediate thought directly shapes subsequent actions. Small deviations in these thoughts can therefore propagate into unsafe behaviors, yet existing guardrails typically operate only on final outputs or require intrusive model modifications. We introduce Thought-Aligner, a lightweight plug-in safety model that performs causal correction on unsafe thoughts before action execution, without altering the underlying agent. The corrected thoughts are fed back into the agent, steering its decision process and tool use toward safer trajectories. Because it operates solely at the thought level, Thought-Aligner is model-agnostic and can be integrated into diverse agent frameworks. We train Thought-Aligner via two-stage contrastive learning on paired safe and unsafe thoughts generated across ten risk scenarios. Experiments on diverse agent-safety benchmarks and six LLMs show that Thought-Aligner increases behavioral safety from about 50% without protection to around 90% on average, exceeding state-ofthe-art guardrails by roughly 23%, while also improving helpfulness by about 5%. The method incurs low per-step latency and minimal overhead, enabling scalable and practical deployment. We publicly release Thought-Aligner-7B at https: //huggingface.co/WhitzardAgent/ Thought-Aligner-7B.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Learning When to Act or Refuse: Guarding Agentic Reasoning Models for Safe Multi-Step Tool UseAradhye Agarwal, Gurdit Singh Siyan, Yash Pandya, Joykirat Singh 等ICML 2026 · 被引用 5 次
- MirrorGuard: Toward Secure Computer-Use Agents via Simulation-to-Real Reasoning CorrectionWenqi Zhang, Yulin Shen, Changyue Jiang, Jiarun Dai 等CCS 2026 · 被引用 3 次
- When "Correct" Is Not Safe: Can We Trust Functionally Correct Patches Generated by Code Agents?Yibo Peng, James Song, Lei Li, Xinyu Yang 等ACL 2026 · 被引用 2 次
- MATE: Policy-Aware Security Auditing for Mobile Agents via Synthesis-Driven Trajectory LearningChangyue Jiang, Jiayi Wang, Xin Wen, Jiarun Dai 等USENIX Security 2026
它引用的顶会 Paper24
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIsYujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu 等ICLR 2024 · 被引用 1,469 次
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou 等ICLR 2024 · 被引用 1,197 次
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou 等ICML 2024 · 被引用 1,031 次
- AgentBench: Evaluating LLMs as AgentsXiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu 等ICLR 2024 · 被引用 748 次
- AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge BasesZhaorun Chen, Zhen Xiang, Chaowei Xiao, Dawn Song 等NeurIPS 2024 · 被引用 539 次
相关 Paper
- SAFEPATH: Preventing Harmful Reasoning in Chain-of-Thought via Early AlignmentWonje Jeung, Sangyeon Yoon, Minsuk Kahng, Albert NoNeurIPS 2025 · 被引用 31 次
- Finding and Reactivating Post-Trained LLMs' Hidden Safety MechanismsMingjie Li, Wai Man Si, Michael Backes, Yang Zhang 等NeurIPS 2025 · 被引用 4 次
- SafeAgent: Safeguarding LLM Agents via an Automated Risk SimulatorXueyang Zhou, Weidong Wang, Lin Lu, Jiawen Shi 等ACL 2026 · 被引用 5 次
- AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement LearningYi Zhang, An Zhang, XiuYu Zhang, Leheng Sheng 等ICLR 2026 · 被引用 15 次
- GuardAgent: Safeguard LLM Agents via Knowledge-Enabled ReasoningZhen Xiang, Linzhi Zheng, Yanjie Li, Junyuan Hong 等ICML 2025
