Learning Efficient Guardrails for Compliance
Xiaofei Wen, Wenjie Mo, Yanan Xie, Peng Qi, Muhao Chen
摘要
Autonomous web agents are increasingly deployed for long-horizon tasks, yet their ability to adhere to real-world policies remains critically underexplored compared to standard safety objectives. To address this gap, we introduce PolicyGuardBench, a benchmark of 60k policy-trajectory pairs designed to evaluate compliance through both full-trajectory and novel prefix-based violation detection tasks. Using this dataset, we train PolicyGuard, a lightweight guardrail model that achieves strong detection accuracy while maintaining high inference efficiency. Notably, our model demonstrates robust generalization capabilities, preserving high performance even on unseen domains. These contributions establish a comprehensive framework for studying policy compliance, showing that accurate and generalizable guardrails are feasible at small scales.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou 等ICLR 2024 · 被引用 1,197 次
- Language Agent Tree Search Unifies Reasoning, Acting, and Planning in Language ModelsAndy Zhou, Kai Yan, Michal Shlapentokh-Rothman, Haohan Wang 等ICML 2024 · 被引用 443 次
- ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web AgentsIdo Levy, Ben wiesel, Sami Marreed, Alon Oved 等ICLR 2026 · 被引用 78 次
相关 Paper
- ShieldAgent: Shielding Agents via Verifiable Safety Policy ReasoningZhaorun Chen, Mintong Kang, Bo LiICML 2025
- GuardAgent: Safeguard LLM Agents via Knowledge-Enabled ReasoningZhen Xiang, Linzhi Zheng, Yanjie Li, Junyuan Hong 等ICML 2025
- GuardBench: A Large-Scale Benchmark for Guardrail ModelsElias Bassani, Ignacio SanchezEMNLP 2024 · 被引用 7 次
- Building a Foundational Guardrail for General Agentic Systems via Synthetic DataYue Huang, Hang Hua, Yujun Zhou, Pengcheng Jing 等ICLR 2026 · 被引用 29 次
- DynaGuard: A Dynamic Guardian Model With User-Defined PoliciesMonte Hoover, Vatsal Baherwani, Neel Jain, Khalid Saifullah 等ICLR 2026 · 被引用 18 次
