DynaGuard: A Dynamic Guardian Model With User-Defined Policies
Monte Hoover, Vatsal Baherwani, Neel Jain, Khalid Saifullah, Joseph James Vincent, Chirag Jain, Melissa Kazemi Rad, C. Bayan Bruss, Ashwinee Panda, Tom Goldstein
摘要
Guardian models play a crucial role in ensuring the safety and ethical behavior of user-facing AI applications by enforcing guardrails and detecting harmful content. While standard guardian models are limited to predefined, static harm categories, we introduce DynaGuard, a suite of dynamic guardian models offering novel flexibility by evaluating text based on user-defined policies, and DynaBench, a dataset for training and evaluating dynamic guardian models. Our models provide both rapid detection of policy violations and a chain-of-thought reasoning option that articulate and justify model outputs. Critically, DynaGuard not only surpasses static models in detection accuracy on traditional safety categories, but is competitive with frontier reasoning models on free-form policy violations, all in a fraction of the time. This makes DynaGuard an critical tool for language model guardrails.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper6
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou 等ICML 2024 · 被引用 1,031 次
- Safe RLHF: Safe Reinforcement Learning from Human FeedbackJosef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji 等ICLR 2024 · 被引用 656 次
- RigorLLM: Resilient Guardrails for Large Language Models against Undesired ContentZhuowen Yuan, Zidi Xiong, Yi Zeng, Ning Yu 等ICML 2024 · 被引用 77 次
- HybridFlow: A Flexible and Efficient RLHF FrameworkGuangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu 等EuroSys 2025 · 被引用 61 次
相关 Paper
- RST-Guarder: Enhancing Long-Context Robustness for Safeguards via RST Parsing and Probabilistic InferenceXu Zhang, Xiaojun WanACL 2026
- FlexGuard: Continuous Risk Scoring for Strictness-Adaptive LLM Content ModerationZhihao Ding, Jinming Li, Ze Lu, Jieming ShiACL 2026 · 被引用 2 次
- Towards Policy-Adaptive Image Guardrail: Benchmark and MethodCaiyong Piao, Zhiyuan Yan, Haoming Xu, Yunzhen Zhao 等CVPR 2026 · 被引用 7 次
- ExpGuard: LLM Content Moderation in Specialized DomainsMinseok Choi, Dongjin Kim, Seungbin Yang, Subin Kim 等ICLR 2026 · 被引用 3 次
- EMBGuard: Constructing Hazard-Aware Guardrails for Safe Planning in Embodied AgentsDongwook Choi, Taeyoon Kwon, Bogyung Jeong, Minju Kim 等ICML 2026
