GRAID: Synthetic Data Generation with Geometric Constraints and Multi-Agentic Reflection for Harmful Content Detection
Melissa Kazemi Rad, Alberto Purpura, Himanshu Kumar, Emily Chen, Mohammad Shahed Sorower
摘要
We address the problem of data scarcity in harmful text classification for guardrailing applications and introduce GRAID (Geometric and Reflective AI-Driven Data Augmentation), a novel pipeline that leverages Large Language Models (LLMs) for dataset augmentation. GRAID consists of two stages: (i) generation of geometrically controlled examples using a constrained LLM, and (ii) augmentation through a multi-agentic reflective process that promotes stylistic diversity and uncovers edge cases. This combination enables both reliable coverage of the input space and nuanced exploration of harmful content. Using two benchmark data sets, we demonstrate that augmenting a harmful text classification dataset with GRAID leads to significant improvements in downstream guardrail model performance. Warning: This paper contains techniques to synthetically generate offensive and malicious content using LLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Tree of Attacks: Jailbreaking Black-Box LLMs AutomaticallyAnay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson 等NeurIPS 2024 · 被引用 835 次
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 被引用 722 次
- Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and InferenceBenjamin Warner, Antoine Chaffin, Benjamin Clavié, Orion Weller 等ACL 2025 · 被引用 552 次
相关 Paper
- ExpGuard: LLM Content Moderation in Specialized DomainsMinseok Choi, Dongjin Kim, Seungbin Yang, Subin Kim 等ICLR 2026 · 被引用 3 次
- HarmAug: Effective Data Augmentation for Knowledge Distillation of Safety Guard ModelsSeanie Lee, Haebin Seong, Dong Bok Lee, Minki Kang 等ICLR 2025
- MrGuard: A Multilingual Reasoning Guardrail for Universal LLM SafetyYahan Yang, Soham Dan, Shuo Li, Dan Roth 等EMNLP 2025
- RigorLLM: Resilient Guardrails for Large Language Models against Undesired ContentZhuowen Yuan, Zidi Xiong, Yi Zeng, Ning Yu 等ICML 2024 · 被引用 77 次
- RST-Guarder: Enhancing Long-Context Robustness for Safeguards via RST Parsing and Probabilistic InferenceXu Zhang, Xiaojun WanACL 2026
