Beyond Static Benchmarks: Synthesizing Harmful Content via Persona-based Simulation for Robust Evaluation
Huije Lee, Jisu Shin, Hoyun Song, Changgeon Ko, Jong C. Park
摘要
Static benchmarks for harmful content detection face limitations in scalability and diversity, and may also be affected by contamination from web-scale pre-training corpora. To address these issues, we propose a framework for synthesizing harmful content, leveraging persona-guided large language model (LLM) agents. Our approach constructs twodimensional user personas by integrating demographic identities and topical interests with situational harmful strategies, enabling the simulation of diverse and contextually grounded harmful interactions. We evaluate the framework along three dimensions: harmfulness, challenge level, and diversity. Both human and LLM-based evaluations confirm that our framework achieves a high harmful generation success rate. Experiments across multiple detection systems reveal that our synthetic scenarios are more challenging to detect than those in existing benchmarks. Furthermore, a multifaceted analysis confirms that our approach achieves linguistic and topical diversity comparable to human-curated datasets, establishing our framework as an effective tool for robust stress-testing of harmful content detection systems 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris 等UIST 2023 · 被引用 1,882 次
- Social Simulacra: Creating Populated Prototypes for Social Computing SystemsJoon Sung Park, Lindsay Popowski, Carrie J. Cai, Meredith Ringel Morris 等UIST 2022 · 被引用 192 次
- Time Travel in LLMs: Tracing Data Contamination in Large Language ModelsShahriar Golchin, Mihai SurdeanuICLR 2024 · 被引用 165 次
- Characterizing Twitter Users Who Engage in Adversarial Interactions against Political CandidatesYiqing Hua, Mor Naaman, Thomas RistenpartCHI 2020 · 被引用 33 次
- On the Challenges of Using Black-Box APIs for Toxicity Evaluation in ResearchLuiza Pozzobon, Beyza Ermis, Patrick Lewis, Sara HookerEMNLP 2023 · 被引用 19 次
相关 Paper
- Can AI-Generated Persuasion Be Detected? Persuaficial Benchmark and AI vs. Human Linguistic DifferencesArkadiusz Modzelewski, Pawel Golik, Anna Kolos, Giovanni Da San MartinoACL 2026 · 被引用 1 次
- HEV Generative Sandbox: A Framework for Assessing Domain-Specific Social Risks Through Human-LLM SimulationYiran Liu, Zhiyi Hou, Xiaoang Xu, Shuo Wang 等AAAI 2026
- Detoxifying Large Language Models via the Diversity of Toxic SamplesYing Zhao, Yuanzhao Guo, Xuemeng Weng, Yuan Tian 等EMNLP 2025
- GRAID: Synthetic Data Generation with Geometric Constraints and Multi-Agentic Reflection for Harmful Content DetectionMelissa Kazemi Rad, Alberto Purpura, Himanshu Kumar, Emily Chen 等EMNLP 2025
- ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMsAdi Simhi, Jonathan Herzig, Martin Tutek, Itay Itzhak 等ICLR 2026 · 被引用 3 次
