On the Challenges of Using Black-Box APIs for Toxicity Evaluation in Research
Luiza Pozzobon, Beyza Ermis, Patrick Lewis, Sara Hooker
摘要
Perception of toxicity evolves over time and often differs between geographies and cultural backgrounds. Similarly, black-box commercially available APIs for detecting toxicity, such as the Perspective API, are not static, but frequently retrained to address any unattended weaknesses and biases. We evaluate the implications of these changes on the reproducibility of findings that compare the relative merits of models and methods that aim to curb toxicity. Our findings suggest that research that relied on inherited automatic toxicity scores to compare models and techniques may have resulted in inaccurate findings. Rescoring all models from HELM, a widely respected living benchmark, for toxicity with the recent version of the API led to a different ranking of widely used foundation models. We suggest caution in applying apples-to-apples comparisons between studies and lay recommendations for a more structured approach to evaluating toxicity over time. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Prometheus: Inducing Fine-Grained Evaluation Capability in Language ModelsSeungone Kim, Jamin Shin, Yejin Choi, Joel Jang 等ICLR 2024 · 被引用 468 次
- OctoPack: Instruction Tuning Code Large Language ModelsNiklas Muennighoff, Qian Liu, Armel Randy Zebaze, Qinkai Zheng 等ICLR 2024 · 被引用 203 次
- Elo Uncovered: Robustness and Best Practices in Language Model EvaluationMeriem Boubdir, Edward Kim, Beyza Ermis, Sara Hooker 等NeurIPS 2024 · 被引用 94 次
- HateDay: Insights from a Global Hate Speech Dataset Representative of a Day on TwitterManuel Tonneau, Diyi Liu, Niyati Malhotra, Scott A. Hale 等ACL 2025 · 被引用 12 次
- Enhancing Reinforcement Learning with Dense Rewards from Language Model CriticMeng Cao, Lei Shu, Lei Yu, Yun Zhu 等EMNLP 2024 · 被引用 7 次
它引用的顶会 Paper9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Plug and Play Language Models: A Simple Approach to Controlled Text GenerationSumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung 等ICLR 2020 · 被引用 1,166 次
- The Psychological Well-Being of Content Moderators: The Emotional Labor of Commercial Moderation and Avenues for Improving SupportMiriah Steiger, Timir J. Bharucha, Sukrit Venkatagiri, Martin J. Riedl 等CHI 2021 · 被引用 168 次
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo 等ACL 2020 · 被引用 93 次
- Exploring the Limits of Domain-Adaptive Training for Detoxifying Large-Scale Language ModelsBoxin Wang, Wei Ping, Chaowei Xiao, Peng Xu 等NeurIPS 2022 · 被引用 89 次
相关 Paper
- T2ISafety: Benchmark for Assessing Fairness, Toxicity, and Privacy in Image GenerationLijun Li, Zhelun Shi, Xuhao Hu, Bowen Dong 等CVPR 2025
- Predicting the Performance of Black-box Language Models with Follow-up QueriesDylan Sam, Marc Finzi, Zico KolterNeurIPS 2025 · 被引用 10 次
- ModelCitizens: Representing Community Voices in Online SafetyAshima Suvarna, Christina Chance, Karolina Naranjo, Hamid Palangi 等EMNLP 2025
- Toxicity Detection is NOT all you Need: Measuring the Gaps to Supporting Volunteer Content Moderators through a User-Centric MethodYang Trista Cao, Lovely-Frances Domingo, Sarah A. Gilbert, Michelle L. Mazurek 等EMNLP 2024 · 被引用 4 次
- ELITE: Enhanced Language-Image Toxicity Evaluation for SafetyWonjun Lee, Doehyeon Lee, Eugene Choi, Sangyoon Yu 等ICML 2025
