KOLD: Korean Offensive Language Dataset
Younghoon Jeong, Juhyun Oh, Jongwon Lee, Jaimeen Ahn, Jihyung Moon, Sungjoon Park, Alice Oh
摘要
Warning: this paper contains content that may be offensive or upsetting. Recent directions for offensive language detection are hierarchical modeling, identifying the type and the target of offensive language, and interpretability with offensive span annotation and prediction. These improvements are focused on English and do not transfer well to other languages because of cultural and linguistic differences. In this paper, we present the Korean Offensive Language Dataset (KOLD) comprising 40,429 comments, which are annotated hierarchically with the type and the target of offensive language, accompanied by annotations of the corresponding text spans. We collect the comments from NAVER news and YouTube platform and provide the titles of the articles and videos as the context information for the annotation process. We use these annotated comments as training data for Korean BERT and RoBERTa models and find that they are effective at offensiveness detection, target classification, and target span detection while having room for improvement for target group classification and offensive span detection. We discover that the target group distribution differs drastically from the existing English datasets, and observe that providing the context information improves the model performance in offensiveness detection (+0.3), target classification (+1.5), and target group classification (+13.1). We publicly release the dataset and baseline models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- A Comprehensive Framework to Operationalize Social Stereotypes for Responsible AI EvaluationsAida Mostafazadeh Davani, Sunipa Dev, Héctor Pérez-Urbina, Vinodkumar PrabhakaranEMNLP 2025 · 被引用 6 次
- PREDICT: Multi-Agent-based Debate Simulation for Generalized Hate Speech DetectionSomeen Park, Jaehoon Kim, Seungwan Jin, Sohyun Park 等EMNLP 2024 · 被引用 5 次
- Using Off-the-Shelf Harmful Content Detection Models: Best Practices for Model ReuseAngela M. Schöpke-Gonzalez, Siqi Wu, Sagar Kumar, Libby HemphillCSCW 2025 · 被引用 4 次
- HateBuffer: Safeguarding Content Moderators' Mental Well-Being through Hate Speech Content ModificationSubin Park, Jeonghyun Kim, Jeanne Choi, Joseph Seering 等CSCW 2025 · 被引用 4 次
- SQuARe: A Large-Scale Dataset of Sensitive Questions and Acceptable Responses Created through Human-Machine CollaborationHwaran Lee, Seokhee Hong, Joonsuk Park, Takyoung Kim 等ACL 2023 · 被引用 3 次
它引用的顶会 Paper9
- Latent Hatred: A Benchmark for Understanding Implicit Hate SpeechMai ElSherief, Caleb Ziems, David Muchlinski, Vaishnavi Anupindi 等EMNLP 2021 · 被引用 159 次
- Hate-Speech and Offensive Language Detection in Roman UrduHammad Rizwan, Muhammad Haroon Shakeel, Asim KarimEMNLP 2020 · 被引用 97 次
- The State and Fate of Linguistic Diversity and Inclusion in the NLP WorldPratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali 等ACL 2020 · 被引用 40 次
- ERASER: A Benchmark to Evaluate Rationalized NLP ModelsJay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric P. Lehman 等ACL 2020 · 被引用 36 次
- From the Detection of Toxic Spans in Online Discussions to the Analysis of Toxic-to-Civil TransferJohn Pavlopoulos, Léo Laugier, Alexandros Xenos, Jeffrey Sorensen 等ACL 2022 · 被引用 35 次
相关 Paper
- COLD: A Benchmark for Chinese Offensive Language DetectionJiawen Deng, Jingyan Zhou, Hao Sun, Chujie Zheng 等EMNLP 2022 · 被引用 82 次
- K/DA: Automated Data Generation Pipeline for Detoxifying Implicitly Offensive Language in KoreanMinkyeong Jeon, Hyemin Jeong, Yerang Kim, Jiyoung Kim 等ACL 2025
- Ruddit: Norms of Offensiveness for English Reddit CommentsRishav Hada, Sohi Sudhir, Pushkar Mishra, Helen Yannakoudakis 等ACL 2021
- XDAC: XAI-Driven Detection and Attribution of LLM-Generated News Comments in KoreanWooyoung Go, Hyoungshick Kim, Alice Oh, Yongdae KimACL 2025
- Spanning the Spectrum of Hatred Detection: A Persian Multi-Label Hate Speech Dataset with Annotator RationalesZahra Delbari, Nafise Sadat Moosavi, Mohammad Taher PilehvarAAAI 2024 · 被引用 11 次
