Improving Generalizability in Implicitly Abusive Language Detection with Concept Activation Vectors
Isar Nejadgholi, Kathleen C. Fraser, Svetlana Kiritchenko
Abstract
Robustness of machine learning models on ever-changing real-world data is critical, especially for applications affecting human well-being such as content moderation. New kinds of abusive language continually emerge in online discussions in response to current events (e.g., COVID-19), and the deployed abuse detection systems should be updated regularly to remain accurate. In this paper, we show that general abusive language classifiers tend to be fairly reliable in detecting out-of-domain explicitly abusive utterances but fail to detect new types of more subtle, implicit abuse. Next, we propose an interpretability technique, based on the Testing Concept Activation Vector (TCAV) method from computer vision, to quantify the sensitivity of a trained model to the human-defined concepts of explicit and implicit abusive language, and use that to explain the generalizability of the model on new data, in this case, COVID-related anti-Asian hate speech. Extending this technique, we introduce a novel metric, Degree of Explicitness, for a single instance and show that the new metric is beneficial in suggesting out-of-domain unlabeled examples to effectively enrich the training data with informative, implicitly abusive texts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7d453805-ba2e-41ff-be4e-2ea02a5ba236Cited by top-tier papers1
Ask how each one uses itBuilds on2
Related papers
- A Community-Centric Perspective for Characterizing and Detecting Anti-Asian Violence-Provoking SpeechGaurav Verma, Rynaa Grover, Jiawei Zhou, Binny Mathew et al.ACL 2024
- GCAV: A Global Concept Activation Vector Framework for Cross-Layer Consistency in InterpretabilityZhenghao He, Sanchit Sinha, Guangzhi Xiong, Aidong ZhangICCV 2025 · 2 citations
- Concept Distillation: Leveraging Human-Centered Explanations for Model ImprovementAvani Gupta, Saurabh Saini, P. J. NarayananNeurIPS 2023 · 18 citations
- Uncovering Safety Risks of Large Language Models through Concept Activation VectorZhihao Xu, Ruixuan Huang, Changyu Chen, Xiting WangNeurIPS 2024 · 83 citations
- ConvAbuse: Data, Analysis, and Benchmarks for Nuanced Detection in Conversational AIAmanda Cercas Curry, Gavin Abercrombie, Verena RieserEMNLP 2021 · 38 citations
