AmpleHate: Amplifying the Attention for Versatile Implicit Hate Detection
Yejin Lee, Joonghyuk Hahn, Hyeseon Ahn, Yo-Sub Han
摘要
Implicit hate speech involves subtle and indirect expressions of prejudice or hostility toward a group.Detecting it is challenging because it relies on nuanced context and implication rather than explicit offensive language.Current approaches rely on contrastive learning, which is shown to be effective on distinguishing hate and non-hate sentences.Humans, however, detect implicit hate speech by first identifying specific targets within the text and subsequently interpreting how these targets relate to their surrounding context.Motivated by this reasoning process, we propose Ample-Hate, a novel approach designed to mirror human inference for implicit hate detection.Am-pleHate identifies explicit targets using a pretrained Named Entity Recognition model and captures implicit target information via [CLS] tokens.It computes attention-based relationships between explicit, implicit targets and sentence context and then, directly injects these relational vectors into the final sentence representation.This amplifies the critical signals of target-context relations for determining implicit hate.Experiments demonstrate that Am-pleHate achieves state-of-the-art performance, outperforming contrastive learning baselines by an average of 82.14% and achieves faster convergence.Qualitative analyses further reveal that attention patterns produced by Am-pleHate closely align with human judgement, underscoring its interpretability and robustness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper7
- Latent Hatred: A Benchmark for Understanding Implicit Hate SpeechMai ElSherief, Caleb Ziems, David Muchlinski, Vaishnavi Anupindi 等EMNLP 2021 · 被引用 159 次
- Social Bias Frames: Reasoning about Social and Power Implications of LanguageMaarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky 等ACL 2020 · 被引用 16 次
- Word Embeddings Are Steers for Language ModelsChi Han, Jialiang Xu, Manling Li, Yi Fung 等ACL 2024 · 被引用 8 次
- Hate Speech Detection with Generalizable Target-aware FairnessTong Chen, Danny Wang, Xurong Liang, Marten Risius 等KDD 2024 · 被引用 7 次
- Learning from the Worst: Dynamically Generated Datasets to Improve Online Hate DetectionBertie Vidgen, Tristan Thrush, Zeerak Waseem, Douwe KielaACL 2021
相关 Paper
- ImpliHateVid: A Benchmark Dataset and Two-stage Contrastive Learning Framework for Implicit Hate Speech Detection in VideosMohammad Zia Ur Rehman, Anukriti Bhatnagar, Omkar Kabde, Shubhi Bansal 等ACL 2025 · 被引用 11 次
- CoSyn: Detecting Implicit Hate Speech in Online Conversations Using a Context Synergized Hyperbolic NetworkSreyan Ghosh, Manan Suri, Purva Chiniya, Utkarsh Tyagi 等EMNLP 2023 · 被引用 9 次
- Causality Guided Representation Learning for Cross-Style Hate Speech DetectionChengshuai Zhao, Shu Wan, Paras Sheth, Karan Patwa 等WWW 2026
- Disentangling Hate in Online MemesRoy Ka-Wei Lee, Rui Cao, Ziqing Fan, Jing Jiang 等ACM MM 2021 · 被引用 85 次
- Pinpointing Fine-Grained Relationships between Hateful Tweets and RepliesAbdullah Albanyan, Eduardo BlancoAAAI 2022 · 被引用 9 次
