AmpleHate: Amplifying the Attention for Versatile Implicit Hate Detection
Yejin Lee, Joonghyuk Hahn, Hyeseon Ahn, Yo-Sub Han
Abstract
Implicit hate speech involves subtle and indirect expressions of prejudice or hostility toward a group.Detecting it is challenging because it relies on nuanced context and implication rather than explicit offensive language.Current approaches rely on contrastive learning, which is shown to be effective on distinguishing hate and non-hate sentences.Humans, however, detect implicit hate speech by first identifying specific targets within the text and subsequently interpreting how these targets relate to their surrounding context.Motivated by this reasoning process, we propose Ample-Hate, a novel approach designed to mirror human inference for implicit hate detection.Am-pleHate identifies explicit targets using a pretrained Named Entity Recognition model and captures implicit target information via [CLS] tokens.It computes attention-based relationships between explicit, implicit targets and sentence context and then, directly injects these relational vectors into the final sentence representation.This amplifies the critical signals of target-context relations for determining implicit hate.Experiments demonstrate that Am-pleHate achieves state-of-the-art performance, outperforming contrastive learning baselines by an average of 82.14% and achieves faster convergence.Qualitative analyses further reveal that attention patterns produced by Am-pleHate closely align with human judgement, underscoring its interpretability and robustness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on7
- Latent Hatred: A Benchmark for Understanding Implicit Hate SpeechMai ElSherief, Caleb Ziems, David Muchlinski, Vaishnavi Anupindi et al.EMNLP 2021 · 159 citations
- Social Bias Frames: Reasoning about Social and Power Implications of LanguageMaarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky et al.ACL 2020 · 16 citations
- Word Embeddings Are Steers for Language ModelsChi Han, Jialiang Xu, Manling Li, Yi Fung et al.ACL 2024 · 8 citations
- Hate Speech Detection with Generalizable Target-aware FairnessTong Chen, Danny Wang, Xurong Liang, Marten Risius et al.KDD 2024 · 7 citations
- Learning from the Worst: Dynamically Generated Datasets to Improve Online Hate DetectionBertie Vidgen, Tristan Thrush, Zeerak Waseem, Douwe KielaACL 2021
Related papers
- ImpliHateVid: A Benchmark Dataset and Two-stage Contrastive Learning Framework for Implicit Hate Speech Detection in VideosMohammad Zia Ur Rehman, Anukriti Bhatnagar, Omkar Kabde, Shubhi Bansal et al.ACL 2025 · 11 citations
- CoSyn: Detecting Implicit Hate Speech in Online Conversations Using a Context Synergized Hyperbolic NetworkSreyan Ghosh, Manan Suri, Purva Chiniya, Utkarsh Tyagi et al.EMNLP 2023 · 9 citations
- Causality Guided Representation Learning for Cross-Style Hate Speech DetectionChengshuai Zhao, Shu Wan, Paras Sheth, Karan Patwa et al.WWW 2026
- Disentangling Hate in Online MemesRoy Ka-Wei Lee, Rui Cao, Ziqing Fan, Jing Jiang et al.ACM MM 2021 · 85 citations
- Pinpointing Fine-Grained Relationships between Hateful Tweets and RepliesAbdullah Albanyan, Eduardo BlancoAAAI 2022 · 9 citations
