Evaluating Large Language Models for Detecting Antisemitism
Jay Patel, Hrudayangam Mehta, Jeremy Blackburn
Abstract
Detecting hateful content is a challenging and important problem. Automated tools, like machine-learning models, can help, but they require continuous training to adapt to the ever-changing landscape of social media. In this work, we evaluate eight open-source LLMs'capability to detect antisemitic content, specifically leveraging in-context definition. We also study how LLMs understand and explain their decisions given a moderation policy as a guideline. First, we explore various prompting techniques and design a new CoT-like prompt, Guided-CoT, and find that injecting domain-specific thoughts increases performance and utility. Guided-CoT handles the in-context policy well, improving performance and utility by reducing refusals across all evaluated models, regardless of decoding configuration, model size, or reasoning capability. Notably, Llama 3.1 70B outperforms fine-tuned GPT-3.5. Additionally, we examine LLM errors and introduce metrics to quantify semantic divergence in model-generated rationales, revealing notable differences and paradoxical behaviors among LLMs. Our experiments highlight the differences observed across LLMs'utility, explainability, and reliability. Code and resources available at: https://github.com/idramalab/quantify-llm-explanations
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9eabcc85-2790-4e8e-9462-3922a13b8109Builds on12
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formattingMelanie Sclar, Yejin Choi, Yulia Tsvetkov, Alane SuhrICLR 2024 · 682 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
- "Go eat a bat, Chang!": On the Emergence of Sinophobic Behavior on Web Communities in the Face of COVID-19Fatemeh Tahmasbi, Leonard Schild, Chen Ling, Jeremy Blackburn et al.WWW 2021 · 92 citations
Related papers
- Moderating New Waves of Online Hate with Chain-of-Thought Reasoning in Large Language ModelsNishant Vishwamitra, Keyan Guo, Farhan Tajwar Romit, Isabelle Ondracek et al.S&P 2024 · 29 citations
- SAHSD: Enhancing Hate Speech Detection in LLM-Powered Web Applications via Sentiment Analysis and Few-Shot LearningYulong Wang, Hong Li, Ni WeiWWW 2025 · 2 citations
- Read as You See: Guiding Unimodal LLMs for Low-Resource Explainable Harmful Meme DetectionFengjun Pan, Xiaobao Wu, Tho Quan, Anh Tuan LuuWWW 2026 · 2 citations
- Robust Adaptation of Large Multimodal Models for Retrieval Augmented Hateful Meme DetectionJingbiao Mei, Jinghong Chen, Guangyu Yang, Weizhe Lin et al.EMNLP 2025 · 2 citations
- From Meme to Threat: On the Hateful Meme Understanding and Induced Hateful Content Generation in Open-Source Vision Language ModelsYihan Ma, Xinyue Shen, Yiting Qu, Ning Yu et al.USENIX Security 2025
