Self-Supervised Euphemism Detection and Identification for Content Moderation
Wanzheng Zhu, Hongyu Gong, Rohan Bansal, Zachary Weinberg, Nicolas Christin, Giulia Fanti, Suma Bhat
Abstract
Fringe groups and organizations have a long history of using euphemisms—ordinary-sounding words with a secret meaning—to conceal what they are discussing. Nowadays, one common use of euphemisms is to evade content moderation policies enforced by social media platforms. Existing tools for enforcing policy automatically rely on keyword searches for words on a "ban list", but these are notoriously imprecise: even when limited to swearwords, they can still cause embarrassing false positives [1]. When a commonly used ordinary word acquires a euphemistic meaning, adding it to a keyword-based ban list is hopeless: consider "pot" (storage container or marijuana?) or "heater" (household appliance or firearm?) The current generation of social media companies instead hire staff to check posts manually, but this is expensive, inhumane, and not much more effective. It is usually apparent to a human moderator that a word is being used euphemistically, but they may not know what the secret meaning is, and therefore whether the message violates policy. Also, when a euphemism is banned, the group that used it need only invent another one, leaving moderators one step behind.This paper will demonstrate unsupervised algorithms that, by analyzing words in their sentence-level context, can both detect words being used euphemistically, and identify the secret meaning of each word. Compared to the existing state of the art, which uses context-free word embeddings, our algorithm for detecting euphemisms achieves 30–400% higher detection accuracies of unlabeled euphemisms in a text corpus. Our algorithm for revealing euphemistic meanings of words is the first of its kind, as far as we are aware. In the arms race between content moderators and policy evaders, our algorithms may help shift the balance in the direction of the moderators.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e16de4dc-b9c7-48e4-9507-1bfb9b599988Cited by top-tier papers11
- DarkBERT: A Language Model for the Dark Side of the InternetYoungjin Jin, Eugene Jang, Jian Cui, Jin-Woo Chung et al.ACL 2023 · 41 citations
- "Get in Researchers; We're Measuring Reproducibility": A Reproducibility Study of Machine Learning Papers in Tier 1 Security ConferencesDaniel Olszewski, Allison Lu, Carson Stillman, Kevin Warren et al.CCS 2023 · 19 citations
- Detecting and Understanding the Promotion of Illicit Goods and Services on TwitterHongyu Wang, Ying Li, Ronghong Huang, Xianghang MiWWW 2025 · 6 citations
- Making FETCH! Happen: Finding Emergent Dog Whistles Through Common HabitatsKuleen Sasse, Carlos Alejandro Aguirre, Isabel Cachola, Sharon Levy et al.ACL 2025 · 3 citations
- Euphemism Identification via Feature Fusion and IndividualizationYuxue Hu, Mingmin Wu, Zhongqiang Huang, Junsong Li et al.WWW 2024 · 1 citation
Builds on15
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Big Self-Supervised Models are Strong Semi-Supervised LearnersTing Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi et al.NeurIPS 2020 · 2,611 citations
- S4L: Self-Supervised Semi-Supervised LearningLucas Beyer, Xiaohua Zhai, Avital Oliver, Alexander KolesnikovICCV 2019 · 854 citations
- Skew-Fit: State-Covering Self-Supervised Reinforcement LearningVitchyr Pong, Murtaza Dalal, Steven Lin, Ashvin Nair et al.ICML 2020 · 303 citations
Related papers
- Covering Cracks in Content Moderation: Delexicalized Distant Supervision for Illicit Drug Jargon DetectionMinkyoo Song, Eugene Jang, Jaehan Kim, Seungwon ShinKDD 2025
- Uncovering and Mitigating the Hidden Chasm: A Study on the Text-Text Domain Gap in Euphemism IdentificationYuxue Hu, Junsong Li, Mingmin Wu, Zhongqiang Huang et al.AAAI 2024 · 1 citation
- From Dogwhistles to Bullhorns: Unveiling Coded Rhetoric with Language ModelsJulia Mendelsohn, Ronan Le Bras, Yejin Choi, Maarten SapACL 2023 · 14 citations
- Silent Signals, Loud Impact: LLMs for Word-Sense Disambiguation of Coded Dog WhistlesJulia Kruk, Michela Marchini, Rijul Magu, Caleb Ziems et al.ACL 2024 · 2 citations
- New Terms, New Toxicity: Consensus-based Chinese Neologism Toxicity Detection via Search-Augmented LLMsShiyao Cui, Qinglin Zhang, Di Wang, Yida Lu et al.ACL 2026
