Euphemism Identification via Feature Fusion and Individualization
Yuxue Hu, Mingmin Wu, Zhongqiang Huang, Junsong Li, Xing Ge, Ying Sha
Abstract
Euphemisms are indirect words to convey sensitive concepts. For instance, "ice" serves as a euphemism for the target keyword "methamphetamine" in illicit transactions. Euphemisms are widely used on social media and darknet marketplaces to evade moderation and supervision. Thus, euphemism identification which aims to map the euphemism to its secret meaning (target keyword) is a crucial task in ensuring social network security. However, this task poses significant challenges, including resource limitations due to the unavailable of annotated datasets and linguistic challenges arising from subtle differences in meaning between target keywords. Existing methods have employed self-supervised schemes to automatically construct labeled training data, addressing the resource limitations. Yet, these methods rely on static embedding methods that fail to distinguish between literal and euphemistic senses, leading to confusion between target keywords with similar meanings. In addition, we observe that different euphemisms in similar contexts confuse the identification results. To overcome these obstacles, we propose a feature fusion and individualization (FFI) method for euphemism identification. First, we reformulate the task as a cloze task, making it more feasible. Next, we develop a feature fusion module to capture both dynamic global and static local features, enhancing discrimination between different euphemisms in similar contexts. Additionally, we employ a feature individualization module to ensure each target keyword has a unique feature representation by projecting features into their orthogonal space. As a result, FFI can effectively identify subtle semantic differences between similar euphemisms that refer to target keywords with similar meanings. Experimental results demonstrate that our method outperforms state-of-the-art methods and large language models (GPT3.5, Llama2, mPLUG-Owl, etc.), providing robust support for its effectiveness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3622277d-12e9-4338-89a9-03dbc2f78723Builds on4
- On the Sentence Embeddings from Pre-trained Language ModelsBohan Li, Hao Zhou, Junxian He, Mingxuan Wang et al.EMNLP 2020 · 538 citations
- Feature Projection for Improved Text ClassificationQi Qin, Wenpeng Hu, Bing LiuACL 2020 · 66 citations
- Reading Thieves' Cant: Automatically Identifying and Understanding Dark Jargons from Cybercrime MarketplacesKan Yuan, Haoran Lu, Xiaojing Liao, XiaoFeng WangUSENIX Security 2018 · 56 citations
- Self-Supervised Euphemism Detection and Identification for Content ModerationWanzheng Zhu, Hongyu Gong, Rohan Bansal, Zachary Weinberg et al.S&P 2021 · 56 citations
Related papers
- Uncovering and Mitigating the Hidden Chasm: A Study on the Text-Text Domain Gap in Euphemism IdentificationYuxue Hu, Junsong Li, Mingmin Wu, Zhongqiang Huang et al.AAAI 2024 · 1 citation
- Covering Cracks in Content Moderation: Delexicalized Distant Supervision for Illicit Drug Jargon DetectionMinkyoo Song, Eugene Jang, Jaehan Kim, Seungwon ShinKDD 2025
- Euphemistic Abuse - A New Dataset and Classification Experiments for Implicitly Abusive LanguageMichael Wiegand, Jana Kampfmeier, Elisabeth Eder, Josef RuppenhoferEMNLP 2023 · 1 citation
- Cognitive Linguistic Identity Fusion Score (CLIFS): A Scalable Cognition-Informed Approach to Quantifying Identity Fusion from TextDevin R. Wright, Jisun An, Yong-Yeol AhnEMNLP 2025 · 1 citation
- Just a Scratch: Enhancing LLM Capabilities for Self-harm Detection through Intent Differentiation and Emoji InterpretationSoumitra Ghosh, Gopendra Vikram Singh, Shambhavi, Sabarna Choudhury et al.ACL 2025
