Euphemism Identification via Feature Fusion and Individualization
Yuxue Hu, Mingmin Wu, Zhongqiang Huang, Junsong Li, Xing Ge, Ying Sha
摘要
Euphemisms are indirect words to convey sensitive concepts. For instance, "ice" serves as a euphemism for the target keyword "methamphetamine" in illicit transactions. Euphemisms are widely used on social media and darknet marketplaces to evade moderation and supervision. Thus, euphemism identification which aims to map the euphemism to its secret meaning (target keyword) is a crucial task in ensuring social network security. However, this task poses significant challenges, including resource limitations due to the unavailable of annotated datasets and linguistic challenges arising from subtle differences in meaning between target keywords. Existing methods have employed self-supervised schemes to automatically construct labeled training data, addressing the resource limitations. Yet, these methods rely on static embedding methods that fail to distinguish between literal and euphemistic senses, leading to confusion between target keywords with similar meanings. In addition, we observe that different euphemisms in similar contexts confuse the identification results. To overcome these obstacles, we propose a feature fusion and individualization (FFI) method for euphemism identification. First, we reformulate the task as a cloze task, making it more feasible. Next, we develop a feature fusion module to capture both dynamic global and static local features, enhancing discrimination between different euphemisms in similar contexts. Additionally, we employ a feature individualization module to ensure each target keyword has a unique feature representation by projecting features into their orthogonal space. As a result, FFI can effectively identify subtle semantic differences between similar euphemisms that refer to target keywords with similar meanings. Experimental results demonstrate that our method outperforms state-of-the-art methods and large language models (GPT3.5, Llama2, mPLUG-Owl, etc.), providing robust support for its effectiveness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- On the Sentence Embeddings from Pre-trained Language ModelsBohan Li, Hao Zhou, Junxian He, Mingxuan Wang 等EMNLP 2020 · 被引用 538 次
- Feature Projection for Improved Text ClassificationQi Qin, Wenpeng Hu, Bing LiuACL 2020 · 被引用 66 次
- Reading Thieves' Cant: Automatically Identifying and Understanding Dark Jargons from Cybercrime MarketplacesKan Yuan, Haoran Lu, Xiaojing Liao, XiaoFeng WangUSENIX Security 2018 · 被引用 56 次
- Self-Supervised Euphemism Detection and Identification for Content ModerationWanzheng Zhu, Hongyu Gong, Rohan Bansal, Zachary Weinberg 等S&P 2021 · 被引用 56 次
相关 Paper
- Uncovering and Mitigating the Hidden Chasm: A Study on the Text-Text Domain Gap in Euphemism IdentificationYuxue Hu, Junsong Li, Mingmin Wu, Zhongqiang Huang 等AAAI 2024 · 被引用 1 次
- Covering Cracks in Content Moderation: Delexicalized Distant Supervision for Illicit Drug Jargon DetectionMinkyoo Song, Eugene Jang, Jaehan Kim, Seungwon ShinKDD 2025
- Euphemistic Abuse - A New Dataset and Classification Experiments for Implicitly Abusive LanguageMichael Wiegand, Jana Kampfmeier, Elisabeth Eder, Josef RuppenhoferEMNLP 2023 · 被引用 1 次
- Cognitive Linguistic Identity Fusion Score (CLIFS): A Scalable Cognition-Informed Approach to Quantifying Identity Fusion from TextDevin R. Wright, Jisun An, Yong-Yeol AhnEMNLP 2025 · 被引用 1 次
- Just a Scratch: Enhancing LLM Capabilities for Self-harm Detection through Intent Differentiation and Emoji InterpretationSoumitra Ghosh, Gopendra Vikram Singh, Shambhavi, Sabarna Choudhury 等ACL 2025
