Uncovering and Mitigating the Hidden Chasm: A Study on the Text-Text Domain Gap in Euphemism Identification
Yuxue Hu, Junsong Li, Mingmin Wu, Zhongqiang Huang, Gang Chen, Ying Sha
Abstract
Euphemisms are commonly used on social media and darknet marketplaces to evade platform regulations by masking their true meanings with innocent ones. For instance, “weed” is used instead of “marijuana” for illicit transactions. Thus, euphemism identification, i.e., mapping a given euphemism (“weed”) to its specific target word (“marijuana”), is essential for improving content moderation and combating underground markets. Existing methods employ self-supervised schemes to automatically construct labeled training datasets for euphemism identification. However, they overlook the text-text domain gap caused by the discrepancy between the constructed training data and the test data, leading to performance deterioration. In this paper, we present the text-text domain gap and explain how it forms in terms of the data distribution and the cone effect. Moreover, to bridge this gap, we introduce a feature alignment network (FA-Net), which can both align the in-domain and cross-domain features, thus mitigating the domain gap from training data to test data and improving the performance of the base models for euphemism identification. We apply this FA-Net to the base models, obtaining markedly better results, and creating a state-of-the-art model which beats the large language models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f070aafc-7dd2-4e92-bb85-521dd6690941Cited by top-tier papers2
- McHirc: A Multimodal Benchmark for Chinese Idiom Reading ComprehensionTongguan Wang, Mingmin Wu, Guixin Su, Dongyu Su et al.AAAI 2025 · 4 citations
- Chinese Two-part Allegorical Sayings Reading Comprehension: Exploration from Reasoning to MetaphorDongyu Su, Yimin Xiao, Tongguan Wang, Feiyue Xue et al.AAAI 2026
Builds on8
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen et al.ICLR 2021 · 1,731 citations
- VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text UnderstandingHu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko et al.EMNLP 2021 · 399 citations
- Invariant Risk Minimization GamesKartik Ahuja, Karthikeyan Shanmugam, Kush R. Varshney, Amit DhurandharICML 2020 · 289 citations
- Reading Thieves' Cant: Automatically Identifying and Understanding Dark Jargons from Cybercrime MarketplacesKan Yuan, Haoran Lu, Xiaojing Liao, XiaoFeng WangUSENIX Security 2018 · 56 citations
Related papers
- Euphemism Identification via Feature Fusion and IndividualizationYuxue Hu, Mingmin Wu, Zhongqiang Huang, Junsong Li et al.WWW 2024 · 1 citation
- Self-Supervised Euphemism Detection and Identification for Content ModerationWanzheng Zhu, Hongyu Gong, Rohan Bansal, Zachary Weinberg et al.S&P 2021 · 56 citations
- Covering Cracks in Content Moderation: Delexicalized Distant Supervision for Illicit Drug Jargon DetectionMinkyoo Song, Eugene Jang, Jaehan Kim, Seungwon ShinKDD 2025
- Unsupervised Intra-Domain Adaptation for Semantic Segmentation Through Self-SupervisionFei Pan, Inkyu Shin, François Rameau, Seokju Lee et al.CVPR 2020
- Enhancing Fake News Detection in Social Media via Label Propagation on Cross-modal Tweet GraphWanqing Zhao, Yuta Nakashima, Haiyuan Chen, Noboru BabaguchiACM MM 2023 · 10 citations
