Uncovering and Mitigating the Hidden Chasm: A Study on the Text-Text Domain Gap in Euphemism Identification
Yuxue Hu, Junsong Li, Mingmin Wu, Zhongqiang Huang, Gang Chen, Ying Sha
摘要
Euphemisms are commonly used on social media and darknet marketplaces to evade platform regulations by masking their true meanings with innocent ones. For instance, “weed” is used instead of “marijuana” for illicit transactions. Thus, euphemism identification, i.e., mapping a given euphemism (“weed”) to its specific target word (“marijuana”), is essential for improving content moderation and combating underground markets. Existing methods employ self-supervised schemes to automatically construct labeled training datasets for euphemism identification. However, they overlook the text-text domain gap caused by the discrepancy between the constructed training data and the test data, leading to performance deterioration. In this paper, we present the text-text domain gap and explain how it forms in terms of the data distribution and the cone effect. Moreover, to bridge this gap, we introduce a feature alignment network (FA-Net), which can both align the in-domain and cross-domain features, thus mitigating the domain gap from training data to test data and improving the performance of the base models for euphemism identification. We apply this FA-Net to the base models, obtaining markedly better results, and creating a state-of-the-art model which beats the large language models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- McHirc: A Multimodal Benchmark for Chinese Idiom Reading ComprehensionTongguan Wang, Mingmin Wu, Guixin Su, Dongyu Su 等AAAI 2025 · 被引用 4 次
- Chinese Two-part Allegorical Sayings Reading Comprehension: Exploration from Reasoning to MetaphorDongyu Su, Yimin Xiao, Tongguan Wang, Feiyue Xue 等AAAI 2026
它引用的顶会 Paper8
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen 等ICLR 2021 · 被引用 1,731 次
- VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text UnderstandingHu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko 等EMNLP 2021 · 被引用 399 次
- Invariant Risk Minimization GamesKartik Ahuja, Karthikeyan Shanmugam, Kush R. Varshney, Amit DhurandharICML 2020 · 被引用 289 次
- Reading Thieves' Cant: Automatically Identifying and Understanding Dark Jargons from Cybercrime MarketplacesKan Yuan, Haoran Lu, Xiaojing Liao, XiaoFeng WangUSENIX Security 2018 · 被引用 56 次
相关 Paper
- Euphemism Identification via Feature Fusion and IndividualizationYuxue Hu, Mingmin Wu, Zhongqiang Huang, Junsong Li 等WWW 2024 · 被引用 1 次
- Self-Supervised Euphemism Detection and Identification for Content ModerationWanzheng Zhu, Hongyu Gong, Rohan Bansal, Zachary Weinberg 等S&P 2021 · 被引用 56 次
- Covering Cracks in Content Moderation: Delexicalized Distant Supervision for Illicit Drug Jargon DetectionMinkyoo Song, Eugene Jang, Jaehan Kim, Seungwon ShinKDD 2025
- Unsupervised Intra-Domain Adaptation for Semantic Segmentation Through Self-SupervisionFei Pan, Inkyu Shin, François Rameau, Seokju Lee 等CVPR 2020
- Enhancing Fake News Detection in Social Media via Label Propagation on Cross-modal Tweet GraphWanqing Zhao, Yuta Nakashima, Haiyuan Chen, Noboru BabaguchiACM MM 2023 · 被引用 10 次
