USENIX Security2018Top-tier venue
Reading Thieves' Cant: Automatically Identifying and Understanding Dark Jargons from Cybercrime Marketplaces
Kan Yuan, Haoran Lu, Xiaojing Liao, XiaoFeng Wang
Abstract
Underground communication is invaluable for understanding cybercrimes. However, it is often obfuscated by the extensive use of dark jargons, innocently-looking terms like "popcorn" that serves sinister purposes (buying/selling drug, seeking crimeware, etc.). Discovery and understanding of these jargons have so far relied on manual effort, which is error-prone and cannot catch up with the fast evolving underground ecosystem. In this paper, we present the first technique, called Cantreader, to automatically detect and understand dark jargon. Our approach employs a neural-network based embedding technique to analyze the semantics of words, detecting those whose contexts in legitimate documents are significantly different from those in underground communication. For this purpose, we enhance the existing word embedding model to support semantic comparison across good and bad corpora, which leads to the detection of dark jargons. To further understand them, our approach utilizes projection learning to identify a jargon's hypernym that sheds light on its true meaning when used in underground communication. Running Cantreader over one million traces collected from four underground forums, our approach automatically reported 3,462 dark jargons and their hypernyms, including 2,491 never known before. The study further reveals how these jargons are used (by 25% of the traces) and evolve and how they help cybercriminals communicate on legitimate forums. Reading thieves' cant: challenges. With their pervasiveness in underground communication, dark jargons are surprisingly elusive and difficult to catch, due to their innocent-looking disguises and the ways they are used,
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b44899e6-f3c1-4da8-a464-9b5fc80ea3cdCited by top-tier papers14
- Self-Supervised Euphemism Detection and Identification for Content ModerationWanzheng Zhu, Hongyu Gong, Rohan Bansal, Zachary Weinberg et al.S&P 2021 · 56 citations
- DarkBERT: A Language Model for the Dark Side of the InternetYoungjin Jin, Eugene Jang, Jian Cui, Jin-Woo Chung et al.ACL 2023 · 41 citations
- Why Should Adversarial Perturbations be Imperceptible? Rethink the Research Paradigm in Adversarial NLPYangyi Chen, Hongcheng Gao, Ganqu Cui, Fanchao Qi et al.EMNLP 2022 · 28 citations
- Scalable Detection of Promotional Website Defacements in Black Hat SEO CampaignsRonghai Yang, Xianbo Wang, Cheng Chi, Dawei Wang et al.USENIX Security 2021 · 27 citations
- "Get in Researchers; We're Measuring Reproducibility": A Reproducibility Study of Machine Learning Papers in Tier 1 Security ConferencesDaniel Olszewski, Allison Lu, Carson Stillman, Kevin Warren et al.CCS 2023 · 19 citations
Builds on3
- Acing the IOC Game: Toward Automatic Discovery and Analysis of Open-Source Cyber Threat IntelligenceXiaojing Liao, Kan Yuan, XiaoFeng Wang, Zhou Li et al.CCS 2016 · 308 citations
- How to Learn Klingon without a Dictionary: Detection and Measurement of Black Keywords Used by the Underground EconomyHao Yang, Xiulin Ma, Kun Du, Zhou Li et al.S&P 2017 · 48 citations
- Seeking Nonsense, Looking for Trouble: Efficient Promotional-Infection Detection through Semantic Inconsistency SearchXiaojing Liao, Kan Yuan, XiaoFeng Wang, Zhongyu Pei et al.S&P 2016 · 41 citations
Related papers
- Breaking Free from Ivory Tower: Evaluating and Enhancing Real-world Chinese Underground Adversarial Jargon DetectionZhifan Jiang, Mingxuan Liu, Yue Qin, Baojun LiuS&P 2026 · 2 citations
- Covering Cracks in Content Moderation: Delexicalized Distant Supervision for Illicit Drug Jargon DetectionMinkyoo Song, Eugene Jang, Jaehan Kim, Seungwon ShinKDD 2025
- DarkGram: A Large-Scale Analysis of Cybercriminal Activity Channels on TelegramSayak Saha Roy, Elham Pourabbas Vafa, Kobra Khanmohamaddi, Shirin NilizadehUSENIX Security 2025
- VendorLink: An NLP approach for Identifying & Linking Vendor Migrants & Potential Aliases on Darknet MarketsVageesh Saxena, Nils Rethmeier, Gijs van Dijck, Gerasimos SpanakisACL 2023 · 6 citations
- Understanding and Analyzing Appraisal Systems in the Underground MarketplacesZhengyi Li, Xiaojing LiaoNDSS 2024
