How to Learn Klingon without a Dictionary: Detection and Measurement of Black Keywords Used by the Underground Economy
Hao Yang, Xiulin Ma, Kun Du, Zhou Li, Hai-Xin Duan, XiaoDong Su, Guang Liu, Zhifeng Geng, Jianping Wu
摘要
used vocabularies. Second, many black keywords are created in ungrammatical or even obfuscated forms (e.g., letter 'l' could be replaced by digit '1' in a word) while the NLP techniques are suitable for well-written text [4].
On the other hand, based on our prior experience in investigating underground economy, we found this challenge can be addressed through a pure data-driven approach. Many underground merchants rely on blackhat SEO (search engine optimization) to promote their business. Usually, plenty of black keywords are stuffed into one SEO page inside certain HTML tags (e.g., anchor tags) to fool the search engines, yet making themselves distinguishable under content analysis. We could extract them from the SEO pages but the irrelevant texts have to be pruned. It turns out that the search results associated with a candidate keyword can be leveraged to determine whether the keyword is "black": as revealed by our study, querying a black keyword usually returns multiple links alarmed by the existing scanners, so we can use the result as the main indicator. Additionally, we found our list of black keywords can be extended through related search, a feature presented by major search engines to correlate similar search terms based on users' searching behaviors. After these steps, a lot of black keywords can be discovered, but a large portion of them are long-tail keywords which contain words not of our interest (e.g., words except "heroin" in "where to buy heroin in Beijing"). To extract the core words (e.g., "heroin" in the above example), we devised a substring matching algorithm which can process the keywords very efficiently.
We developed KDES (Keywords Detection and Expansion System) and evaluated it on more than 2 million pages related to SEO, porn and gambling. We discovered 478,879 black keywords in total and extracted 1,522 core words (433,335 black keywords are covered). After sampling the detected keywords, we found that the accuracy can achieve 94.3%, suggesting KDES is effective. We applied our findings to Baidu and the feedback was very encouraging. Many of the detected keywords have been added into their internal blacklist.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Reading Thieves' Cant: Automatically Identifying and Understanding Dark Jargons from Cybercrime MarketplacesKan Yuan, Haoran Lu, Xiaojing Liao, XiaoFeng WangUSENIX Security 2018 · 被引用 56 次
- Self-Supervised Euphemism Detection and Identification for Content ModerationWanzheng Zhu, Hongyu Gong, Rohan Bansal, Zachary Weinberg 等S&P 2021 · 被引用 56 次
- Scalable Detection of Promotional Website Defacements in Black Hat SEO CampaignsRonghai Yang, Xianbo Wang, Cheng Chi, Dawei Wang 等USENIX Security 2021 · 被引用 27 次
- Detecting and Understanding the Promotion of Illicit Goods and Services on TwitterHongyu Wang, Ying Li, Ronghong Huang, Xianghang MiWWW 2025 · 被引用 6 次
- Making FETCH! Happen: Finding Emergent Dog Whistles Through Common HabitatsKuleen Sasse, Carlos Alejandro Aguirre, Isabel Cachola, Sharon Levy 等ACL 2025 · 被引用 3 次
它引用的顶会 Paper6
- Fast, Lean, and Accurate: Modeling Password Guessability Using Neural NetworksWilliam Melicher, Blase Ur, Sean M. Segreti, Saranga Komanduri 等USENIX Security 2016 · 被引用 331 次
- Cloak of Visibility: Detecting When Machines Browse a Different WebLuca Invernizzi, Kurt Thomas, Alexandros Kapravelos, Oxana Comanescu 等S&P 2016 · 被引用 93 次
- Investigating Commercial Pay-Per-Install and the Distribution of Unwanted SoftwareKurt Thomas, Juan A. Elices Crespo, Ryan Rasti, Jean-Michel Picod 等USENIX Security 2016 · 被引用 77 次
- Measuring PUP Prevalence and PUP Distribution through Pay-Per-Install ServicesPlaton Kotzias, Leyla Bilge, Juan CaballeroUSENIX Security 2016 · 被引用 74 次
- Seeking Nonsense, Looking for Trouble: Efficient Promotional-Infection Detection through Semantic Inconsistency SearchXiaojing Liao, Kan Yuan, XiaoFeng Wang, Zhongyu Pei 等S&P 2016 · 被引用 41 次
相关 Paper
- NOKEScam: Understanding and Rectifying Non-Sense Keywords Spear Scam in Search EnginesMingxuan Liu, Yunyi Zhang, Lijie Wu, Baojun Liu 等USENIX Security 2025
- Game of Missuggestions: Semantic Analysis of Search-Autocomplete ManipulationsPeng Wang, Xianghang Mi, Xiaojing Liao, XiaoFeng Wang 等NDSS 2018 · 被引用 32 次
- Into the Dark: Unveiling Internal Site Search Abused for Black Hat SEOYunyi Zhang, Mingxuan Liu, Baojun Liu, Yiming Zhang 等USENIX Security 2024 · 被引用 1 次
- The Ever-Changing Labyrinth: A Large-Scale Analysis of Wildcard DNS Powered Blackhat SEOKun Du, Hao Yang, Zhou Li, Hai-Xin Duan 等USENIX Security 2016 · 被引用 40 次
- Measuring and Analyzing Search Engine Poisoning of Linguistic CollisionsMatthew Joslin, Neng Li, Shuang Hao, Minhui Xue 等S&P 2019 · 被引用 18 次
