Seeking Nonsense, Looking for Trouble: Efficient Promotional-Infection Detection through Semantic Inconsistency Search
Xiaojing Liao, Kan Yuan, XiaoFeng Wang, Zhongyu Pei, Hao Yang, Jianjun Chen, Hai-Xin Duan, Kun Du, Eihal Alowaisheq, Sumayah A. Alrwais, Luyi Xing, Raheem A. Beyah
Abstract
Promotional infection is an attack in which the adversary exploits a website's weakness to inject illicit advertising content. Detection of such an infection is challenging due to its similarity to legitimate advertising activities. An interesting observation we make in our research is that such an attack almost always incurs a great semantic gap between the infected domain (e.g., a university site) and the content it promotes (e.g., selling cheap viagra). Exploiting this gap, we developed a semantic-based technique, called Semantic Inconsistency Search (SEISE), for efficient and accurate detection of the promotional injections on sponsored top-level domains (sTLD) with explicit semantic meanings. Our approach utilizes Natural Language Processing (NLP) to identify the bad terms (those related to illicit activities like fake drug selling, etc.) most irrelevant to an sTLD's semantics. These terms, which we call irrelevant bad terms (IBTs), are used to query search engines under the sTLD for suspicious domains. Through a semantic analysis on the results page returned by the search engines, SEISE is able to detect those truly infected sites and automatically collect new IBTs from the titles/URLs/snippets of their search result items for finding new infections. Running on 403 sTLDs with an initial 30 seed IBTs, SEISE analyzed 100K fully qualified domain names (FQDN), and along the way automatically gathered nearly 600 IBTs. In the end, our approach detected 11K infected FQDN with a false detection rate of 1.5% and over 90% coverage. Our study shows that by effective detection of infected sTLDs, the bar to promotion infections can be substantially raised, since other non-sTLD vulnerable domains typically have much lower Alexa ranks and are therefore much less attractive for underground advertising. Our findings further bring to light the stunning impacts of such promotional attacks, which compromise FQDNs under 3% of .edu, .gov domains and over one thousand gov.cn domains, including those of leading universities such as stanford.edu, mit.edu, princeton.edu, havard.edu and government institutes such as nsf.gov and nih.gov. We further demonstrate the potential to extend our current technique to protect generic domains such as .com and .org.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ff75d4e8-f7f4-419f-b8a6-635898f8e847Cited by top-tier papers18
- Acing the IOC Game: Toward Automatic Discovery and Analysis of Open-Source Cyber Threat IntelligenceXiaojing Liao, Kan Yuan, XiaoFeng Wang, Zhou Li et al.CCS 2016 · 308 citations
- Resident Evil: Understanding Residential IP Proxy as a Dark ServiceXianghang Mi, Xuan Feng, Xiaojing Liao, Baojun Liu et al.S&P 2019 · 80 citations
- Reading Thieves' Cant: Automatically Identifying and Understanding Dark Jargons from Cybercrime MarketplacesKan Yuan, Haoran Lu, Xiaojing Liao, XiaoFeng WangUSENIX Security 2018 · 56 citations
- How to Learn Klingon without a Dictionary: Detection and Measurement of Black Keywords Used by the Underground EconomyHao Yang, Xiulin Ma, Kun Du, Zhou Li et al.S&P 2017 · 48 citations
- Understanding Malicious Cross-library Data Harvesting on AndroidJice Wang, Yue Xiao, Xueqiang Wang, Yuhong Nan et al.USENIX Security 2021 · 41 citations
Related papers
- MAWSEO: Adversarial Wiki Search Poisoning for Illicit Online PromotionZilong Lin, Zhengyi Li, Xiaojing Liao, XiaoFeng Wang et al.S&P 2024 · 16 citations
- Scalable Detection of Promotional Website Defacements in Black Hat SEO CampaignsRonghai Yang, Xianbo Wang, Cheng Chi, Dawei Wang et al.USENIX Security 2021 · 27 citations
- Into the Dark: Unveiling Internal Site Search Abused for Black Hat SEOYunyi Zhang, Mingxuan Liu, Baojun Liu, Yiming Zhang et al.USENIX Security 2024 · 1 citation
- Measuring and Analyzing Search Engine Poisoning of Linguistic CollisionsMatthew Joslin, Neng Li, Shuang Hao, Minhui Xue et al.S&P 2019 · 18 citations
- Ranking Manipulation for Conversational Search EnginesSamuel Pfrommer, Yatong Bai, Tanmay Gautam, Somayeh SojoudiEMNLP 2024 · 4 citations
