Game of Missuggestions: Semantic Analysis of Search-Autocomplete Manipulations
Peng Wang, Xianghang Mi, Xiaojing Liao, XiaoFeng Wang, Kan Yuan, Feng Qian, Raheem A. Beyah
Abstract
As a new type of blackhat Search Engine Optimization (SEO), autocomplete manipulations are increasingly utilized by miscreants and promotion companies alike to advertise desired suggestion terms when related trigger terms are entered by the user into a search engine. Like other illicit SEO, such activities game the search engine, mislead the querier, and in some cases, spread harmful content. However, little has been done to understand this new threat, in terms of its scope, impact and techniques, not to mention any serious effort to detect such manipulated terms on a large scale. Systematic analysis of autocomplete manipulation is challenging, due to the scale of the problem (tens or even hundreds of millions suggestion terms and their search results) and the heavy burdens it puts on the search engines. In this paper, we report the first technique that addresses these challenges, making a step toward better understanding and ultimately eliminating this new threat. Our technique, called Sacabuche, takes a semantics-based, two-step approach to minimize its performance impact: it utilizes Natural Language Processing (NLP) to analyze a large number of trigger and suggestion combinations, without querying search engines, to filter out the vast majority of legitimate suggestion terms; only a small set of suspicious suggestions are run against the search engines to get query results for identifying truly abused terms. This approach achieves a 96.23% precision and 95.63% recall, and its scalability enables us to perform a measurement study on 114 millions of suggestion terms, an unprecedented scale for this type of studies. The findings of the study bring to light the magnitude of the threat (0.48% Google suggestion terms we collected manipulated), and its significant security implications never reported before (e.g., exceedingly long lifetime of campaigns, sophisticated techniques and channels for spreading malware and phishing content).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 72631f8d-32f0-4d8d-b0ce-cdbd40263537Cited by top-tier papers8
- When Are Search Completion Suggestions Problematic?Alexandra Olteanu, Fernando Diaz, Gabriella KazaiCSCW 2020 · 38 citations
- Scalable Detection of Promotional Website Defacements in Black Hat SEO CampaignsRonghai Yang, Xianbo Wang, Cheng Chi, Dawei Wang et al.USENIX Security 2021 · 27 citations
- "Get in Researchers; We're Measuring Reproducibility": A Reproducibility Study of Machine Learning Papers in Tier 1 Security ConferencesDaniel Olszewski, Allison Lu, Carson Stillman, Kevin Warren et al.CCS 2023 · 19 citations
- Measuring and Analyzing Search Engine Poisoning of Linguistic CollisionsMatthew Joslin, Neng Li, Shuang Hao, Minhui Xue et al.S&P 2019 · 18 citations
- Detecting and Measuring Misconfigured Manifests in Android AppsYuqing Yang, Mohamed Elsabagh, Chaoshun Zuo, Ryan Johnson et al.CCS 2022 · 11 citations
Builds on3
- Fake Co-visitation Injection Attacks to Recommender SystemsGuolei Yang, Neil Zhenqiang Gong, Ying CaiNDSS 2017 · 126 citations
- Under the Shadow of Sunshine: Understanding and Detecting Bulletproof Hosting on Legitimate Service Provider NetworksSumayah A. Alrwais, Xiaojing Liao, Xianghang Mi, Peng Wang et al.S&P 2017 · 51 citations
- Seeking Nonsense, Looking for Trouble: Efficient Promotional-Infection Detection through Semantic Inconsistency SearchXiaojing Liao, Kan Yuan, XiaoFeng Wang, Zhongyu Pei et al.S&P 2016 · 41 citations
Related papers
- How to Learn Klingon without a Dictionary: Detection and Measurement of Black Keywords Used by the Underground EconomyHao Yang, Xiulin Ma, Kun Du, Zhou Li et al.S&P 2017 · 48 citations
- Into the Dark: Unveiling Internal Site Search Abused for Black Hat SEOYunyi Zhang, Mingxuan Liu, Baojun Liu, Yiming Zhang et al.USENIX Security 2024 · 1 citation
- NOKEScam: Understanding and Rectifying Non-Sense Keywords Spear Scam in Search EnginesMingxuan Liu, Yunyi Zhang, Lijie Wu, Baojun Liu et al.USENIX Security 2025
- Exposing the Hidden Layer: Software Repositories in the Service of Seo ManipulationMengying Wu, Geng Hong, Wuyuao Mai, Xinyi Wu et al.ICSE 2025 · 1 citation
- The Ever-Changing Labyrinth: A Large-Scale Analysis of Wildcard DNS Powered Blackhat SEOKun Du, Hao Yang, Zhou Li, Hai-Xin Duan et al.USENIX Security 2016 · 40 citations
