Measuring and Analyzing Search Engine Poisoning of Linguistic Collisions
Matthew Joslin, Neng Li, Shuang Hao, Minhui Xue, Haojin Zhu
Abstract
Misspelled keywords have become an appealing target in search poisoning, since they are less competitive to promote than the correct queries and account for a considerable amount of search traffic. Search engines have adopted several countermeasure strategies, e.g., Google applies automated corrections on queried keywords and returns search results of the corrected versions directly. However, a sophisticated class of attack, which we term as linguistic-collision misspelling, can evade auto-correction and poison search results. Cybercriminals target special queries where the misspelled terms are existent words, even in other languages (e.g., "idobe", a misspelling of the English word "adobe", is a legitimate word in the Nigerian language). In this paper, we perform the first large-scale analysis on linguistic-collision search poisoning attacks. In particular, we check 1.77 million misspelled search terms on Google and Baidu and analyze both English and Chinese languages, which are the top two languages used by Internet users [1] . We leverage edit distance operations and linguistic properties to generate misspelling candidates. To more efficiently identify linguisticcollision search terms, we design a deep learning model that can improve collection rate by 2.84x compared to random sampling. Our results show that the abuse is prevalent: around 1.19% of linguistic-collision search terms on Google and Baidu have results on the first page directing to malicious websites. We also find that cybercriminals mainly target categories of gambling, drugs, and adult content. Mobile-device users disproportionately search for misspelled keywords, presumably due to small screen for input. Our work highlights this new class of search engine poisoning and provides insights to help mitigate the threat.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f07dbf5e-d006-4173-9530-eae855f271afCited by top-tier papers8
- Towards More Practical Threat Models in Artificial Intelligence SecurityKathrin Grosse, Lukas Bieringer, Tarek R. Besold, Alexandre AlahiUSENIX Security 2024 · 27 citations
- Scalable Detection of Promotional Website Defacements in Black Hat SEO CampaignsRonghai Yang, Xianbo Wang, Cheng Chi, Dawei Wang et al.USENIX Security 2021 · 27 citations
- "Get in Researchers; We're Measuring Reproducibility": A Reproducibility Study of Machine Learning Papers in Tier 1 Security ConferencesDaniel Olszewski, Allison Lu, Carson Stillman, Kevin Warren et al.CCS 2023 · 19 citations
- When Search Goes Wrong: Red-Teaming Web-Augmented Large Language ModelsHaoran Ou, Kangjie Chen, Xingshuo Han, Gelei Deng et al.ICML 2026 · 2 citations
- Exposing the Hidden Layer: Software Repositories in the Service of Seo ManipulationMengying Wu, Geng Hong, Wuyuao Mai, Xinyi Wu et al.ICSE 2025 · 1 citation
Builds on7
- DeepLog: Anomaly Detection and Diagnosis from System Logs through Deep LearningMin Du, Feifei Li, Guineng Zheng, Vivek SrikumarCCS 2017 · 1,823 citations
- Automated Website Fingerprinting through Deep LearningVera Rimmer, Davy Preuveneers, Marc Juarez, Tom van Goethem et al.NDSS 2018 · 399 citations
- Automated Crowdturfing Attacks and Defenses in Online Review SystemsYuanshun Yao, Bimal Viswanath, Jenna Cryan, Haitao Zheng et al.CCS 2017 · 169 citations
- Hiding in Plain Sight: A Longitudinal Study of Combosquatting AbusePanagiotis Kintis, Najmeh Miramirkhani, Charles Lever, Yizheng Chen et al.CCS 2017 · 166 citations
- Dial One for Scam: A Large-Scale Analysis of Technical Support ScamsNajmeh Miramirkhani, Oleksii Starov, Nick NikiforakisNDSS 2017 · 116 citations
Related papers
- Enhancing Chinese Offensive Language Detection with Homophonic PerturbationJunqi Wu, Shujie Ji, Kang Zhong, Huiling Peng et al.EMNLP 2025
- Seeking Nonsense, Looking for Trouble: Efficient Promotional-Infection Detection through Semantic Inconsistency SearchXiaojing Liao, Kan Yuan, XiaoFeng Wang, Zhongyu Pei et al.S&P 2016 · 41 citations
- Adversarial Semantic CollisionsCongzheng Song, Alexander M. Rush, Vitaly ShmatikovEMNLP 2020 · 31 citations
- How to Learn Klingon without a Dictionary: Detection and Measurement of Black Keywords Used by the Underground EconomyHao Yang, Xiulin Ma, Kun Du, Zhou Li et al.S&P 2017 · 48 citations
- Game of Missuggestions: Semantic Analysis of Search-Autocomplete ManipulationsPeng Wang, Xianghang Mi, Xiaojing Liao, XiaoFeng Wang et al.NDSS 2018 · 32 citations
