Cloak of Visibility: Detecting When Machines Browse a Different Web
Luca Invernizzi, Kurt Thomas, Alexandros Kapravelos, Oxana Comanescu, Jean-Michel Picod, Elie Bursztein
Abstract
The contentious battle between web services and miscreants involved in blackhat search engine optimization and malicious advertisements has driven the underground to develop increasingly sophisticated techniques that hide the true nature of malicious sites. These web cloaking techniques hinder the effectiveness of security crawlers and potentially expose Internet users to harmful content. In this work, we study the spectrum of blackhat cloaking techniques that target browser, network, or contextual cues to detect organic visitors. As a starting point, we investigate the capabilities of ten prominent cloaking services marketed within the underground. This includes a first look at multiple IP blacklists that contain over 50 million addresses tied to the top five search engines and tens of anti-virus and security crawlers. We use our findings to develop an anti-cloaking system that detects split-view content returned to two or more distinct browsing profiles with an accuracy of 95.5% and a false positive rate of 0.9% when tested on a labeled dataset of 94,946 URLs. We apply our system to an unlabeled set of 135,577 search and advertisement URLs keyed on high-risk terms (e.g., luxury products, weight loss supplements) to characterize the prevalence of threats in the wild and expose variations in cloaking techniques across traffic sources. Our study provides the first broad perspective of cloaking as it affects Google Search and Google Ads and underscores the minimum capabilities necessary of security crawlers to bypass the state of the art in mobile, rDNS, and IP cloaking. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b485cf88-a3e4-48c2-a842-fa2337e3f902Cited by top-tier papers35
- Tranco: A Research-Oriented Top Sites Ranking Hardened Against ManipulationVictor Le Pochat, Tom van Goethem, Samaneh Tajalizadehkhoob, Maciej Korczynski et al.NDSS 2019 · 826 citations
- Data Breaches, Phishing, or Malware?: Understanding the Risks of Stolen CredentialsKurt Thomas, Frank Li, Ali Zand, Jacob Barrett et al.CCS 2017 · 248 citations
- Phishpedia: A Hybrid Deep Learning Based Approach to Visually Identify Phishing WebpagesYun Lin, Ruofan Liu, Dinil Mon Divakaran, Jun Yang Ng et al.USENIX Security 2021 · 164 citations
- PhishFarm: A Scalable Framework for Measuring the Effectiveness of Evasion Techniques against Browser Phishing BlacklistsAdam Oest, Yeganeh Safaei, Adam Doupé, Gail-Joon Ahn et al.S&P 2019 · 129 citations
- Mystique: Uncovering Information Leakage from Browser ExtensionsQuan Chen, Alexandros KapravelosCCS 2018 · 88 citations
Related papers
- Where are you taking me?Understanding Abusive Traffic Distribution SystemsJanos Szurdi, Meng Luo, Brian Kondracki, Nick Nikiforakis et al.WWW 2021 · 13 citations
- I'm SPARTACUS, No, I'm SPARTACUS: Proactively Protecting Users from Phishing by Intentionally Triggering Cloaking BehaviorPenghui Zhang, Zhibo Sun, Sukwha Kyung, Hans Walter Behrens et al.CCS 2022 · 18 citations
- CrawlPhish: Large-scale Analysis of Client-side Cloaking Techniques in PhishingPenghui Zhang, Adam Oest, Haehyun Cho, Zhibo Sun et al.S&P 2021 · 1 citation
- How to Learn Klingon without a Dictionary: Detection and Measurement of Black Keywords Used by the Underground EconomyHao Yang, Xiulin Ma, Kun Du, Zhou Li et al.S&P 2017 · 48 citations
- The Ever-Changing Labyrinth: A Large-Scale Analysis of Wildcard DNS Powered Blackhat SEOKun Du, Hao Yang, Zhou Li, Hai-Xin Duan et al.USENIX Security 2016 · 40 citations
