USENIX Security2021Top-tier venue
Phishpedia: A Hybrid Deep Learning Based Approach to Visually Identify Phishing Webpages
Yun Lin, Ruofan Liu, Dinil Mon Divakaran, Jun Yang Ng, Qing Zhou Chan, Yiwen Lu, Yuxuan Si, Fan Zhang, Jin Song Dong
Abstract
Recent years have seen the development of phishing detection and identification approaches to defend against phishing attacks. Phishing detection solutions often report binary results, i.e., phishing or not, without any explanation. In contrast, phishing identification approaches identify phishing webpages by visually comparing webpages with predefined legitimate references and report phishing along with its target brand, thereby having explainable results. However, there are technical challenges in visual analyses that limit existing solutions from being effective (with high accuracy) and efficient (with low runtime overhead), to be put to practical use. In this work, we design a hybrid deep learning system, Phishpedia, to address two prominent technical challenges in phishing identification, i.e., (i) accurate recognition of identity logos on webpage screenshots, and (ii) matching logo variants of the same brand. Phishpedia achieves both high accuracy and low runtime overhead. And very importantly, different from common approaches, Phishpedia does not require training on any phishing samples. We carry out extensive experiments using real phishing data; the results demonstrate that Phishpedia significantly outperforms baseline identification approaches (EMD, PhishZoo, and LogoSENSE) in accurately and efficiently identifying phishing pages. We also deployed Phishpedia with CertStream service and discovered 1,704 new real phishing websites within 30 days, significantly more than other solutions; moreover, 1,133 of them are not reported by any engines in VirusTotal.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7b8c8ddd-d5f7-4d5e-9528-220d314b2322Cited by top-tier papers35
- KnowPhish: Large Language Models Meet Multimodal Knowledge Graphs for Enhancing Reference-Based Phishing DetectionYuexin Li, Chengyu Huang, Shumin Deng, Mei Lin Lock et al.USENIX Security 2024 · 70 citations
- Catching Phishers By Their Bait: Investigating the Dutch Phishing Landscape through Phishing Kit DetectionHugo L. J. Bijmans, Tim M. Booij, Anneke Schwedersky, Aria Nedgabat et al.USENIX Security 2021 · 61 citations
- Less Defined Knowledge and More True Alarms: Reference-based Phishing Detection without a Pre-defined Reference ListRuofan Liu, Yun Lin, Xiwen Teoh, Gongshen Liu et al.USENIX Security 2024 · 39 citations
- TxPhishScope: Towards Detecting and Understanding Transaction-based Phishing on EthereumBowen He, Yuan Chen, Zhuo Chen, Xiaohui Hu et al.CCS 2023 · 38 citations
- "Get in Researchers; We're Measuring Reproducibility": A Reproducibility Study of Machine Learning Papers in Tier 1 Security ConferencesDaniel Olszewski, Allison Lu, Carson Stillman, Kevin Warren et al.CCS 2023 · 19 citations
Builds on7
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha et al.S&P 2016 · 3,275 citations
- Data Breaches, Phishing, or Malware?: Understanding the Risks of Stolen CredentialsKurt Thomas, Frank Li, Ali Zand, Jacob Barrett et al.CCS 2017 · 248 citations
- Cloak of Visibility: Detecting When Machines Browse a Different WebLuca Invernizzi, Kurt Thomas, Alexandros Kapravelos, Oxana Comanescu et al.S&P 2016 · 93 citations
- Logo-2K+: A Large-Scale Logo Dataset for Scalable Logo ClassificationJing Wang, Weiqing Min, Sujuan Hou, Shengnan Ma et al.AAAI 2020 · 58 citations
- Recovering fitness gradients for interprocedural Boolean flags in search-based testingYun Lin, Jun Sun, Gordon Fraser, Ziheng Xiu et al.ISSTA 2020 · 25 citations
Related papers
- Inferring Phishing Intention via Webpage Appearance and Dynamics: A Deep Vision Based ApproachRuofan Liu, Yun Lin, Xianglin Yang, Siang Hwee Ng et al.USENIX Security 2022
- Knowledge Expansion and Counterfactual Interaction for Reference-Based Phishing DetectionRuofan Liu, Yun Lin, Yifan Zhang, Penn Han Lee et al.USENIX Security 2023
- It Doesn't Look Like Anything to Me: Using Diffusion Model to Subvert Visual Phishing DetectorsQingying Hao, Nirav Diwan, Ying Yuan, Giovanni Apruzzese et al.USENIX Security 2024 · 11 citations
- Evaluating the Effectiveness and Robustness of Visual Similarity-based Phishing Detection ModelsFujiao Ji, Kiho Lee, Hyungjoon Koo, Wenhao You et al.USENIX Security 2025
- AI2TALE: An Innovative Information Theory-based Approach for Learning to Localize Phishing AttacksVan Nguyen, Tingmin Wu, Xingliang Yuan, Marthie Grobler et al.ICLR 2025
