PhishLang: A Real-Time, Fully Client-Side Phishing Detection Framework Using MobileBERT
Sayak Saha Roy, Shirin Nilizadeh
摘要
We present PhishLang, the first fully client-side anti-phishing framework implemented as a Chromium-based browser extension. PhishLang enables real-time, on-device detection of phishing websites by utilizing a lightweight language model (MobileBERT). Unlike traditional heuristic or static feature-based models that struggle with evasive threats, and deep learning approaches that are too resource-intensive for client-side use, PhishLang analyzes the contextual structure of a page’s source code, achieving detection performance on par with several state-of-the-art models while consuming up to 7 times less memory than comparable architectures. Over a 3.5-month period, we deployed the framework in real-time, successfully identifying approximately 26k phishing URLs, many of which were undetected by popular antiphishing blocklists, thus demonstrating PhishLang's potential to aid current detection measures. On the other hand, the browser extension outperformed several anti-phishing tools, detecting over 91% of the threats during zero-day. PhishLang also showed strong adversarial robustness, resisting 16 categories of realistic problem space evasions through a combination of parser-level defenses and adversarial retraining. To aid both end-users and the research community, we have open-sourced both the PhishLang framework and the browser extension.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper12
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited DevicesZhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu 等ACL 2020 · 被引用 660 次
- Learning and Evaluating Contextual Embedding of Source CodeAditya Kanade, Petros Maniatis, Gogul Balakrishnan, Kensen ShiICML 2020 · 被引用 438 次
- Semantics-Aware BERT for Language UnderstandingZhuosheng Zhang, Yuwei Wu, Hai Zhao, Zuchao Li 等AAAI 2020 · 被引用 396 次
- Intriguing Properties of Adversarial ML Attacks in the Problem SpaceFabio Pierazzi, Feargus Pendlebury, Jacopo Cortellazzi, Lorenzo CavallaroS&P 2020 · 被引用 334 次
相关 Paper
- A Good Fishman Knows All the Angles: A Critical Evaluation of Google's Phishing Page ClassifierChangqing Miao, Jianan Feng, Wei You, Wenchang Shi 等CCS 2023 · 被引用 2 次
- From Chatbots to Phishbots?: Phishing Scam Generation in Commercial Large Language ModelsSayak Saha Roy, Poojitha Thota, Krishna Vamsi Naragam, Shirin NilizadehS&P 2024 · 被引用 57 次
- PhishFarm: A Scalable Framework for Measuring the Effectiveness of Evasion Techniques against Browser Phishing BlacklistsAdam Oest, Yeganeh Safaei, Adam Doupé, Gail-Joon Ahn 等S&P 2019 · 被引用 129 次
- SoK: PHILTER: Uncovering Security and Functional Gaps in AI-based Phishing Website Detection Literature via an LLM-based Reasoning FrameworkMahbub Alam, Muhammad Lutfor Rahman, Sonjoy Kumar Paul, Amy W. Hays 等USENIX Security 2026
- PhishTime: Continuous Longitudinal Measurement of the Effectiveness of Anti-phishing BlacklistsAdam Oest, Yeganeh Safaei, Penghui Zhang, Brad Wardman 等USENIX Security 2020
