SoK: PHILTER: Uncovering Security and Functional Gaps in AI-based Phishing Website Detection Literature via an LLM-based Reasoning Framework
Mahbub Alam, Muhammad Lutfor Rahman, Sonjoy Kumar Paul, Amy W. Hays, Aftab Hussain, Md Imanul Huq, Nitesh Saxena
摘要
Phishing websites remain a dominant enabler of cybercrime. In response, many academic AI-based phishing website detection methods have been developed, often inspiring the design of real-world systems. Although most studies report high accuracy, it remains unclear whether they meet real-world requirements such as resilience to evolving phishing tactics, robustness on diverse benign pages, interpretability, and privacy. We present PHILTER (PHishing detection literature Inspection via LLMs and Targeted Expert Review), a scalable framework for qualitatively assessing phishing website detection studies across four functionality and three security metrics. PHILTER leverages LLMs to extract evidence and draft rationales, which experts then validate and use to produce the final assessment. Applying it to 55 academic approaches reveals systemic gaps. No study fulfills all functionality and security requirements. None show evidence of effectively addressing diverse phishing tactics. Most approaches struggle to preserve privacy and adapt to evolving attacker strategies, and many risk elevated false alarms in practice due to limited testing on diverse benign pages. We also introduce a taxonomy of detection strategies (feature-based, similarity-based, identity-based, and hybrid) that highlights design trade-offs and helps explain these shortcomings. Our study reveals that accuracy-driven evaluation overlooks blind spots that undermine practical effectiveness and exposes a key open challenge: achieving high accuracy while fulfilling all functionality and security requirements. We provide actionable recommendations to guide the design of future defenses that pursue this simultaneous goal against evolving and adaptive phishing campaigns.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Phishpedia: A Hybrid Deep Learning Based Approach to Visually Identify Phishing WebpagesYun Lin, Ruofan Liu, Dinil Mon Divakaran, Jun Yang Ng 等USENIX Security 2021 · 被引用 164 次
- KnowPhish: Large Language Models Meet Multimodal Knowledge Graphs for Enhancing Reference-Based Phishing DetectionYuexin Li, Chengyu Huang, Shumin Deng, Mei Lin Lock 等USENIX Security 2024 · 被引用 70 次
- Less Defined Knowledge and More True Alarms: Reference-based Phishing Detection without a Pre-defined Reference ListRuofan Liu, Yun Lin, Xiwen Teoh, Gongshen Liu 等USENIX Security 2024 · 被引用 39 次
- PhishPrint: Evading Phishing Detection Crawlers by Prior ProfilingBhupendra Acharya, Phani VadrevuUSENIX Security 2021 · 被引用 37 次
- It Doesn't Look Like Anything to Me: Using Diffusion Model to Subvert Visual Phishing DetectorsQingying Hao, Nirav Diwan, Ying Yuan, Giovanni Apruzzese 等USENIX Security 2024 · 被引用 11 次
相关 Paper
- Evaluating the Effectiveness and Robustness of Visual Similarity-based Phishing Detection ModelsFujiao Ji, Kiho Lee, Hyungjoon Koo, Wenhao You 等USENIX Security 2025
- PhishDecloaker: Detecting CAPTCHA-cloaked Phishing Websites via Hybrid Vision-based Interactive ModelsXiwen Teoh, Yun Lin, Ruofan Liu, Zhiyong Huang 等USENIX Security 2024 · 被引用 13 次
- PiMRef: Deducing Ever-evolving Spear-phishing Emails with Knowledge Base InvariantsRuofan Liu, Yun Lin, Yuxin Wang, Xiwen Teoh 等CCS 2026 · 被引用 4 次
- VisualPhishNet: Zero-Day Phishing Website Detection by Visual SimilaritySahar Abdelnabi, Katharina Krombholz, Mario FritzCCS 2020 · 被引用 5 次
- PhishTime: Continuous Longitudinal Measurement of the Effectiveness of Anti-phishing BlacklistsAdam Oest, Yeganeh Safaei, Penghui Zhang, Brad Wardman 等USENIX Security 2020
