Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm
Yan Pang, Wenlong Meng, Xiaojing Liao, Tianhao Wang
摘要
With the rapid development of large language models, the potential threat of their malicious use, particularly in generating phishing content, is becoming increasingly prevalent. Leveraging the capabilities of LLMs, malicious users can synthesize phishing emails that are free from spelling mistakes and other easily detectable features. Furthermore, such models can generate topic-specific phishing messages, tailoring content to the target domain and increasing the likelihood of success. Detecting such content remains a significant challenge, as LLM-generated phishing emails often lack clear or distinguishable linguistic features. As a result, most existing semantic-level detection approaches struggle to identify them reliably. While certain LLM-based detection methods have shown promise, they suffer from high computational costs and are constrained by the performance of the underlying language model, making them impractical for large-scale deployment. In this work, we aim to address this issue. We propose Paladin, which embeds trigger-tag associations into vanilla LLM using various insertion strategies, creating them into instrumented LLMs. When an instrumented LLM generates content related to phishing, it will automatically include detectable tags, enabling easier identification. Based on the design on implicit and explicit triggers and tags, we consider four distinct scenarios in our work. We evaluate our method from three key perspectives: stealthiness, effectiveness, and robustness, and compare it with existing baseline methods. Experimental results show that our method outperforms the baselines, achieving over 90% detection accuracy across all scenarios. We share our code at [1].
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper48
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen 等ICLR 2024 · 被引用 1,104 次
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou 等ICML 2024 · 被引用 1,031 次
相关 Paper
- A Large-Scale Study of Personalized Phishing using Large Language ModelsStefan Czybik, Anne Josiane Kouam, Peter Heubl, Jan Magnus Nold 等USENIX Security 2026 · 被引用 1 次
- From Chatbots to Phishbots?: Phishing Scam Generation in Commercial Large Language ModelsSayak Saha Roy, Poojitha Thota, Krishna Vamsi Naragam, Shirin NilizadehS&P 2024 · 被引用 57 次
- PiMRef: Deducing Ever-evolving Spear-phishing Emails with Knowledge Base InvariantsRuofan Liu, Yun Lin, Yuxin Wang, Xiwen Teoh 等CCS 2026 · 被引用 4 次
- TrojLLM: A Black-box Trojan Prompt Attack on Large Language ModelsJiaqi Xue, Mengxin Zheng, Ting Hua, Yilin Shen 等NeurIPS 2023 · 被引用 63 次
- Efficient Detection of Toxic Prompts in Large Language ModelsYi Liu, Junzhe Yu, Huijia Sun, Ling Shi 等ASE 2024 · 被引用 6 次
