HFuzzer: Testing Large Language Models for Package Hallucinations via Phrase-based Fuzzing
Yukai Zhao, Menghan Wu, Xing Hu, Xin Xia
Abstract
Large Language Models (LLMs) are widely used for code generation, but they face critical security risks when applied to practical production due to package hallucinations, in which LLMs recommend non-existent packages. These hallucinations can be exploited in software supply chain attacks, where malicious attackers exploit them to register harmful packages. It is critical to test LLMs for package hallucinations to mitigate package hallucinations and defend against potential attacks. Although researchers have proposed testing frameworks for fact-conflicting hallucinations in natural language generation, there is a lack of research on package hallucinations. To fill this gap, we propose HFuzzer, a novel phrase-based fuzzing framework to test LLMs for package hallucinations. HFuzzer adopts fuzzing technology and guides the model to infer a wider range of reasonable information based on phrases, thereby generating enough and diverse coding tasks. Furthermore, HFuzzer extracts phrases from package information or coding tasks to ensure the relevance of phrases and code, thereby improving the relevance of generated tasks and code. We evaluate HFuzzer on multiple LLMs and find that it triggers package hallucinations across all selected models. Compared to the mutational fuzzing framework, HFuzzer identifies 2.60× more unique hallucinated packages and generates more diverse tasks. Additionally, when testing the model GPT-4o, HFuzzer finds 46 unique hallucinated packages. Further analysis reveals that for GPT-4o, LLMs exhibit package hallucinations not only during code generation but also when assisting with environment configuration.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ab0f9d43-9435-4040-8af6-4bb3fd1ee21bCited by top-tier papers1
Ask how each one uses itBuilds on27
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- INSIDE: LLMs' Internal States Retain the Power of Hallucination DetectionChao Chen, Kai Liu, Ze Chen, Yi Gu et al.ICLR 2024 · 281 citations
- Automated Repair of Programs from Large Language ModelsZhiyu Fan, Xiang Gao, Martin Mirchev, Abhik Roychoudhury et al.ICSE 2023 · 213 citations
- Fuzz4All: Universal Fuzzing with Large Language ModelsChunqiu Steven Xia, Matteo Paltenghi, Jia Le Tian, Michael Pradel et al.ICSE 2024 · 155 citations
- Automated Program Repair via Conversation: Fixing 162 out of 337 Bugs for $0.42 Each using ChatGPTChunqiu Steven Xia, Lingming ZhangISSTA 2024 · 105 citations
Related papers
- We Have a Package for You! A Comprehensive Analysis of Package Hallucinations by Code Generating LLMsJoseph Spracklen, Raveen Wijewickrama, A. H. M. Nazmus Sakib, Anindya Maiti et al.USENIX Security 2025
- PackMonitor: Towards Zero Package Hallucinations via Decoding-Time MonitoringXiting Liu, Yuetong Liu, Yitong Zhang, Jia Li et al.ISSTA 2026
- Generating Precise Format Specification for Network Protocols Through Adversarial LLM InteractionsHengdi Ye, Bing Shui, Jielun Wu, Yufan Zhou et al.USENIX Security 2026
- CodeHalu: Investigating Code Hallucinations in LLMs via Execution-based VerificationYuchen Tian, Weixiang Yan, Qian Yang, Xuandong Zhao et al.AAAI 2025 · 41 citations
- LLM Hallucinations in Practical Code Generation: Phenomena, Mechanism, and MitigationZiyao Zhang, Chong Wang, Yanlin Wang, Ensheng Shi et al.ISSTA 2025 · 53 citations
