Hallucinating Certificates: Differential Testing of TLS Certificate Validation Using Generative Language Models
Muhammad Talha Paracha, Kyle Posluns, Kevin Borgolte, Martina Lindorfer, David Choffnes
Abstract
Certificate validation is a crucial step in Transport Layer Security (TLS), the de facto standard network security protocol. Prior research has shown that differentially testing TLS implementations with synthetic certificates can reveal critical security issues, such as accidentally accepting untrusted certificates. Leveraging known techniques, like random input mutations and program coverage guidance, prior work created corpora of synthetic certificates. By testing the certificates with multiple TLS libraries and comparing the validation outcomes, they discovered new bugs. However, they cannot generate the corresponding inputs efficiently, or they require to model the programs and their inputs in ways that scale poorly.
In this paper, we introduce a new approach, MLCerts, to generate synthetic certificates for differential testing that leverages generative language models to more extensively test software implementations. Recently, these models have become (in)famous for their applications in generating content, writing code, and conversing with users, as well as for "hallucinating" syntactically correct yet semantically nonsensical output. In this paper, we provide and leverage two novel insights: (a) TLS certificates can be expressed in natural-like language, namely in the X.509 standard that aids human readability, and (b) differential testing can benefit from hallucinated malformed test cases.
Using our approach MLCerts, we find significantly more distinct discrepancies between the five TLS implementations OpenSSL, Li-breSSL, GnuTLS, MbedTLS, and MatrixSSL than the state-of-the-art benchmark Transcert (+30%; 20 vs 26, out of a maximum possible of 30) and an order of magnitude more than the seminal work Frankencerts (+1,200%; 2 vs 26). Finally, we show that the diversity of MLCerts-generated certificates reveals a range of previously unobserved and interesting behavior with security implications.
• Software and its engineering → Software testing and debugging; • Security and privacy → Security protocols.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on9
- Large Language Models Are Zero-Shot Fuzzers: Fuzzing Deep-Learning Libraries via Large Language ModelsYinlin Deng, Chunqiu Steven Xia, Haoran Peng, Chenyuan Yang et al.ISSTA 2023 · 253 citations
- CodaMosa: Escaping Coverage Plateaus in Test Generation with Pre-trained Large Language ModelsCaroline Lemieux, Jeevana Priya Inala, Shuvendu K. Lahiri, Siddhartha SenICSE 2023 · 221 citations
- NEZHA: Efficient Domain-Independent Differential TestingTheofilos Petsios, Adrian Tang, Salvatore J. Stolfo, Angelos D. Keromytis et al.S&P 2017 · 132 citations
- Large Language Models are Edge-Case Generators: Crafting Unusual Programs for Fuzzing Deep Learning LibrariesYinlin Deng, Chunqiu Steven Xia, Chenyuan Yang, Shizhuo Dylan Zhang et al.ICSE 2024 · 85 citations
- SymCerts: Practical Symbolic Execution for Exposing Noncompliance in X.509 Certificate Validation ImplementationsSze Yiu Chau, Omar Chowdhury, Md. Endadul Hoque, Huangyi Ge et al.S&P 2017 · 67 citations
Related papers
- SADT: Syntax-Aware Differential Testing of Certificate Validation in SSL/TLS ImplementationsLili Quan, Qianyu Guo, Hongxu Chen, Xiaofei Xie et al.ASE 2020 · 10 citations
- SBDT: Search-Based Differential Testing of Certificate Parsers in SSL/TLS ImplementationsChu Chen, Pinghong Ren, Zhenhua Duan, Cong Tian et al.ISSTA 2023 · 13 citations
- RAT: Retrieval-Augmented Testing of Certificate Revocation List Parsers in TLS ImplementationsChu Chen, Qianxin Cheng, Pinghong Ren, Hairong Yu et al.FSE 2026
- Validating Network Protocol Parsers with Traceable RFC Document InterpretationMingwei Zheng, Danning Xie, Qingkai Shi, Chengpeng Wang et al.ISSTA 2025 · 4 citations
- HVLearn: Automated Black-Box Analysis of Hostname Verification in SSL/TLS ImplementationsSuphannee Sivakorn, George Argyros, Kexin Pei, Angelos D. Keromytis et al.S&P 2017 · 88 citations
