Hallucinating Certificates: Differential Testing of TLS Certificate Validation Using Generative Language Models
Muhammad Talha Paracha, Kyle Posluns, Kevin Borgolte, Martina Lindorfer, David Choffnes
摘要
Certificate validation is a crucial step in Transport Layer Security (TLS), the de facto standard network security protocol. Prior research has shown that differentially testing TLS implementations with synthetic certificates can reveal critical security issues, such as accidentally accepting untrusted certificates. Leveraging known techniques, like random input mutations and program coverage guidance, prior work created corpora of synthetic certificates. By testing the certificates with multiple TLS libraries and comparing the validation outcomes, they discovered new bugs. However, they cannot generate the corresponding inputs efficiently, or they require to model the programs and their inputs in ways that scale poorly.
In this paper, we introduce a new approach, MLCerts, to generate synthetic certificates for differential testing that leverages generative language models to more extensively test software implementations. Recently, these models have become (in)famous for their applications in generating content, writing code, and conversing with users, as well as for "hallucinating" syntactically correct yet semantically nonsensical output. In this paper, we provide and leverage two novel insights: (a) TLS certificates can be expressed in natural-like language, namely in the X.509 standard that aids human readability, and (b) differential testing can benefit from hallucinated malformed test cases.
Using our approach MLCerts, we find significantly more distinct discrepancies between the five TLS implementations OpenSSL, Li-breSSL, GnuTLS, MbedTLS, and MatrixSSL than the state-of-the-art benchmark Transcert (+30%; 20 vs 26, out of a maximum possible of 30) and an order of magnitude more than the seminal work Frankencerts (+1,200%; 2 vs 26). Finally, we show that the diversity of MLCerts-generated certificates reveals a range of previously unobserved and interesting behavior with security implications.
• Software and its engineering → Software testing and debugging; • Security and privacy → Security protocols.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Large Language Models Are Zero-Shot Fuzzers: Fuzzing Deep-Learning Libraries via Large Language ModelsYinlin Deng, Chunqiu Steven Xia, Haoran Peng, Chenyuan Yang 等ISSTA 2023 · 被引用 253 次
- CodaMosa: Escaping Coverage Plateaus in Test Generation with Pre-trained Large Language ModelsCaroline Lemieux, Jeevana Priya Inala, Shuvendu K. Lahiri, Siddhartha SenICSE 2023 · 被引用 221 次
- NEZHA: Efficient Domain-Independent Differential TestingTheofilos Petsios, Adrian Tang, Salvatore J. Stolfo, Angelos D. Keromytis 等S&P 2017 · 被引用 132 次
- Large Language Models are Edge-Case Generators: Crafting Unusual Programs for Fuzzing Deep Learning LibrariesYinlin Deng, Chunqiu Steven Xia, Chenyuan Yang, Shizhuo Dylan Zhang 等ICSE 2024 · 被引用 85 次
- SymCerts: Practical Symbolic Execution for Exposing Noncompliance in X.509 Certificate Validation ImplementationsSze Yiu Chau, Omar Chowdhury, Md. Endadul Hoque, Huangyi Ge 等S&P 2017 · 被引用 67 次
相关 Paper
- SADT: Syntax-Aware Differential Testing of Certificate Validation in SSL/TLS ImplementationsLili Quan, Qianyu Guo, Hongxu Chen, Xiaofei Xie 等ASE 2020 · 被引用 10 次
- SBDT: Search-Based Differential Testing of Certificate Parsers in SSL/TLS ImplementationsChu Chen, Pinghong Ren, Zhenhua Duan, Cong Tian 等ISSTA 2023 · 被引用 13 次
- RAT: Retrieval-Augmented Testing of Certificate Revocation List Parsers in TLS ImplementationsChu Chen, Qianxin Cheng, Pinghong Ren, Hairong Yu 等FSE 2026
- Validating Network Protocol Parsers with Traceable RFC Document InterpretationMingwei Zheng, Danning Xie, Qingkai Shi, Chengpeng Wang 等ISSTA 2025 · 被引用 4 次
- HVLearn: Automated Black-Box Analysis of Hostname Verification in SSL/TLS ImplementationsSuphannee Sivakorn, George Argyros, Kexin Pei, Angelos D. Keromytis 等S&P 2017 · 被引用 88 次
