Checked-In Secret Detection: Strings Are All You Need
Zhengdong Huang, Kevin Li, Jinqiu Yang, Yepang Liu, Lili Wei
Abstract
Hardcoded secrets in source code pose critical security vulnerabilities which can be easily exploited by malicious adversaries. Existing regex-based detection approaches suffer from fundamental limitations, as secrets often lack identifiable patterns, resulting in poor precision and recall. Recent studies have explored context-aware detection methods, as surrounding code can reveal the purpose of candidate strings. However, these methods confront three key challenges: (1) obfuscation robustness where models over-rely on easily obfuscated identifiers, (2) cross-language generalization difficulties due to uneven training data distribution, and (3) lengthy and noisy context that introduces excessive irrelevant tokens and slows inference.
We observe that strings serve as a critical information source for code semantics, offering superior contextual density, obfuscation robustness, and language independence. Based on this insight, we propose StringGroup, a new context extraction algorithm that mines strings surrounding potential secrets. By introducing a relatively simple modification to existing patterns that narrows the analysis specifically to string literals, the method achieves significant gains. With only 33.2% of the original context, it preserves over 80% of semantic information and significantly improves the signal-to-noise ratio for secret detection. We further design a context-aware secret detection tool, Secretron, based on StringGroup methods and Transformer model. Evaluation on the SecretBench dataset demonstrates high accuracy with 98.74% F1score and strong robustness under obfuscation and cross-language scenarios, outperforming state-of-the-art LLM-based baselines. We deploy our tool in real-world environments and successfully detect 48 previously unknown secret keys from 26 applications, demonstrating the practical effectiveness of our approach.
CCS Concepts: • Security and privacy → Software and application security.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c066c337-9e28-49ca-a8aa-ce25d6effe92Builds on7
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
- How Bad Can It Git? Characterizing Secret Leakage in Public GitHub RepositoriesMichael Meli, Matthew R. McNiece, Bradley ReavesNDSS 2019 · 130 citations
- Why Does Your Data Leak? Uncovering the Data Leakage in Cloud from Mobile AppsChaoshun Zuo, Zhiqiang Lin, Yinqian ZhangS&P 2019 · 123 citations
- Automated Detection of Password Leakage from Public GitHub RepositoriesRunhan Feng, Ziyang Yan, Shiyan Peng, Yuanyuan ZhangICSE 2022 · 36 citations
Related papers
- Hey, Your Secrets Leaked! Detecting and Characterizing Secret Leakage in the WildJiawei Zhou, Zidong Zhang, Lingyun Ying, Huajun Chai et al.S&P 2025
- CTX-Coder: Cross-Attention Architectures Empower LLMs for Long-Context Vulnerability DetectionJujie Wang, Kangfeng Zheng, Bin Wu, Chunhua Wu et al.AAAI 2026
- Argus: A Multi-agent Sensitive Information Leakage Detection Framework Based on Hierarchical Reference RelationshipsWang Bin, Hui Li, Liyang Zhang, Qijia Zhuang et al.ICSE 2026
- AssetHarvester: A Static Analysis Tool for Detecting Secret-Asset Pairs in Software ArtifactsSetu Kumar Basak, K. Virgil English, Ken Ogura, Vitesh Kambara et al.ICSE 2025 · 1 citation
- Decoding Secret Memorization in Code LLMs Through Token-Level CharacterizationYuqing Nie, Chong Wang, Kailong Wang, Guoai Xu et al.ICSE 2025 · 10 citations
