Automated Detection of Password Leakage from Public GitHub Repositories
Runhan Feng, Ziyang Yan, Shiyan Peng, Yuanyuan Zhang
摘要
The prosperity of the GitHub community has raised new concerns about data security in public repositories. Practitioners who manage authentication secrets such as textual passwords and API keys in the source code may accidentally leave these texts in the public repositories, resulting in secret leakage. If such leakage in the source code can be automatically detected in time, potential damage would be avoided. With existing approaches focusing on detecting secrets with distinctive formats (e.g., API keys, cryptographic keys in PEM format), textual passwords, which are ubiquitously used for authentication, fall through the crack. Given that textual passwords could be virtually any strings, a naive detection scheme based on regular expression performs poorly. This paper presents PassFinder, an automated approach to effectively detecting password leakage from public repositories that involve various programming languages on a large scale. PassFinder utilizes deep neural networks to unveil the intrinsic characteristics of textual passwords and understand the semantics of the code snippets that use textual passwords for authentication, i.e., the contextual information of the passwords in the source code. Using this new technique, we performed the first large-scale and longitudinal analysis of password leakage on GitHub. We inspected newly uploaded public code files on GitHub for 75 days and found that password leakage is pervasive, affecting over sixty thousand repositories. Our work contributes to a better understanding of password leakage on GitHub, and we believe our technique could promote the security of the open-source ecosystem. CCS CONCEPTS • Security and privacy → Authentication.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Unveiling Memorization in Code ModelsZhou Yang, Zhipeng Zhao, Chenyu Wang, Jieke Shi 等ICSE 2024 · 被引用 33 次
- Your Code Secret Belongs to Me: Neural Code Completion Tools Can Memorize Hard-Coded CredentialsYizhan Huang, Yichen Li, Weibin Wu, Jianping Zhang 等FSE 2024 · 被引用 22 次
- SoK: A Comprehensive Analysis and Evaluation of Docker Container Attack and Defense MechanismsMd. Sadun Haq, Thien Duc Nguyen, Ali Saman Tosun, Franziska Vollmer 等S&P 2024 · 被引用 17 次
- Don't Leak Your Keys: Understanding, Measuring, and Exploiting the AppSecret Leaks in Mini-ProgramsYue Zhang, Yuqing Yang, Zhiqiang LinCCS 2023 · 被引用 14 次
- More Haste, Less Speed: Cache Related Security Threats in Continuous Integration ServicesYacong Gu, Lingyun Ying, Huajun Chai, Yingyuan Pu 等S&P 2024 · 被引用 4 次
它引用的顶会 Paper12
- Fast, Lean, and Accurate: Modeling Password Guessability Using Neural NetworksWilliam Melicher, Blase Ur, Sean M. Segreti, Saranga Komanduri 等USENIX Security 2016 · 被引用 331 次
- Order Matters: Semantic-Aware Neural Networks for Binary Code Similarity DetectionZeping Yu, Rui Cao, Qiyi Tang, Sen Nie 等AAAI 2020 · 被引用 265 次
- Let's Go in for a Closer Look: Observing Passwords in Their Natural HabitatSarah Pearman, Jeremy Thomas, Pardis Emami Naeini, Hana Habib 等CCS 2017 · 被引用 168 次
- Big code != big vocabulary: open-vocabulary models for source codeRafael-Michael Karampatsis, Hlib Babii, Romain Robbes, Charles Sutton 等ICSE 2020 · 被引用 140 次
- How Bad Can It Git? Characterizing Secret Leakage in Public GitHub RepositoriesMichael Meli, Matthew R. McNiece, Bradley ReavesNDSS 2019 · 被引用 130 次
相关 Paper
- A Large-Scale Empirical Study of Secret Key Leakage in Hugging Face SpacesShaoxuan Yun, Yuchao Zhang, Zhikun Shi, Liu Wang 等ICSE 2026
- Hey, Your Secrets Leaked! Detecting and Characterizing Secret Leakage in the WildJiawei Zhou, Zidong Zhang, Lingyun Ying, Huajun Chai 等S&P 2025
- AssetHarvester: A Static Analysis Tool for Detecting Secret-Asset Pairs in Software ArtifactsSetu Kumar Basak, K. Virgil English, Ken Ogura, Vitesh Kambara 等ICSE 2025 · 被引用 1 次
- Leaky Apps: Large-scale Analysis of Secrets Distributed in Android and iOS AppsDavid Schmidt, Sebastian Schrittwieser, Edgar R. WeipplCCS 2025
- VulSim: Leveraging Similarity of Multi-Dimensional Neighbor Embeddings for Vulnerability DetectionSamiha Shimmi, Ashiqur Rahman, Mohan Gadde, Hamed Okhravi 等USENIX Security 2024 · 被引用 13 次
