Secrets Unlocked: Evaluating LLMs for Secrets Detection in Android Apps
Marco Alecci, Jordan Samhi, Tegawendé F. Bissyandé, Jacques Klein
Abstract
Mobile apps frequently embed sensitive secrets, such as API keys, access tokens, client secrets, and private keys that support internal functionality or enable integration with external systems and third-party services. Developers frequently embed these secrets into Android apps, which allows attackers to extract them through reverse engineering. Once exposed, attackers can exploit them to access sensitive data, manipulate resources, or abuse APIs, resulting in severe security and potential financial risks.
In this paper, we present the first large-scale empirical evidence that off-the-shelf large language models (LLMs) can automatically identify secrets in Android apps without any domain-specific prior knowledge, thereby substantially lowering the barrier for attackers. On a benchmark of 5135 Android apps from prior work, LLMs rediscovered 93% of previously known secrets and identified 4361 additional valid credentials (+195%). Extending our analysis to 50 000 Google Play apps collected between August and October 2025, we conducted the largest-scale study to date on secret detection in Android apps, identifying secrets in 17 590 apps (35%). Among the 18 908 detected secrets, 1802 remained active at discovery, including, among others, critical credentials such as Stripe payment keys, OpenAI API keys, and GitHub personal access tokens. We responsibly contacted all the affected developers, of whom 170 confirmed the issues and updated their apps accordingly.
Our findings empirically demonstrate the reality of vibe hacking: anyone can now leverage publicly accessible AI models to perform complex offensive security tasks with minimal expertise.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0ba5a541-27cd-48cf-8d92-7d4a80b7b553Builds on6
- How Bad Can It Git? Characterizing Secret Leakage in Public GitHub RepositoriesMichael Meli, Matthew R. McNiece, Bradley ReavesNDSS 2019 · 130 citations
- Why Does Your Data Leak? Uncovering the Data Leakage in Cloud from Mobile AppsChaoshun Zuo, Zhiqiang Lin, Yinqian ZhangS&P 2019 · 123 citations
- Automated Detection of Password Leakage from Public GitHub RepositoriesRunhan Feng, Ziyang Yan, Shiyan Peng, Yuanyuan ZhangICSE 2022 · 36 citations
- Leaky Apps: Large-scale Analysis of Secrets Distributed in Android and iOS AppsDavid Schmidt, Sebastian Schrittwieser, Edgar R. WeipplCCS 2025
- From Large to Mammoth: A Comparative Evaluation of Large Language Models in Vulnerability DetectionJie Lin, David MohaisenNDSS 2025
Related papers
- A Large-Scale Empirical Study of Secret Key Leakage in Hugging Face SpacesShaoxuan Yun, Yuchao Zhang, Zhikun Shi, Liu Wang et al.ICSE 2026
- Hey, Your Secrets Leaked! Detecting and Characterizing Secret Leakage in the WildJiawei Zhou, Zidong Zhang, Lingyun Ying, Huajun Chai et al.S&P 2025
- Beyond Static Pattern Matching? Rethinking Automatic Cryptographic API Misuse Detection in the Era of LLMsYifan Xia, Zichen Xie, Peiyu Liu, Kangjie Lu et al.ISSTA 2025 · 2 citations
- Decoding Secret Memorization in Code LLMs Through Token-Level CharacterizationYuqing Nie, Chong Wang, Kailong Wang, Guoai Xu et al.ICSE 2025 · 10 citations
- How Safe Is Your Screen? Understanding and Detecting Privacy Leaks in Sensitive ActivitiesYoungseok Kim, Sungho Lee, Sungjae HwangISSTA 2026
