Finding Clues for Your Secrets: Semantics-Driven, Learning-Based Privacy Discovery in Mobile Apps
Yuhong Nan, Zhemin Yang, Xiaofeng Wang, Yuan Zhang, Donglai Zhu, Min Yang
摘要
A long-standing challenge in analyzing information leaks within mobile apps is to automatically identify the code operating on sensitive data. With all existing solutions relying on System APIs (e.g., IMEI, GPS location) or features of user interfaces (UI), the content from app servers, like user's Facebook profile, payment history, fall through the crack. Finding such content is important given the fact that most apps today are web applications, whose critical data are often on the server side. In the meantime, operations on the data within mobile apps are often hard to capture, since all server-side information is delivered to the app in the same way, sensitive or not. A unique observation of our research is that in modern apps, a program is essentially a semantics-rich documentation carrying meaningful program elements such as method names, variables and constants that reveal the sensitive data involved, even when the program is under moderate obfuscation. Leveraging this observation, we develop a novel semantics-driven solution for automatic discovery of sensitive user data, including those from the server side. Our approach utilizes natural language processing (NLP) to automatically locate the program elements (variables, methods, etc.) of interest, and then performs a learning-based program structure analysis to accurately identify those indeed carrying sensitive content. Using this new technique, we analyzed 445,668 popular apps, an unprecedented scale for this type of research. Our work brings to light the pervasiveness of information leaks, and the channels through which the leaks happen, including unintentional over-sharing across libraries and aggressive data acquisition behaviors. Further we found that many high-profile apps and libraries are involved in such leaks. Our findings contribute to a better understanding of the privacy risk in mobile apps and also highlight the importance of data protection in today's software composition.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper32
- Charting the Attack Surface of Trigger-Action IoT PlatformsQi Wang, Pubali Datta, Wei Yang, Si Liu 等CCS 2019 · 被引用 162 次
- CryptoGuard: High Precision Detection of Cryptographic Vulnerabilities in Massive-sized Java ProjectsSazzadur Rahaman, Ya Xiao, Sharmin Afrose, Fahad Shaon 等CCS 2019 · 被引用 159 次
- Demystifying Resource Management Risks in Emerging Mobile App-in-App EcosystemsHaoran Lu, Luyi Xing, Yue Xiao, Yifan Zhang 等CCS 2020 · 被引用 48 次
- Understanding Malicious Cross-library Data Harvesting on AndroidJice Wang, Yue Xiao, Xueqiang Wang, Yuhong Nan 等USENIX Security 2021 · 被引用 41 次
- Automated Detection of Password Leakage from Public GitHub RepositoriesRunhan Feng, Ziyang Yan, Shiyan Peng, Yuanyuan ZhangICSE 2022 · 被引用 36 次
它引用的顶会 Paper4
- Reliable Third-Party Library Detection in Android and its Security ApplicationsMichael Backes, Sven Bugiel, Erik DerrCCS 2016 · 被引用 345 次
- Automated Analysis of Privacy Requirements for Mobile AppsSebastian Zimmeck, Ziqi Wang, Lieyong Zou, Roger Iyengar 等NDSS 2017 · 被引用 255 次
- What Mobile Ads Know About Mobile UsersSooel Son, Daehyeok Kim, Vitaly ShmatikovNDSS 2016 · 被引用 101 次
- Free for All! Assessing User Data Exposure to Advertising Libraries on AndroidSoteris Demetriou, Whitney Merrill, Wei Yang, Aston Zhang 等NDSS 2016 · 被引用 95 次
相关 Paper
- Why Does Your Data Leak? Uncovering the Data Leakage in Cloud from Mobile AppsChaoshun Zuo, Zhiqiang Lin, Yinqian ZhangS&P 2019 · 被引用 123 次
- Dark Hazard: Learning-based, Large-Scale Discovery of Hidden Sensitive Operations in Android AppsXiaorui Pan, Xueqiang Wang, Yue Duan, XiaoFeng Wang 等NDSS 2017 · 被引用 69 次
- DocFlow: Extracting Taint Specifications from Software DocumentationMarcos Tileria, Jorge Blasco, Santanu Kumar DashICSE 2024 · 被引用 5 次
- AppBDS: LLM-Powered Description Synthesis for Sensitive Behaviors in Mobile AppsZichen Liu, Xusheng XiaoASE 2025
- Taintmini: Detecting Flow of Sensitive Data in Mini-Programs with Static Taint AnalysisChao Wang, Ronny Ko, Yue Zhang, Yuqing Yang 等ICSE 2023 · 被引用 36 次
