A Machine Learning Approach to Prevent Malicious Calls over Telephony Networks
Huichen Li, Xiaojun Xu, Chang Liu, Teng Ren, Kun Wu, Xuezhi Cao, Weinan Zhang, Yong Yu, Dawn Song
Abstract
Malicious calls, i.e., telephony spams and scams, have been a long-standing challenging issue that causes billions of dollars of annual financial loss worldwide. This work presents the first machine learning-based solution without relying on any particular assumptions on the underlying telephony network infrastructures. The main challenge of this decade-long problem is that it is unclear how to construct effective features without the access to the telephony networks' infrastructures. We solve this problem by combining several innovations. We first develop a TouchPal user interface on top of a mobile App to allow users tagging malicious calls. This allows us to maintain a large-scale call log database. We then conduct a measurement study over three months of call logs, including 9 billion records. We design 29 features based on the results, so that machine learning algorithms can be used to predict malicious calls. We extensively evaluate different state-of-the-art machine learning approaches using the proposed features, and the results show that the best approach can reduce up to 90% unblocked malicious calls while maintaining a precision over 99.99% on the benign call traffic. The results also show the models are efficient to implement without incurring a significant latency overhead. We also conduct ablation analysis, which reveals that using 10 out of the 29 features can reach a performance comparable to using all features.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1286f86e-2888-438f-b987-40da40bc2bdcCited by top-tier papers11
- Privacy Attacks to the 4G and 5G Cellular Paging Protocols Using Side Channel InformationSyed Rafiul Hussain, Mitziu Echeverria, Omar Chowdhury, Ninghui Li et al.NDSS 2019 · 160 citations
- Throwing Darts in the Dark? Detecting Bots with Limited Data using Neural Data AugmentationSteve T. K. Jan, Qingying Hao, Tianrui Hu, Jiameng Pu et al.S&P 2020 · 88 citations
- DeepSQLi: deep semantic learning for testing SQL injectionMuyang Liu, Ke Li, Tao ChenISSTA 2020 · 47 citations
- Lies in the Air: Characterizing Fake-base-station Spam Ecosystem in ChinaYiming Zhang, Baojun Liu, Chaoyi Lu, Zhou Li et al.CCS 2020 · 43 citations
- "Get in Researchers; We're Measuring Reproducibility": A Reproducibility Study of Machine Learning Papers in Tier 1 Security ConferencesDaniel Olszewski, Allison Lu, Carson Stillman, Kevin Warren et al.CCS 2023 · 19 citations
Builds on1
Related papers
- Towards Measuring the Effectiveness of Telephony BlacklistsSharbani Pandit, Roberto Perdisci, Mustaque Ahamad, Payas GuptaNDSS 2018 · 36 citations
- Combating Robocalls with Phone Virtual Assistant Mediated InteractionSharbani Pandit, Krishanu Sarker, Roberto Perdisci, Mustaque Ahamad et al.USENIX Security 2023
- When Scammers Talk Back: Understanding Potentially Unwanted Calls via LLM-Based InteractionZhuoer Lyu, Marzieh Bitaab, Shuyi Huang, Alireza Karimi et al.CCS 2026
- UCBlocker: Unwanted Call Blocking Using Anonymous AuthenticationChanglai Du, Hexuan Yu, Yang Xiao, Y. Thomas Hou et al.USENIX Security 2023
- Understanding and Detecting International Revenue Share FraudMerve Sahin, Aurélien FrancillonNDSS 2021
