Towards a Lightweight, Hybrid Approach for Detecting DOM XSS Vulnerabilities with Machine Learning
William Melicher, Clement Fung, Lujo Bauer, Limin Jia
摘要
Client-side cross-site scripting (DOM XSS) vulnerabilities in web applications are common, hard to identify, and difficult to prevent. Taint tracking is the most promising approach for detecting DOM XSS with high precision and recall, but is too computationally expensive for many practical uses. We investigate whether machine learning (ML) classifiers can replace or augment taint tracking when detecting DOM XSS vulnerabilities. Through a large-scale web crawl, we collect over 18 billion JavaScript functions and use taint tracking to label over 180,000 functions as potentially vulnerable. With this data, we train a deep neural network (DNN) to analyze a JavaScript function and predict if it is vulnerable to DOM XSS. We experiment with a range of hyperparameters and present a low-latency, high-recall classifier that could serve as a pre-filter to taint tracking, reducing the cost of stand-alone taint tracking by 3.43× while detecting 94.5% of unique vulnerabilities. We argue that this combination of a DNN and taint tracking is efficient enough for a range of use cases for which taint tracking by itself is not, including in-browser run-time DOM XSS detection and analyzing large codebases. CCS CONCEPTS • Security and privacy → Web application security; • Information systems → World Wide Web; • Computing methodologies → Machine learning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- An Empirical Study of Knowledge Distillation for Code Understanding TasksRuiqi Wang, Zezhou Yang, Cuiyun Gao, Xin Xia 等ICSE 2026 · 被引用 1 次
- DOM-XSS Detection via Webpage Interaction Fuzzing and URL Component SynthesisNuno Sabino, Darion Cassel, Rui Abreu, Pedro Adão 等NDSS 2026 · 被引用 1 次
- Probe the Proto: Measuring Client-Side Prototype Pollution Vulnerabilities of One Million Real-world WebsitesZifeng Kang, Song Li, Yinzhi CaoNDSS 2022
它引用的顶会 Paper4
- Fast, Lean, and Accurate: Modeling Password Guessability Using Neural NetworksWilliam Melicher, Blase Ur, Sean M. Segreti, Saranga Komanduri 等USENIX Security 2016 · 被引用 331 次
- Content Security Problems?: Evaluating the Effectiveness of Content Security Policy in the WildStefano Calzavara, Alvise Rabitti, Michele BugliesiCCS 2016 · 被引用 71 次
- Neutaint: Efficient Dynamic Taint Analysis with Neural NetworksDongdong She, Yizheng Chen, Abhishek Shah, Baishakhi Ray 等S&P 2020 · 被引用 54 次
- VulDeePecker: A Deep Learning-Based System for Vulnerability DetectionZhen Li, Deqing Zou, Shouhuai Xu, Xinyu Ou 等NDSS 2018
相关 Paper
- Riding out DOMsday: Towards Detecting and Preventing DOM Cross-Site ScriptingWilliam Melicher, Anupam Das, Mahmood Sharif, Lujo Bauer 等NDSS 2018 · 被引用 84 次
- Don't Trust The Locals: Investigating the Prevalence of Persistent Client-Side Cross-Site Scripting in the WildMarius Steffens, Christian Rossow, Martin Johns, Ben StockNDSS 2019 · 被引用 84 次
- Dancer in the Dark: Synthesizing and Evaluating Polyglots for Blind Cross-Site ScriptingRobin Kirchner, Jonas Möller, Marius Musch, David Klein 等USENIX Security 2024 · 被引用 9 次
- Black Widow: Blackbox Data-driven Web ScanningBenjamin Eriksson, Giancarlo Pellegrino, Andrei SabelfeldS&P 2021 · 被引用 65 次
- In the DOM We Trust: Exploring the Hidden Dangers of Reading from the DOM on the WebJan Drescher, Sepehr Mirzaei, Soheil Khodayari, David Klein 等CCS 2025
