CrossFit: Demystifying VM Callback Bugs in Interpreters
Chibin Zhang, Qiang Liu, Mathias Payer
Abstract
Scripting languages like Python, Ruby, or PHP are integral to modern software development. Despite security measures like memory safety and sandboxing, vulnerabilities within these engines can lead to critical issues such as remote code execution or sandbox escapes. A particularly pervasive class of vulnerabilities is callback bugs, which occur when user-defined callbacks violate runtime invariants, such as freeing an object still in use (can be reached through live pointers) or modifying a data structure during traversal. These violations can result in severe consequences, including use-after-free, null-pointer dereferences, and type confusion, often leading to crashes, memory corruption, or even exploitable vulnerabilities. Detecting callback bugs remains challenging due to a lack of general understanding, as they have not been formally characterized or systematically studied. As such, existing tools lack the ability to ( 1) establish clear links between script-side callbacks and their native-side invokers, and (2) generate scripts that systematically satisfy preconditions required to trigger these bugs.
We propose CrossFit, a novel 2-tier approach combining static analysis and targeted fuzzing to systematically discover callback bugs. CrossFit first establishes links between script-side callbacks and their native-side invokers through context link analysis, enabling targeted exploration of high-risk code paths. It then generates proof-of-concept scripts with custom classes and magic methods, introducing side-effect operations to violate runtime invariants. Our evaluation shows that CrossFit effectively outperforms existing tools by up to 12.04% in terms of callsite coverage (i.e., potential sites where callback bugs may occur). We also identified 20 new bugs in Python, Ruby, and PHP, many of which are severe memory corruptions. Moreover, we provide a comprehensive benchmark totaling 150 proof-of-concepts to improve interpreter security.
CCS Concepts: • Software and its engineering → Software testing and debugging.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 617e954b-1fcb-49e2-9032-8198f3bd56feBuilds on15
- Skyfire: Data-Driven Seed Generation for FuzzingJunjie Wang, Bihuan Chen, Lei Wei, Yang LiuS&P 2017 · 382 citations
- NAUTILUS: Fishing for Deep Bugs with GrammarsCornelius Aschermann, Tommaso Frassetto, Thorsten Holz, Patrick Jauernig et al.NDSS 2019 · 291 citations
- CodeAlchemist: Semantics-Aware Code Generation to Find Vulnerabilities in JavaScript EnginesHyungSeok Han, DongHyeon Oh, Sang Kil ChaNDSS 2019 · 178 citations
- Fuzzing JavaScript Engines with Aspect-preserving MutationSoyeon Park, Wen Xu, Insu Yun, Daehee Jang et al.S&P 2020 · 126 citations
- One Engine to Fuzz 'em All: Generic Language Processor Testing with Semantic ValidationYongheng Chen, Rui Zhong, Hong Hu, Hangfan Zhang et al.S&P 2021 · 70 citations
Related papers
- PyRTFuzz: Detecting Bugs in Python Runtimes via Two-Level Collaborative FuzzingWen Li, Haoran Yang, Xiapu Luo, Long Cheng et al.CCS 2023 · 14 citations
- XSSky: Detecting XSS Vulnerabilities through Local Path-Persistent FuzzingYoukun Shi, Yuan Zhang, Tianhao Bai, Feng Xue et al.USENIX Security 2025
- JIT-Picking: Differential Fuzzing of JavaScript EnginesLukas Bernhard, Tobias Scharnowski, Moritz Schloegel, Tim Blazytko et al.CCS 2022 · 42 citations
- PyXray: Practical Cross-Language Call Graph Construction through Object Layout AnalysisGeorgios Alexopoulos, Thodoris Sotiropoulos, Georgios Gousios, Zhendong Su et al.ICSE 2026
- Fuzzing the PHP Interpreter via Dataflow FusionYuancheng Jiang, Chuqi Zhang, Bonan Ruan, Jiahao Liu et al.USENIX Security 2025
