A Highly Scalable, Hybrid, Cross-Platform Timing Analysis Framework Providing Accurate Differential Throughput Estimation via Instruction-Level Tracing
Min-Yih Hsu, Felicitas Hetzelt, David Gens, Michael Maitland, Michael Franz
Abstract
Differential throughput estimation, i.e., predicting the performance impact of software changes, is critical when developing applications that rely on accurate timing bounds, such as automotive, avionic, or industrial control systems. However, developers often lack access to the target hardware to perform on-device measurements, and hence rely on instruction throughput estimation tools to evaluate performance impacts.
State-of-the-art techniques broadly fall into two categories: dynamic and static. Dynamic approaches emulate program execution using cycle-accurate microarchitectural simulators resulting in high precision at the cost of long turnaround times and convoluted setups. Static approaches reduce overhead by predicting cycle counts outside of a concrete runtime environment. However, they are limited by the lack of dynamic runtime information and mostly focus on predictions over single basic blocks which requires developers to manually construct critical instruction sequences.
We present MCAD 1 , a hybrid timing analysis framework that combines the advantages of dynamic and static approaches. Instead of relying on heavyweight cycle-accurate emulation, MCAD collects instruction traces along with dynamic runtime information from QEMU and streams them to a static throughput estimator. This allows developers to accurately estimate the performance impact of software changes for complete programs within minutes, reducing turnaround times by orders of magnitude compared to existing approaches with similar accuracy. Our evaluation shows that MCAD scales to real-world applications such as FFmpeg and Clang with millions of instructions, achieving < 3% geo. mean error * Both authors contributed equally to this research. † Now affiliated with Cerebras 1 The actual name has been redacted for anonymous review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 12cbbcb5-f23c-4db9-94f3-5fefb0cbdaaaBuilds on2
- Unleashing the hidden power of compiler optimization on binary code difference: an empirical studyXiaolei Ren, Michael Ho, Jiang Ming, Yu Lei et al.PLDI 2021 · 57 citations
- DeepBinDiff: Learning Program-Wide Code Representations for Binary DiffingYue Duan, Xuezixiang Li, Jinghan Wang, Heng YinNDSS 2020
Related papers
- Warping cache simulation of polyhedral programsCanberk Morelli, Jan ReinekePLDI 2022 · 3 citations
- Concorde: Fast and Accurate CPU Performance Modeling with Compositional Analytical-ML FusionArash Nasr-Esfahany, Mohammad Alizadeh, Victor Lee, Hanna Alam et al.ISCA 2025 · 1 citation
- AnICA: analyzing inconsistencies in microarchitectural code analyzersFabian Ritter, Sebastian HackOOPSLA 2022 · 2 citations
- CRISP: critical slice prefetchingHeiner Litz, Grant Ayers, Parthasarathy RanganathanASPLOS 2022 · 33 citations
- Enhancing Semantic-Aware Binary Diffing with High-Confidence Dynamic Instruction AlignmentChengfeng Ye, Anshunkang Zhou, Charles ZhangNDSS 2026 · 2 citations
