InstrSem: Automatically and Generically Inferring Semantics of (Undocumented) CPU Instructions
Lorenz Hetterich, Fabian Thomas, Tristan Hornetz, Michael Schwarz
摘要
Modern CPUs implement complex Instruction Set Architectures (ISAs), yet machine-readable semantics are often incomplete. Worse, many CPUs support undocumented instructions, i.e., bitstrings that execute on hardware but are absent from specifications, leading to potential security vulnerabilities. In this paper, we present InstrSem, an ISA-agnostic, modular, fully automated approach to infer instruction semantics from execution behavior alone and provide semantics that are understandable by both, humans and machines. Starting from a raw encoding, InstrSem executes it under systematically varied architectural states and synthesizes compact mathematical functions that explain every changed state component. By mutating encoding bits and correlating induced behavioral changes with bit positions, InstrSem then generalizes from a single encoding to a full instruction, recovering register and immediate fields. In contrast to prior work focusing on a single ISA, InstrSem is generic. It requires only a lightweight ISA model and a per-architecture user-space runner and supports fixed- and variable-length encodings (RISC and CISC), memory accesses, and conditional behavior. We evaluate InstrSem on RV64I, AArch64, and LA64, and additionally showcase CISC applicability on a Logitech macro language and partial x86-64. InstrSem automatically recovers correct semantics for over 97.81 % of the RV64I base instruction set, and 136 instructions covering 1 009 055 744 instruction encodings within 77 h for the LA64 instruction set. InstrSem discovers undocumented vector instructions, inconsistencies between QEMU and Loongson hardware, and instructions that crash QEMU. InstrSem enables scalable recovery of instruction semantics, substantially automating reverse engineering across commodity and niche targets and strengthening the foundations for emulation, verification, and security analysis. With minimal requirements to support new architectures, its modular design, and human-readable output, InstrSem can aid future security analysis.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Spectre Attacks: Exploiting Speculative ExecutionPaul Kocher, Jann Horn, Anders Fogh, Daniel Genkin 等S&P 2019 · 被引用 2,435 次
- Meltdown: Reading Kernel Memory from User SpaceMoritz Lipp, Michael Schwarz, Daniel Gruss, Thomas Prescher 等USENIX Security 2018 · 被引用 1,456 次
- Foreshadow: Extracting the Keys to the Intel SGX Kingdom with Transient Out-of-Order ExecutionJo Van Bulck, Marina Minkin, Ofir Weisse, Daniel Genkin 等USENIX Security 2018 · 被引用 1,175 次
- Fallout: Leaking Data on Meltdown-resistant CPUsClaudio Canella, Daniel Genkin, Lukas Giner, Daniel Gruss 等CCS 2019 · 被引用 289 次
- LVI: Hijacking Transient Execution through Microarchitectural Load Value InjectionJo Van Bulck, Daniel Moghimi, Michael Schwarz, Moritz Lipp 等S&P 2020 · 被引用 275 次
相关 Paper
- libLISA: Instruction Discovery and Analysis on x86-64Jos Craaijo, Freek Verbeek, Binoy RavindranOOPSLA 2024 · 被引用 1 次
- iDEV: exploring and exploiting semantic deviations in ARM instruction processingShisong Qin, Chao Zhang, Kaixiang Chen, Zheming LiISSTA 2021 · 被引用 7 次
- D-ARM: Disassembling ARM Binaries by Lightweight Superset Instruction Interpretation and Graph ModelingYapeng Ye, Zhuo Zhang, Qingkai Shi, Yousra Aafer 等S&P 2023
- ArchSem: Reusable Rigorous Semantics of Relaxed ArchitecturesThibaut Pérami, Thomas Bauereiss, Brian Campbell, Zongyuan Liu 等POPL 2026 · 被引用 1 次
- Semantics-Guided Control-Flow Reconstruction for Firmware Binaries via Static AnalysisFengjuan Gao, Qingjie Zhu, Yi Zhang, Yu Wang 等FSE 2026
