USENIX Security2026Top-tier venue
Towards Generality: Task-Adaptive Binary Analysis via Semantic Retrieval and Verifiable Reasoning
Yuzhe Liu, Zhijie Liu, Zhengmin Yu, Shu Wang, Ling Jiang, Sen Nie, Shi Wu, Zhanyong Tang, Yuan Zhang
Abstract
Stripped binaries dominate real-world security analysis: COTS software, firmware, and malware are released without symbols, obscuring program semantics. Recent LLMbased approaches target narrowly predefined tasks, limiting applicability in intent-driven analysis. Toward general binary analysis, agentic LLMs provide a natural direction, yet achieving generality remains challenging: query intent is hard to ground in stripped binaries, and effective tool-driven analysis is difficult.
We present BINREX, the first agentic framework for general, fully automated static binary analysis. BINREX employs a dual-encoder pipeline to enrich stripped functions with semantic context, paired with hierarchical planning to decompose intents into executable subtasks. It leverages code synthesis to translate subtasks into IDAPython scripts, executed with iterative validation to produce outputs aggregated into the final report.
We evaluate BINREX on BinREval, a benchmark of 72 tasks across seven COTS categories with machine-checkable oracles. BINREX achieves 83.3% success rate, outperforming baselines (Codex: 27.8%, Codex with IDA: 43.1%) while reducing total analysis time from 87.9 to 36.1 hours. On domain benchmarks (Juliet, NYU CTF), BINREX is competitive with task-specific methods on their respective evaluation settings. Human study and industrial deployment confirm practical value: BINREX identifies 395 unknown malware samples with 96× average efficiency gain over experts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b2a2e17c-6d0e-4e8e-96c0-a2e2f53cac73Builds on37
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- Understanding the Mirai BotnetManos Antonakakis, Tim April, Michael D. Bailey, Matt Bernhard et al.USENIX Security 2017 · 2,003 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 1,085 citations
- Angora: Efficient Fuzzing by Principled SearchPeng Chen, Hao ChenS&P 2018 · 616 citations
Related papers
- Beyond Classification: Inferring Function Names in Stripped Binaries via Domain Adapted LLMsLinxi Jiang, Xin Jin, Zhiqiang LinNDSS 2025
- BinQuery: A Novel Framework for Natural Language-Based Binary Code RetrievalBolun Zhang, Zeyu Gao, Hao Wang, Yuxin Cui et al.ISSTA 2025 · 1 citation
- SymLM: Predicting Function Names in Stripped Binaries via Context-Sensitive Execution-Aware Code EmbeddingsXin Jin, Kexin Pei, Jun Yeon Won, Zhiqiang LinCCS 2022 · 56 citations
- BinStruct: Binary Structure Recovery Combining Static Analysis and SemanticsYiran Zhang, Zhengzi Xu, Zhe Lang, Chengyue Liu et al.ASE 2025
- BinRAG: An RAG-Based Decompilation Framework Fusing Name Prediction and Calling ContextWai Kin Wong, Daoyuan Wu, Zhibo Liu, Huaijin Wang et al.ISSTA 2026
