Towards Generality: Task-Adaptive Binary Analysis via Semantic Retrieval and Verifiable Reasoning
Yuzhe Liu, Zhijie Liu, Zhengmin Yu, Shu Wang, Ling Jiang, Sen Nie, Shi Wu, Zhanyong Tang, Yuan Zhang
摘要
Stripped binaries dominate real-world security analysis: COTS software, firmware, and malware are released without symbols, obscuring program semantics. Recent LLMbased approaches target narrowly predefined tasks, limiting applicability in intent-driven analysis. Toward general binary analysis, agentic LLMs provide a natural direction, yet achieving generality remains challenging: query intent is hard to ground in stripped binaries, and effective tool-driven analysis is difficult.
We present BINREX, the first agentic framework for general, fully automated static binary analysis. BINREX employs a dual-encoder pipeline to enrich stripped functions with semantic context, paired with hierarchical planning to decompose intents into executable subtasks. It leverages code synthesis to translate subtasks into IDAPython scripts, executed with iterative validation to produce outputs aggregated into the final report.
We evaluate BINREX on BinREval, a benchmark of 72 tasks across seven COTS categories with machine-checkable oracles. BINREX achieves 83.3% success rate, outperforming baselines (Codex: 27.8%, Codex with IDA: 43.1%) while reducing total analysis time from 87.9 to 36.1 hours. On domain benchmarks (Juliet, NYU CTF), BINREX is competitive with task-specific methods on their respective evaluation settings. Human study and industrial deployment confirm practical value: BINREX identifies 395 unknown malware samples with 96× average efficiency gain over experts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper37
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- Understanding the Mirai BotnetManos Antonakakis, Tim April, Michael D. Bailey, Matt Bernhard 等USENIX Security 2017 · 被引用 2,003 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 被引用 1,085 次
- Angora: Efficient Fuzzing by Principled SearchPeng Chen, Hao ChenS&P 2018 · 被引用 616 次
相关 Paper
- Beyond Classification: Inferring Function Names in Stripped Binaries via Domain Adapted LLMsLinxi Jiang, Xin Jin, Zhiqiang LinNDSS 2025
- BinQuery: A Novel Framework for Natural Language-Based Binary Code RetrievalBolun Zhang, Zeyu Gao, Hao Wang, Yuxin Cui 等ISSTA 2025 · 被引用 1 次
- SymLM: Predicting Function Names in Stripped Binaries via Context-Sensitive Execution-Aware Code EmbeddingsXin Jin, Kexin Pei, Jun Yeon Won, Zhiqiang LinCCS 2022 · 被引用 56 次
- BinStruct: Binary Structure Recovery Combining Static Analysis and SemanticsYiran Zhang, Zhengzi Xu, Zhe Lang, Chengyue Liu 等ASE 2025
- BinRAG: An RAG-Based Decompilation Framework Fusing Name Prediction and Calling ContextWai Kin Wong, Daoyuan Wu, Zhibo Liu, Huaijin Wang 等ISSTA 2026
