Virtual Compiler Is All You Need For Assembly Code Search
Zeyu Gao, Hao Wang, Yuanda Wang, Chao Zhang
摘要
Assembly code search is vital for reducing the burden on reverse engineers, allowing them to quickly identify specific functions using natural language within vast binary programs. Despite its significance, this critical task is impeded by the complexities involved in building highquality datasets. This paper explores training a Large Language Model (LLM) to emulate a general compiler. By leveraging Ubuntu packages to compile a dataset of 20 billion tokens, we further continue pre-train CodeLlama as a Virtual Compiler (ViC), capable of compiling any source code of any language to assembly code. This approach allows for virtual compilation across a wide range of programming languages without the need for a real compiler, preserving semantic equivalency and expanding the possibilities for assembly code dataset construction. Furthermore, we use ViC to construct a sufficiently large dataset for assembly code search. Employing this extensive dataset, we achieve a substantial improvement in assembly code search performance, with our model surpassing the leading baseline by 26%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- BinQuery: A Novel Framework for Natural Language-Based Binary Code RetrievalBolun Zhang, Zeyu Gao, Hao Wang, Yuxin Cui 等ISSTA 2025 · 被引用 1 次
- EquiBench: Benchmarking Large Language Models' Reasoning about Program Semantics via Equivalence CheckingAnjiang Wei, Jiannan Cao, Ran Li, Hongyu Chen 等EMNLP 2025
- Selective Knowledge Distillation: Fusing LLM Semantic Strengths with DNN Efficiency for Binary Code Similarity DetectionShize Zhou, Peiyu Liu, Lirong Fu, Tong Ye 等ACL 2026
它引用的顶会 Paper9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng 等ICLR 2021 · 被引用 1,644 次
- CodeT5+: Open Code Large Language Models for Code Understanding and GenerationYue Wang, Hung Le, Akhilesh Gotmare, Nghi D. Q. Bui 等EMNLP 2023 · 被引用 339 次
- Retrieval-Augmented Generation for Code Summarization via Hybrid GNNShangqing Liu, Yu Chen, Xiaofei Xie, Jing Kai Siow 等ICLR 2021 · 被引用 194 次
- PalmTree: Learning an Assembly Language Model for Instruction EmbeddingXuezixiang Li, Yu Qu, Heng YinCCS 2021 · 被引用 139 次
相关 Paper
- Nova: Generative Language Models for Assembly Code with Hierarchical Attention and Contrastive LearningNan Jiang, Chengxiao Wang, Kevin Liu, Xiangzhe Xu 等ICLR 2025
- Can LLMs Obfuscate Code? A Systematic Analysis of Large Language Models into Assembly Code ObfuscationSeyedreza Mohseni, Seyedali Mohammadi, Deepa Tilwani, Yash Saxena 等AAAI 2025 · 被引用 6 次
- Asm2Vec: Boosting Static Representation Robustness for Binary Clone Search against Code Obfuscation and Compiler OptimizationSteven H. H. Ding, Benjamin C. M. Fung, Philippe CharlandS&P 2019 · 被引用 447 次
- QiMeng-NeuComBack: Self-Evolving Translation from IR to Assembly CodeHainan Fang, Yuanbo Wen, Jun Bi, Yihan Wang 等NeurIPS 2025
- Multi-modal Learning for WebAssembly Reverse EngineeringHanxian Huang, Jishen ZhaoISSTA 2024 · 被引用 1 次
