CLAP: Learning Transferable Binary Code Representations with Natural Language Supervision
Hao Wang, Zeyu Gao, Chao Zhang, Zihan Sha, Mingyang Sun, Yuchen Zhou, Wenyu Zhu, Wenju Sun, Han Qiu, Xi Xiao
摘要
Binary code representation learning has shown significant performance in binary analysis tasks. But existing solutions often have poor transferability, particularly in few-shot and zero-shot scenarios where few or no training samples are available for the tasks. To address this problem, we present CLAP (Contrastive Language-Assembly Pre-training), which employs natural language supervision to learn better representations of binary code (i.e., assembly code) and get better transferability. At the core, our approach boosts superior transfer learning capabilities by effectively aligning binary code with their semantics explanations (in natural language), resulting a model able to generate better embeddings for binary code. To enable this alignment training, we then propose an efficient dataset engine that could automatically generate a large and diverse dataset comprising of binary code and corresponding natural language explanations. We have generated 195 million pairs of binary code and explanations and trained a prototype of CLAP. The evaluations of CLAP across various downstream tasks in binary analysis all demonstrate exceptional performance. Notably, without any task-specific training, CLAP is often competitive with a fully supervised baseline, showing excellent transferability. We release our pre-trained model and code at https://github.com/Hustcw/CLAP . CCS CONCEPTS • Security and privacy → Software reverse engineering; • Computing methodologies → Machine learning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- Improving ML-based Binary Function Similarity Detection by Assessing and Deprioritizing Control Flow Graph FeaturesJialai Wang, Chao Zhang, Longfei Chen, Yi Rong 等USENIX Security 2024 · 被引用 15 次
- SK2Decompile: LLM-based Two-Phase Binary Decompilation from Skeleton to SkinHanzhuo Tan, Weihao Li, Xiaolong Tian, Siyi Wang 等ICLR 2026 · 被引用 10 次
- ShieldedCode: Learning Robust Representations for Virtual Machine Protected CodeMingqiao Mo, Yunlong Tan, Hao Zhang, Heng Zhang 等ICLR 2026 · 被引用 10 次
- Virtual Compiler Is All You Need For Assembly Code SearchZeyu Gao, Hao Wang, Yuanda Wang, Chao ZhangACL 2024 · 被引用 3 次
- vSim: Semantics-Aware Value Extraction for Efficient Binary Code Similarity AnalysisHuaijin Wang, Zhiqiang LinNDSS 2026 · 被引用 3 次
它引用的顶会 Paper11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Neural Network-based Graph Embedding for Cross-Platform Binary Code Similarity DetectionXiaojun Xu, Chang Liu, Qian Feng, Heng Yin 等CCS 2017 · 被引用 682 次
- Asm2Vec: Boosting Static Representation Robustness for Binary Clone Search against Code Obfuscation and Compiler OptimizationSteven H. H. Ding, Benjamin C. M. Fung, Philippe CharlandS&P 2019 · 被引用 447 次
- Order Matters: Semantic-Aware Neural Networks for Binary Code Similarity DetectionZeping Yu, Rui Cao, Qiyi Tang, Sen Nie 等AAAI 2020 · 被引用 265 次
- Neural Nets Can Learn Function Type Signatures From BinariesZheng Leong Chua, Shiqi Shen, Prateek Saxena, Zhenkai LiangUSENIX Security 2017 · 被引用 175 次
相关 Paper
- Understanding Transferable Representation Learning and Zero-shot Transfer in CLIPZixiang Chen, Yihe Deng, Yuanzhi Li, Quanquan GuICLR 2024 · 被引用 21 次
- Is a Caption Worth a Thousand Images? A Study on Representation LearningShibani Santurkar, Yann Dubois, Rohan Taori, Percy Liang 等ICLR 2023 · 被引用 9 次
- PalmTree: Learning an Assembly Language Model for Instruction EmbeddingXuezixiang Li, Yu Qu, Heng YinCCS 2021 · 被引用 139 次
- Advancing Multi-grained Alignment for Contrastive Language-Audio Pre-trainingYiming Li, Zhifang Guo, Xiangdong Wang, Hong LiuACM MM 2024 · 被引用 9 次
- CompA: Addressing the Gap in Compositional Reasoning in Audio-Language ModelsSreyan Ghosh, Ashish Seth, Sonal Kumar, Utkarsh Tyagi 等ICLR 2024 · 被引用 53 次
