Nova: Generative Language Models for Assembly Code with Hierarchical Attention and Contrastive Learning
Nan Jiang, Chengxiao Wang, Kevin Liu, Xiangzhe Xu, Lin Tan, Xiangyu Zhang, Petr Babkin
摘要
Binary code analysis is the foundation of crucial tasks in the security domain; thus building effective binary analysis techniques is more important than ever. Large language models (LLMs) although have brought impressive improvement to source code tasks, do not directly generalize to assembly code due to the unique challenges of assembly: (1) the low information density of assembly and (2) the diverse optimizations in assembly code. To overcome these challenges, this work proposes a hierarchical attention mechanism that builds attention summaries to capture the semantics more effectively, and designs contrastive learning objectives to train LLMs to learn assembly optimization. Equipped with these techniques, this work develops Nova, a generative LLM for assembly code. Nova outperforms existing techniques on binary code decompilation by up to 14.84 -21.58% (absolute percentage point improvement) higher Pass@1 and Pass@10, and outperforms the latest binary code similarity detection techniques by up to 6.17% Recall@1, showing promising abilities on both assembly generation and understanding tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- DecLLM: LLM-Augmented Recompilable Decompilation for Enabling Programmatic Use of Decompiled CodeWai Kin Wong, Daoyuan Wu, Huaijin Wang, Zongjie Li 等ISSTA 2025 · 被引用 8 次
- Idioms: A Simple and Effective Framework for Turbo-Charging Local Neural Decompilation with Well-Defined TypesLuke Dramko, Claire Le Goues, Edward J. SchwartzNDSS 2026 · 被引用 7 次
它引用的顶会 Paper30
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- Neural Network-based Graph Embedding for Cross-Platform Binary Code Similarity DetectionXiaojun Xu, Chang Liu, Qian Feng, Heng Yin 等CCS 2017 · 被引用 682 次
相关 Paper
- Beyond Classification: Inferring Function Names in Stripped Binaries via Domain Adapted LLMsLinxi Jiang, Xin Jin, Zhiqiang LinNDSS 2025
- Transforming Generic Coder LLMs to Effective Binary Code Embedding Models for Similarity DetectionLitao Li, Leo Song, Steven H. H. Ding, Benjamin C. M. Fung 等NeurIPS 2025 · 被引用 2 次
- WaDec: Decompiling WebAssembly Using Large Language ModelXinyu She, Yanjie Zhao, Haoyu WangASE 2024 · 被引用 9 次
- LLM4Decompile: Decompiling Binary Code with Large Language ModelsHanzhuo Tan, Qi Luo, Jing Li, Yuqun ZhangEMNLP 2024 · 被引用 31 次
- Uncovering LLM-Generated Code: A Zero-Shot Synthetic Code Detector via Code RewritingTong Ye, Yangkai Du, Tengfei Ma, Lingfei Wu 等AAAI 2025 · 被引用 21 次
