Nova: Generative Language Models for Assembly Code with Hierarchical Attention and Contrastive Learning
Nan Jiang, Chengxiao Wang, Kevin Liu, Xiangzhe Xu, Lin Tan, Xiangyu Zhang, Petr Babkin
Abstract
Binary code analysis is the foundation of crucial tasks in the security domain; thus building effective binary analysis techniques is more important than ever. Large language models (LLMs) although have brought impressive improvement to source code tasks, do not directly generalize to assembly code due to the unique challenges of assembly: (1) the low information density of assembly and (2) the diverse optimizations in assembly code. To overcome these challenges, this work proposes a hierarchical attention mechanism that builds attention summaries to capture the semantics more effectively, and designs contrastive learning objectives to train LLMs to learn assembly optimization. Equipped with these techniques, this work develops Nova, a generative LLM for assembly code. Nova outperforms existing techniques on binary code decompilation by up to 14.84 -21.58% (absolute percentage point improvement) higher Pass@1 and Pass@10, and outperforms the latest binary code similarity detection techniques by up to 6.17% Recall@1, showing promising abilities on both assembly generation and understanding tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3be0a6ab-d708-43d2-8df4-a121b372af99Cited by top-tier papers2
- DecLLM: LLM-Augmented Recompilable Decompilation for Enabling Programmatic Use of Decompiled CodeWai Kin Wong, Daoyuan Wu, Huaijin Wang, Zongjie Li et al.ISSTA 2025 · 8 citations
- Idioms: A Simple and Effective Framework for Turbo-Charging Local Neural Decompilation with Well-Defined TypesLuke Dramko, Claire Le Goues, Edward J. SchwartzNDSS 2026 · 7 citations
Builds on30
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
- Neural Network-based Graph Embedding for Cross-Platform Binary Code Similarity DetectionXiaojun Xu, Chang Liu, Qian Feng, Heng Yin et al.CCS 2017 · 682 citations
Related papers
- Beyond Classification: Inferring Function Names in Stripped Binaries via Domain Adapted LLMsLinxi Jiang, Xin Jin, Zhiqiang LinNDSS 2025
- Transforming Generic Coder LLMs to Effective Binary Code Embedding Models for Similarity DetectionLitao Li, Leo Song, Steven H. H. Ding, Benjamin C. M. Fung et al.NeurIPS 2025 · 2 citations
- WaDec: Decompiling WebAssembly Using Large Language ModelXinyu She, Yanjie Zhao, Haoyu WangASE 2024 · 9 citations
- LLM4Decompile: Decompiling Binary Code with Large Language ModelsHanzhuo Tan, Qi Luo, Jing Li, Yuqun ZhangEMNLP 2024 · 31 citations
- Uncovering LLM-Generated Code: A Zero-Shot Synthetic Code Detector via Code RewritingTong Ye, Yangkai Du, Tengfei Ma, Lingfei Wu et al.AAAI 2025 · 21 citations
