MiSum: Multi-modality Heterogeneous Code Graph Learning for Multi-intent Binary Code Summarization
Kangchen Zhu, Zhiliang Tian, Shangwen Wang, Weiguo Chen, Zixuan Dong, Mingyue Leng, Xiaoguang Mao
摘要
The current landscape of binary code summarization predominantly focuses on generating a single summary, which limits its utility and understanding for reverse engineers. Existing approaches often fail to meet the diverse needs of users, such as providing detailed insights into usage patterns, implementation nuances, and design rationale, as observed in the field of source code summarization. This highlights the need for multi-intent binary code summarization to enhance the effectiveness of reverse engineering processes. To address this gap, we propose MiSum, a novel method that leverages multi-modality heterogeneous code graph alignment and learning to integrate both assembly code and pseudo-code. MiSum introduces a unified multi-modality heterogeneous code graph (MM-HCG) that aligns assembly code graphs with pseudo-code graphs, capturing both low-level execution details and high-level structural information. We further propose multi-modality heterogeneous graph learning with heterogeneous mutual attention and message passing, which highlights important code blocks and discovers inter-dependencies across different code forms. Additionally, an intent-aware summary generator with an intent-aware attention mechanism is introduced to produce customized summaries tailored to multiple intents. Extensive experiments, including evaluations across various architectures and optimization levels, demonstrate that MiSum outperforms state-of-the-art baselines in BLEU, METEOR, and ROUGE-L metrics. Human evaluations validate its capability to effectively support reverse engineers in understanding diverse binary code intents, marking a significant advancement in binary code analysis.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Three Heads Are Better Than One: A Multi-perspective Reasoning Framework for Enhanced Vulnerability DetectionXin Peng, Bo Lin, Jing Wang, Xiaoling Li 等FSE 2026 · 被引用 1 次
- Atomizer: An LLM-based Collaborative Multi-Agent Framework for Intent-Driven Commit UntanglingKangchen Zhu, Zhiliang Tian, Shangwen Wang, Mingyue Leng 等ICSE 2026
- Towards Generality: Task-Adaptive Binary Analysis via Semantic Retrieval and Verifiable ReasoningYuzhe Liu, Zhijie Liu, Zhengmin Yu, Shu Wang 等USENIX Security 2026
它引用的顶会 Paper14
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- WizardCoder: Empowering Code Large Language Models with Evol-InstructZiyang Luo, Can Xu, Pu Zhao, Qingfeng Sun 等ICLR 2024 · 被引用 945 次
- A Survey on In-context LearningQingxiu Dong, Lei Li, Damai Dai, Ce Zheng 等EMNLP 2024 · 被引用 479 次
- CodeGen: An Open Large Language Model for Code with Multi-Turn Program SynthesisErik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu 等ICLR 2023 · 被引用 234 次
相关 Paper
- CP-BCS: Binary Code Summarization Guided by Control Flow Graph and Pseudo CodeTong Ye, Lingfei Wu, Tengfei Ma, Xuhong Zhang 等EMNLP 2023 · 被引用 4 次
- Multi-modal Learning for WebAssembly Reverse EngineeringHanxian Huang, Jishen ZhaoISSTA 2024 · 被引用 1 次
- Source Code Foundation Models are Transferable Binary Analysis Knowledge BasesZian Su, Xiangzhe Xu, Ziyang Huang, Kaiyuan Zhang 等NeurIPS 2024 · 被引用 17 次
- HexT5: Unified Pre-Training for Stripped Binary Code Information InferenceJiaqi Xiong, Guoqiang Chen, Kejiang Chen, Han Gao 等ASE 2023 · 被引用 8 次
- Retrieval-Augmented Generation for Code Summarization via Hybrid GNNShangqing Liu, Yu Chen, Xiaofei Xie, Jing Kai Siow 等ICLR 2021 · 被引用 194 次
