From Similarity Ranking to Definitive Verdict: LLM-Enhanced Source-to-Binary Function Localization
Jingyi Shi, Chengyue Liu, Zhengzi Xu, Yang Xiao, Xingchu Chen, Yeting Li, Wei Huo, Yang Liu
摘要
Locating a known source function in a stripped binary is a prerequisite for many security and software engineering tasks, including Software Composition Analysis (SCA) false-positive elimination, patch presence verification, malware analysis, code plagiarism detection, and license compliance auditing. We formalize this need as source-to-binary function localization: given the source code of a target function and its encompassing source package, determine whether the function is present in a stripped binary and, if so, report its address. Two fundamental challenges arise: cross-modal alignment, as source code and stripped binary reside in vastly different representation spaces; and similar function disambiguation, as compilation erases the symbolic features that distinguish functionally similar functions. We present XLoc, a recall-then-verify framework built on two insights. First, cross-modal alignment does not require costly and error-prone compilation; it only demands token-level alignment, a process that can be reliably approximated. Second, the information needed to disambiguate similar functions is already available on the source side and can be extracted ahead of time to guide verification. Building on these insights, XLoc implements a multi-stage recall module in which an LLM transforms source code into pseudo-decompiled representations aligned with binary decompilation output, bridging the cross-modal gap. For verification, XLoc identifies potentially confusing similar functions, extracts differential summaries, and uses them to guide the verification process toward the specific distinguishing evidence for each candidate, producing definitive accept/reject verdicts rather than similarity rankings. We evaluate XLoc on two complementary datasets spanning 196 CVEs, 480 vulnerable functions, and 756 binaries. XLoc achieves up to 84.4% localization accuracy (4.2× over the best baseline) and HM=87.1% for positive/negative discrimination (vs. 35.1% for the best baseline). These results demonstrate that XLoc can locate target functions with high accuracy, reliably discriminate between positive and negative cases, and produce definitive verdicts.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- BinaryAI: Binary Software Composition Analysis via Intelligent Binary Source Code MatchingLing Jiang, Junwen An, Huihui Huang, Qiyi Tang 等ICSE 2024 · 被引用 43 次
- Advancing Binary Code Similarity Detection via Context-Content Fusion and LLM VerificationChaopeng Dong, Jingdong Guo, Shouguo Yang, Yi Li 等ASE 2025
- SBridge: Identifying Source-to-Binary Function Similarity via Cross-Domain Control Block MatchingHeedong Yang, Jeongwoo Lee, Hajin Yun, Seunghoon WooFSE 2026
- Beyond Similarity Scores: Evidence-Based Third-Party Library Detection for C/C++ BinariesChengyue Liu, Zhengzi Xu, Lyuye Zhang, Jiahui Wu 等ISSTA 2026
- Understanding the Limitations of C/C++ Binary Third-Party Library Detection Tool: An Empirical Study at ScaleChengyue Liu, Zhengzi Xu, Kaixuan Li, Jiahui Wu 等FSE 2026
