The Deception Delta: Adversarial Evaluation of LLM-Based Smart Contract Bytecode Forensics
Timo Schefold
Abstract
Large language models are increasingly used in blockchain forensic investigations to interpret unverified smart contract bytecode. Their robustness has not been systematically tested against contracts adversarially designed to mislead analysis.
We evaluate 22 frontier models on 13 purpose-built contracts (9 deception vectors, 4 controls) across six prompt strategies, yielding 8,528 analyzable non-refusal runs against contracts with EVMverified ground truth. A calibrated LLM-as-judge pipeline, supported by two judge-independent metrics and 50 human goldstandard labels, shows that adversarial deception reduces drain detection by 20.0 percentage points (95% CI: [17.2, 22.8]) relative to functionally matched controls.
Structural camouflage via multi-hop call chains, XOR-masked selectors, and storage-loaded drain parameters resists detection across nearly all models. Beyond non-detection, we identify rationalization: models correctly describe the hidden drain mechanism but accept the contract's deceptive framing and dismiss it as benign, yielding positive but incorrect evidence of safety. Simple guard instructions provide no aggregate benefit and destabilize individual models in both directions.
Structural deception is largely insensitive across the six tested prompt strategies, more consistent with a capability limitation than with a simple prompting problem. Only five models from two providers exceed 50% detection.
Under our single-shot, raw-bytecode-only protocol, current LLMs are not reliable standalone forensic tools. Our central claim does not extend to multi-turn, tool-augmented, source-aware, or decompiler-in-the-loop workflows; a source-code boundary check is reported as an explicit subset analysis rather than as part of the main evaluation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 09a17757-6f18-4335-86eb-18cc034c43e2Builds on14
- Making Smart Contracts SmarterLoi Luu, Duc-Hiep Chu, Hrishi Olickel, Prateek Saxena et al.CCS 2016 · 2,306 citations
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 1,792 citations
- Securify: Practical Security Analysis of Smart ContractsPetar Tsankov, Andrei Marian Dan, Dana Drachsler-Cohen, Arthur Gervais et al.CCS 2018 · 1,108 citations
- Towards Understanding Sycophancy in Language ModelsMrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud et al.ICLR 2024 · 762 citations
- Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code ContributionsHammond Pearce, Baleegh Ahmad, Benjamin Tan, Brendan Dolan-Gavitt et al.S&P 2022 · 725 citations
Related papers
- LH-DECEPTION: Simulating and Understanding LLM Deceptive Behaviors in Long-Horizon InteractionsYang Xu, Xuanming Zhang, Samuel (Min-Hsuan) Yeh, Jwala Dhamala et al.ICLR 2026 · 7 citations
- OctopusGuard: K-Line Enhanced Token Scam Detector Powered by Multimodal LLMsLitong Sun, YangTian Mi, Xiapu Luo, Weigang WuICSE 2026
- Beyond Prompt-Induced Lies: Investigating LLM Deception on Benign PromptsZhaomin Wu, Mingzhe Du, See-Kiong Ng, Bingsheng HeICLR 2026 · 11 citations
- Trust Me, I Know This Function: Hijacking LLM Static Analysis using BiasShir Bernstein, David Beste, Daniel Ayzenshteyn, Lea Schönherr et al.NDSS 2026 · 7 citations
- PRISON: Unmasking the Criminal Potential of Large Language ModelsXinyi Wu, Geng Hong, Pei Chen, Yueyue Chen et al.ICLR 2026 · 3 citations
