TrojanSQL: SQL Injection against Natural Language Interface to Database
Jinchuan Zhang, Yan Zhou, Binyuan Hui, Yaxin Liu, Ziming Li, Songlin Hu
Abstract
The technology of text-to-SQL has significantly enhanced the efficiency of accessing and manipulating databases. However, limited research has been conducted to study its vulnerabilities emerging from malicious user interaction. By proposing TrojanSQL, a backdoor-based SQL injection framework for text-to-SQL systems, we show how state-of-the-art text-to-SQL parsers can be easily misled to produce harmful SQL statements that can invalidate user queries or compromise sensitive information about the database. The study explores two specific injection attacks, namely boolean-based injection and union-based injection, which use different types of triggers to achieve distinct goals in compromising the parser. Experimental results demonstrate that both medium-sized models based on fine-tuning and LLM-based parsers using prompting techniques are vulnerable to this type of attack, with attack success rates as high as 99% and 89%, respectively. We hope that this study will raise more concerns about the potential security risks of building natural language interfaces to databases.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext de2ad27d-3d2c-4aed-86cf-9f351e70e658Cited by top-tier papers4
- SkillScope: Toward Fine-Grained Least-Privilege Enforcement for Agent SkillsJiangrong Wu, Yuhong Nan, Yixi Lin, Huaijin Wang et al.CCS 2026 · 5 citations
- Are Your LLM-based Text-to-SQL Models Secure? Exploring SQL Injection via Backdoor AttacksMeiyu Lin, Haichuan Zhang, Jiale Lao, Renyuan Li et al.SIGMOD 2026 · 4 citations
- Holistic Automated Red Teaming for Large Language Models through Top-Down Test Case Generation and Multi-turn InteractionJinchuan Zhang, Yan Zhou, Yaxin Liu, Ziming Li et al.EMNLP 2024 · 2 citations
- VisPoison: An Effective Backdoor Attack Framework for Tabular Data Visualization ModelsShuaimin Li, Chen Jason Zhang, Xuanang Chen, Anni Peng et al.ICDE 2026
Builds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 1,085 citations
- Poisoning Language Models During Instruction TuningAlexander Wan, Eric Wallace, Sheng Shen, Dan KleinICML 2023 · 319 citations
- Weight Poisoning Attacks on Pretrained ModelsKeita Kurita, Paul Michel, Graham NeubigACL 2020 · 312 citations
Related papers
- AIA: Autoregression-Based Injection Attacks Against Text2SQL ModelsDeyin Li, Xiang Ling, Changjiang Li, Xiang Chen et al.AAAI 2025
- Structure-Guided Large Language Models for Text-to-SQL GenerationQinggang Zhang, Hao Chen, Junnan Dong, Shengyuan Chen et al.ICML 2025
- Prompt-to-SQL Injections in LLM-Integrated Web Applications: Risks and DefensesRodrigo Pedro, Miguel E. Coimbra, Daniel Castro, Paulo Carreira et al.ICSE 2025 · 14 citations
- "What Do You Mean by That?" A Parser-Independent Interactive Approach for Enhancing Text-to-SQLYuntao Li, Bei Chen, Qian Liu, Yan Gao et al.EMNLP 2020 · 17 citations
- SchemaRAG: A Schema-aware Retrieval-Augmented Generation Framework for Text-to-SQLDi Wu, Zetong Tang, Yi He, Xin LuoSIGMOD 2026 · 9 citations
