TranSQL + : Serving Large Language Models with SQL on Low-Resource Hardware
Wenbo Sun, Qiming Guo, Wenlu Wang, Rihan Hai
Abstract
Deploying Large Language Models (LLMs) on resource-constrained devices remains challenging due to limited memory, lack of GPUs, and the complexity of existing runtimes. In this paper, we introduce TranSQL + , a template-based code generator that translates LLM computation graphs into pure SQL queries for execution in relational databases. Without relying on external libraries, TranSQL + , leverages mature database features-such as vectorized execution and out-of-core processing-for efficient inference. We further propose a row-to-column (ROW2COL) optimization that improves join efficiency in matrix operations. Evaluated on Llama3-8B and DeepSeekMoE models, TranSQL + achieves up to 20× lower prefill latency and 4× higher decoding speed compared to DeepSpeed Inference and Llama.cpp in low-memory and CPU-only configurations. Our results highlight relational databases as a practical environment for LLMs on low-resource hardware.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e8f6ba3f-2f6a-4e3f-87a1-670a53496b87Builds on8
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPUYing Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li et al.ICML 2023 · 683 citations
- DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model ServingYinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu et al.OSDI 2024 · 646 citations
- LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation DatasetLianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li et al.ICLR 2024 · 419 citations
- DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language ModelsDamai Dai, Chengqi Deng, Chenggang Zhao, R. X. Xu et al.ACL 2024 · 171 citations
Related papers
- Reliable Answers for Recurring Questions: Boosting Text-to-SQL Accuracy with Template Constrained DecodingSmit Jivani, Sarvam Maheshwari, Sunita SarawagiSIGMOD 2026 · 1 citation
- TransMLA: Migrating GQA Models to MLA with Full DeepSeek Compatibility and SpeedupFanxu Meng, Pingzhi Tang, Zengwei Yao, Xing Sun et al.NeurIPS 2025 · 5 citations
- BOND: A Co-Designed Framework for LLM-Powered Analytics Over Relational DataLixiang Chen, Qin Zheng, Zhicheng Pan, Chengcheng Yang et al.ICDE 2026
- LightRetriever: A LLM-based Text Retrieval Architecture with Extremely Faster Query InferenceGuangyuan Ma, Yongliang Ma, Xuanrui Gou, Zhenpeng Su et al.ICLR 2026 · 3 citations
- LLM4Hint: Leveraging Large Language Models for Hint Recommendation in Offline Query OptimizationSuchen Liu, Yang Lin, Yinjun Han, Jun GaoICDE 2026 · 1 citation
