Lune

ICDE2026Top-tier venue

CYANSQL: Unlock the Power of NL2SQL Via Clustering-Based Test-Time Scaling

Haoyu Qin, Tonghui Ren, Zhenying He, X. Sean Wang, Jiashu Xing, Yanghuan Ye, Shifei Huang, Jinbao Li

2026Year

Abstract

Large language models (LLMs) are playing an increasingly important role in the Natural Language to SQL (NL2SQL) task. Although LLMs exhibit strong inherent reasoning abilities, they often fail to generate correct SQL because they lack knowledge of the target logical operator composition. Few-shot learning is a commonly employed approach to provide relevant knowledge to LLMs, and many studies focus on how to select effective demonstrations. However, these methods are often limited by the quality of the selected demonstrations, resulting in a failure to fully exploit the rich logical operation information contained in the demonstrations. In this work, we propose CYANSQL (Cluster-aware Yielded Augmentation for NL2SQL), a method that leverages parallel test-time scaling to strengthen LLM reasoning. We first cluster the demonstrations based on SQL logical operator composition, then perform multi-path generation by injecting demonstrations from each cluster during LLM inference, followed by a ranking stage to identify the optimal SQL. CYANSQL achieves an execution accuracy of 73.47% and a state-of-the-art recall of 87.22% on the development set of the popular NL2SQL benchmark BIRD. The source code of this project is available at: https://github.com/zyddqhy/CYANSQL

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines