jqBench: a benchmark for reading and editing JSON from natural language and/or examples
Gust Verbruggen, Chris Parnin, Vu Le, Sumit Gulwani
Abstract
We introduce JQBENCH, a new benchmark for evaluating language models on JSON querying and transformation tasks, where the intent can be given specified using natural language and/or examples. Whereas JQBENCH is mainly aimed at using the jq tool, it can be used to evaluate other programming languages that query and/or transform JSON. Benchmarks are automatically created from two rich sources of data: Stack Overflow discussions (1496 instances with instructions and examples, called JQSTACK) and the Spider dataset for SQL generation from natural language (859 instances with instructions and JSON Schema, called JQSPIDER). We describe and analyze the automated pipeline for benchmark creation, and perform extensive baseline experiments on different models to analyze the complexity and failure modes. Using implicit feedback, the best model (Opus 4.1) scores 76% on the JQSTACK benchmarks and 81% on the JQSPIDER benchmarks. Additionally, we show (1) that access to the documentation surprisingly does not help, (2) jq lags behind Python, and (3) that automatic feedback (and therefore examples) is crucial. Besides the challenging benchmarks, we release 13K converted but filtered cases for training purposes. → → ## Example 1 Command: jq '.foo?'* *Input**: "foo": 42* Output*: 42# # Example 2 **Command**: jq '.foo?'* Input*: "bar": 1* *Output**: null`1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on8
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 1,085 citations
- Synthesizing Natural Language to Visualization (NL2VIS) Benchmarks from NL2SQL BenchmarksYuyu Luo, Nan Tang, Guoliang Li, Chengliang Chai et al.SIGMOD 2021 · 90 citations
- Is Programming by Example Solved by LLMs?Wen-Ding Li, Kevin EllisNeurIPS 2024 · 45 citations
- Knowledge Transfer from High-Resource to Low-Resource Programming Languages for Code LLMsFederico Cassano, John Gouwar, Francesca Lucchetti, Claire Schlesinger et al.OOPSLA 2024 · 33 citations
Related papers
- ScienceBenchmark: A Complex Real-World Benchmark for Evaluating Natural Language to SQL SystemsYi Zhang, Jan Deriu, George Katsogiannis-Meimarakis, Catherine Kosten et al.VLDB 2024 · 65 citations
- GQLBench: A Large-Scale Cross-Domain, Cross-Dialect Benchmark for NL2GQLYanning Su, Yuhang Zhou, Yang Fang, Sen Liu et al.ACL 2026
- NL2SQLBench: A Modular Benchmarking Framework for LLM-Enabled NL2SQL SolutionsShizheng Hou, Wenqi Pei, Nuo Chen, Quang-Trung Ta et al.VLDB 2026 · 1 citation
- Evaluating Cross-Domain Text-to-SQL Models and BenchmarksMohammadreza Pourreza, Davood RafieiEMNLP 2023 · 14 citations
- BETZE: Benchmarking Data Exploration Tools with (Almost) Zero EffortNico Schäfer, Sebastian MichelICDE 2022 · 2 citations
