jqBench: a benchmark for reading and editing JSON from natural language and/or examples
Gust Verbruggen, Chris Parnin, Vu Le, Sumit Gulwani
摘要
We introduce JQBENCH, a new benchmark for evaluating language models on JSON querying and transformation tasks, where the intent can be given specified using natural language and/or examples. Whereas JQBENCH is mainly aimed at using the jq tool, it can be used to evaluate other programming languages that query and/or transform JSON. Benchmarks are automatically created from two rich sources of data: Stack Overflow discussions (1496 instances with instructions and examples, called JQSTACK) and the Spider dataset for SQL generation from natural language (859 instances with instructions and JSON Schema, called JQSPIDER). We describe and analyze the automated pipeline for benchmark creation, and perform extensive baseline experiments on different models to analyze the complexity and failure modes. Using implicit feedback, the best model (Opus 4.1) scores 76% on the JQSTACK benchmarks and 81% on the JQSPIDER benchmarks. Additionally, we show (1) that access to the documentation surprisingly does not help, (2) jq lags behind Python, and (3) that automatic feedback (and therefore examples) is crucial. Besides the challenging benchmarks, we release 13K converted but filtered cases for training purposes. → → ## Example 1 Command: jq '.foo?'* *Input**: "foo": 42* Output*: 42# # Example 2 **Command**: jq '.foo?'* Input*: "bar": 1* *Output**: null`1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 被引用 1,085 次
- Synthesizing Natural Language to Visualization (NL2VIS) Benchmarks from NL2SQL BenchmarksYuyu Luo, Nan Tang, Guoliang Li, Chengliang Chai 等SIGMOD 2021 · 被引用 90 次
- Is Programming by Example Solved by LLMs?Wen-Ding Li, Kevin EllisNeurIPS 2024 · 被引用 45 次
- Knowledge Transfer from High-Resource to Low-Resource Programming Languages for Code LLMsFederico Cassano, John Gouwar, Francesca Lucchetti, Claire Schlesinger 等OOPSLA 2024 · 被引用 33 次
相关 Paper
- ScienceBenchmark: A Complex Real-World Benchmark for Evaluating Natural Language to SQL SystemsYi Zhang, Jan Deriu, George Katsogiannis-Meimarakis, Catherine Kosten 等VLDB 2024 · 被引用 65 次
- GQLBench: A Large-Scale Cross-Domain, Cross-Dialect Benchmark for NL2GQLYanning Su, Yuhang Zhou, Yang Fang, Sen Liu 等ACL 2026
- NL2SQLBench: A Modular Benchmarking Framework for LLM-Enabled NL2SQL SolutionsShizheng Hou, Wenqi Pei, Nuo Chen, Quang-Trung Ta 等VLDB 2026 · 被引用 1 次
- Evaluating Cross-Domain Text-to-SQL Models and BenchmarksMohammadreza Pourreza, Davood RafieiEMNLP 2023 · 被引用 14 次
- BETZE: Benchmarking Data Exploration Tools with (Almost) Zero EffortNico Schäfer, Sebastian MichelICDE 2022 · 被引用 2 次
