Beyond Traditional Benchmarks: Analyzing Behaviors of Open LLMs on Data-to-Text Generation
Zdenek Kasner, Ondrej Dusek
Abstract
We analyze the behaviors of open large language models (LLMs) on the task of data-totext (D2T) generation, i.e., generating coherent and relevant text from structured data. To avoid the issue of LLM training data contamination with standard benchmarks, we design QUINTD -a tool for collecting novel structured data records from public APIs. We find that open LLMs (Llama 2, Mistral, and Zephyr) can generate fluent and coherent texts in zero-shot settings from data in common formats collected with QUINTD. However, we show that the semantic accuracy of the outputs is a major issue: both according to human annotators and our reference-free metric based on GPT-4, more than 80% of the outputs of open LLMs contain at least one semantic error. We publicly release the code, data, and model outputs. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b4afd85b-a5b0-4397-a0db-11d9c3eb3f2eCited by top-tier papers2
- UPME: An Unsupervised Peer Review Framework for Multimodal Large Language Model EvaluationQihui Zhang, Munan Ning, Zheyuan Liu, Yue Huang et al.CVPR 2025
- Justice or Prejudice? Quantifying Biases in LLM-as-a-JudgeJiayi Ye, Yanbo Wang, Yue Huang, Dongping Chen et al.ICLR 2025
Builds on9
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang et al.EMNLP 2023 · 549 citations
- Can Large Language Models Be an Alternative to Human Evaluations?David Cheng-Han Chiang, Hung-yi LeeACL 2023 · 254 citations
- Time Travel in LLMs: Tracing Data Contamination in Large Language ModelsShahriar Golchin, Mihai SurdeanuICLR 2024 · 165 citations
Related papers
- Exploring Precision and Recall to assess the quality and diversity of LLMsFlorian Le Bronnec, Alexandre Verine, Benjamin Négrevergne, Yann Chevaleyre et al.ACL 2024 · 11 citations
- Quality Matters: Evaluating Synthetic Data for Tool-Using LLMsShadi Iskander, Sofia Tolmach, Ori Shapira, Nachshon Cohen et al.EMNLP 2024 · 2 citations
- StrucText-Eval: Evaluating Large Language Model's Reasoning Ability in Structure-Rich TextZhouhong Gu, Haoning Ye, Xingzhou Chen, Zeyang Zhou et al.ACL 2025
- An Empirical Study of Many-to-Many Summarization with Large Language ModelsJiaan Wang, Fandong Meng, Zengkui Sun, Yunlong Liang et al.ACL 2025
- DNA-GPT: Divergent N-Gram Analysis for Training-Free Detection of GPT-Generated TextXianjun Yang, Wei Cheng, Yue Wu, Linda Ruth Petzold et al.ICLR 2024 · 173 citations
