A Theoretical Framework for Statistical Evaluability of Generative Models
Shashaank Aiyer, Yishay Mansour, Shay Moran, Han Shao
摘要
Statistical evaluation aims to estimate the generalization performance of a model using held-out i.i.d. test data sampled from the ground-truth distribution. In supervised learning settings such as classification, performance metrics such as error rate are well-defined, and test error reliably approximates population error given sufficiently large datasets. In contrast, evaluation is more challenging for generative models due to their open-ended nature: it is unclear which metrics are appropriate and whether such metrics can be reliably evaluated from finite samples. In this work, we introduce a theoretical framework for evaluating language models and establish evaluability results for commonly used metrics. We study two categories of metrics: test-based metrics, including integral probability metrics (IPMs), and similarity-based metrics, including Rényi and KL divergences. We show that IPMs with respect to any bounded test class can be evaluated from finite samples up to multiplicative and additive approximation errors. Moreover, when the test class has finite fat-shattering dimension, IPMs can be evaluated with arbitrary precision. In contrast, similarity-based metrics, including Rényi and KL divergences, are not evaluable from finite samples, as their values can be critically determined by rare events. We also analyze the potential and limitations of perplexity as an evaluation method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos 等ICML 2024 · 被引用 1,212 次
- MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence FrontiersKrishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun 等NeurIPS 2021 · 被引用 606 次
- The Coverage Principle: How Pre-Training Enables Post-TrainingFan Chen, Audrey Huang, Noah Golowich, Sadhika Malladi 等ICLR 2026 · 被引用 28 次
- Sparse Text GenerationPedro Henrique Martins, Zita Marinho, André F. T. MartinsEMNLP 2020 · 被引用 20 次
相关 Paper
- A Unifying Information-theoretic Perspective on Evaluating Generative ModelsAlexis Fox, Samarth Swarup, Abhijin AdigaAAAI 2025 · 被引用 1 次
- Feature Likelihood Score: Evaluating the Generalization of Generative Models Using SamplesMarco Jiralerspong, Avishek Joey Bose, Ian Gemp, Chongli Qin 等NeurIPS 2023 · 被引用 39 次
- On the Usefulness of Embeddings, Clusters and Strings for Text Generation EvaluationTiago Pimentel, Clara Meister, Ryan CotterellICLR 2023
- Evaluating Evaluation Metrics: A Framework for Analyzing NLG Evaluation Metrics using Measurement TheoryZiang Xiao, Susu Zhang, Vivian Lai, Q. Vera LiaoEMNLP 2023 · 被引用 6 次
- Evaluating Distributional Distortion in Neural Language ModelingBenjamin LeBrun, Alessandro Sordoni, Timothy J. O'DonnellICLR 2022 · 被引用 26 次
