Label-Efficient Model Selection for Text Generation
Shir Ashury-Tahan, Ariel Gera, Benjamin Sznajder, Leshem Choshen, Liat Ein-Dor, Eyal Shnarch
摘要
Model selection for a given target task can be costly, as it may entail extensive annotation of the quality of outputs of different models. We introduce DiffUse, an efficient method to make an informed decision between candidate text generation models based on preference annotations. DiffUse reduces the required amount of annotations, thus saving valuable time and resources in performing evaluation. DiffUse intelligently selects instances by clustering embeddings that represent the semantic differences between model outputs. Thus, it is able to identify a subset of examples that are more informative for preference decisions. Our method is model-agnostic, and can be applied to any text generation model for selecting between models, prompts and configurations. Moreover, we propose a practical iterative approach for dynamically determining how many instances to annotate. In a series of experiments over hundreds of model pairs, we demonstrate that DiffUse can dramatically reduce the required number of annotations -by up to 75% -while maintaining high evaluation reliability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Efficient multi-prompt evaluation of LLMsFelipe Maia Polo, Ronald Xu, Lucas Weber, Mírian Silva 等NeurIPS 2024 · 被引用 93 次
- Scaling Up Active Testing to Large Language ModelsGabrielle Berrada, Jannik Kossen, Freddie Bickford Smith, Muhammed Razzak 等NeurIPS 2025 · 被引用 11 次
- ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI EvaluationYizheng Huang, Wenjun Zeng, Aditi Kumaresan, Zi WangICML 2026 · 被引用 1 次
- JuStRank: Benchmarking LLM Judges for System RankingAriel Gera, Odellia Boni, Yotam Perlitz, Roy Bar-Haim 等ACL 2025
- Active Evaluation Acquisition for Efficient LLM BenchmarkingYang Li, Jie Ma, Miguel Ballesteros, Yassine Benajiba 等ICML 2025
它引用的顶会 Paper11
- Emergent and Predictable Memorization in Large Language ModelsStella Biderman, USVSN Sai Prashanth, Lintang Sutawika, Hailey Schoelkopf 等NeurIPS 2023 · 被引用 205 次
- Efficient multi-prompt evaluation of LLMsFelipe Maia Polo, Ronald Xu, Lucas Weber, Mírian Silva 等NeurIPS 2024 · 被引用 93 次
- Active Testing: Sample-Efficient Model EvaluationJannik Kossen, Sebastian Farquhar, Yarin Gal, Tom RainforthICML 2021 · 被引用 81 次
- Corpus Wide Argument Mining - A Working SolutionLiat Ein-Dor, Eyal Shnarch, Lena Dankin, Alon Halfon 等AAAI 2020 · 被引用 70 次
- A Survey of Active Learning for Natural Language ProcessingZhisong Zhang, Emma Strubell, Eduard H. HovyEMNLP 2022 · 被引用 60 次
相关 Paper
- FlashEval: Towards Fast and Accurate Evaluation of Text-to-Image Diffusion Generative ModelsLin Zhao, Tianchen Zhao, Zinan Lin, Xuefei Ning 等CVPR 2024 · 被引用 2 次
- DICE: Distilling Classifier-Free Guidance into Text EmbeddingsZhenyu Zhou, Defang Chen, Can Wang, Chun Chen 等AAAI 2026 · 被引用 2 次
- Reusing Computation in Text-to-Image Diffusion for Efficient Generation of Image SetsDale Decatur, Thibault Groueix, Wang Yifan, Rana Hanocka 等ICCV 2025
- DyMO: Training-Free Diffusion Model Alignment with Dynamic Multi-Objective SchedulingXin Xie, Dong GongCVPR 2025
- Self-Supervised Direct Preference Optimization for Text-to-Image Diffusion ModelsLiang Peng, Boxi Wu, Haoran Cheng, Yibo Zhao 等NeurIPS 2025 · 被引用 2 次
