Label-Efficient Model Selection for Text Generation
Shir Ashury-Tahan, Ariel Gera, Benjamin Sznajder, Leshem Choshen, Liat Ein-Dor, Eyal Shnarch
Abstract
Model selection for a given target task can be costly, as it may entail extensive annotation of the quality of outputs of different models. We introduce DiffUse, an efficient method to make an informed decision between candidate text generation models based on preference annotations. DiffUse reduces the required amount of annotations, thus saving valuable time and resources in performing evaluation. DiffUse intelligently selects instances by clustering embeddings that represent the semantic differences between model outputs. Thus, it is able to identify a subset of examples that are more informative for preference decisions. Our method is model-agnostic, and can be applied to any text generation model for selecting between models, prompts and configurations. Moreover, we propose a practical iterative approach for dynamically determining how many instances to annotate. In a series of experiments over hundreds of model pairs, we demonstrate that DiffUse can dramatically reduce the required number of annotations -by up to 75% -while maintaining high evaluation reliability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6474bdeb-fca5-4ac3-99e1-e50ba35c93beCited by top-tier papers5
- Efficient multi-prompt evaluation of LLMsFelipe Maia Polo, Ronald Xu, Lucas Weber, Mírian Silva et al.NeurIPS 2024 · 93 citations
- Scaling Up Active Testing to Large Language ModelsGabrielle Berrada, Jannik Kossen, Freddie Bickford Smith, Muhammed Razzak et al.NeurIPS 2025 · 11 citations
- ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI EvaluationYizheng Huang, Wenjun Zeng, Aditi Kumaresan, Zi WangICML 2026 · 1 citation
- JuStRank: Benchmarking LLM Judges for System RankingAriel Gera, Odellia Boni, Yotam Perlitz, Roy Bar-Haim et al.ACL 2025
- Active Evaluation Acquisition for Efficient LLM BenchmarkingYang Li, Jie Ma, Miguel Ballesteros, Yassine Benajiba et al.ICML 2025
Builds on11
- Emergent and Predictable Memorization in Large Language ModelsStella Biderman, USVSN Sai Prashanth, Lintang Sutawika, Hailey Schoelkopf et al.NeurIPS 2023 · 205 citations
- Efficient multi-prompt evaluation of LLMsFelipe Maia Polo, Ronald Xu, Lucas Weber, Mírian Silva et al.NeurIPS 2024 · 93 citations
- Active Testing: Sample-Efficient Model EvaluationJannik Kossen, Sebastian Farquhar, Yarin Gal, Tom RainforthICML 2021 · 81 citations
- Corpus Wide Argument Mining - A Working SolutionLiat Ein-Dor, Eyal Shnarch, Lena Dankin, Alon Halfon et al.AAAI 2020 · 70 citations
- A Survey of Active Learning for Natural Language ProcessingZhisong Zhang, Emma Strubell, Eduard H. HovyEMNLP 2022 · 60 citations
Related papers
- FlashEval: Towards Fast and Accurate Evaluation of Text-to-Image Diffusion Generative ModelsLin Zhao, Tianchen Zhao, Zinan Lin, Xuefei Ning et al.CVPR 2024 · 2 citations
- DICE: Distilling Classifier-Free Guidance into Text EmbeddingsZhenyu Zhou, Defang Chen, Can Wang, Chun Chen et al.AAAI 2026 · 2 citations
- Reusing Computation in Text-to-Image Diffusion for Efficient Generation of Image SetsDale Decatur, Thibault Groueix, Wang Yifan, Rana Hanocka et al.ICCV 2025
- DyMO: Training-Free Diffusion Model Alignment with Dynamic Multi-Objective SchedulingXin Xie, Dong GongCVPR 2025
- Self-Supervised Direct Preference Optimization for Text-to-Image Diffusion ModelsLiang Peng, Boxi Wu, Haoran Cheng, Yibo Zhao et al.NeurIPS 2025 · 2 citations
