Learning More from Less: Unlocking Internal Representations for Benchmark Compression
Yueqi Zhang, Jin Hu, Shaoxiong Feng, Peiwen Yuan, Xinglin Wang, Yiwei Li, Jiayi Shi, Chuyi Tan, Ji Zhang, Boyuan Pan, Yao Hu, Kan Li
Abstract
The prohibitive cost of evaluating Large Language Models (LLMs) necessitates efficient alternatives to full-scale benchmarking. Prevalent approaches address this by identifying a small coreset of items to approximate full-benchmark performance. However, existing methods must estimate a reliable item profile from response patterns across many source models, which becomes statistically unstable when the source pool is small. This dependency is particularly limiting for newly released benchmarks with minimal historical evaluation data. We argue that discrete correctness labels are a lossy view of the model's decision process and fail to capture information encoded in hidden states. To address this, we introduce RepCore, which aligns heterogeneous hidden states into a unified latent space to construct representative coresets. Using these subsets for performance extrapolation, RepCore achieves precise estimation accuracy with as few as ten source models. Experiments on five benchmarks and over 200 models show consistent gains over output-based baselines in ranking correlation and estimation accuracy. Spectral analysis further indicates that the aligned representations contain separable components reflecting broad response tendencies and task-specific reasoning patterns.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on4
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister et al.NeurIPS 2023 · 1,549 citations
- tinyBenchmarks: evaluating LLMs with fewer examplesFelipe Maia Polo, Lucas Weber, Leshem Choshen, Yuekai Sun et al.ICML 2024 · 212 citations
- Discovering Latent Knowledge in Language Models Without SupervisionCollin Burns, Haotian Ye, Dan Klein, Jacob SteinhardtICLR 2023 · 45 citations
- Large Language Models Have Intrinsic Meta-Cognition, but Need a Good LensZiyang Ma, Qingyue Yuan, Zhenglin Wang, Deyu ZhouEMNLP 2025 · 7 citations
Related papers
- metabench - A Sparse Benchmark of Reasoning and Knowledge in Large Language ModelsAlexander Kipnis, Konstantinos Voudouris, Luca M. Schulze Buschoff, Eric SchulzICLR 2025
- MetaEval: Measuring the Discrimination of Benchmarks for Efficient LLM EvaluationZhuo Wang, Wen Wu, Guoqing Wang, Guangze Ye et al.AAAI 2026 · 1 citation
- SubLIME: Subset Selection via Rank Correlation Prediction for Data-Efficient LLM EvaluationGayathri Saranathan, Cong Xu, Mahammad Parwez Alam, Tarun Kumar et al.ACL 2025 · 3 citations
- How Benchmark Prediction from Fewer Data Misses the MarkGuanhua Zhang, Florian E. Dorner, Moritz HardtNeurIPS 2025 · 26 citations
- Efficient multi-prompt evaluation of LLMsFelipe Maia Polo, Ronald Xu, Lucas Weber, Mírian Silva et al.NeurIPS 2024 · 93 citations
