SKILL-MIX: a Flexible and Expandable Family of Evaluations for AI Models
Dingli Yu, Simran Kaur, Arushi Gupta, Jonah Brown-Cohen, Anirudh Goyal, Sanjeev Arora
Abstract
With LLMs shifting their role from statistical modeling of language to serving as general-purpose AI agents, how should LLM evaluations change? Arguably, a key ability of an AI agent is to flexibly combine, as needed, the basic skills it has learned. The capability to combine skills plays an important role in (human) pedagogy and also in a paper on emergence phenomena (Arora & Goyal, 2023) . This work introduces SKILL-MIX, a new evaluation to measure ability to combine skills. Using a list of N skills the evaluator repeatedly picks random subsets of k skills and asks the LLM to produce text combining that subset of skills. Since the number of subsets grows like N k , for even modest k this evaluation will, with high probability, require the LLM to produce text significantly different from any text in the training set. The paper develops a methodology for (a) designing and administering such an evaluation, and (b) automatic grading (plus spot-checking by humans) of the results using GPT-4 as well as the open LLaMA-2 70B model. Administering a version of SKILL-MIX to popular chatbots gave results that, while generally in line with prior expectations, contained surprises. Sizeable differences exist among model capabilities that are not captured by their ranking on popular LLM leaderboards ("cramming for the leaderboard"). Furthermore, simple probability calculations indicate that GPT-4's reasonable performance on k = 5 is suggestive of going beyond "stochastic parrot" behavior (Bender et al., 2021) , i.e., it combines skills in ways that it had not seen during training. We sketch how the methodology can lead to a SKILL-MIX based eco-system of open evaluations for AI capabilities of future models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 15a5a49f-6a13-42e2-a193-0dfdf46bd24cCited by top-tier papers32
- QuRating: Selecting High-Quality Data for Training Language ModelsAlexander Wettig, Aatmik Gupta, Saumya Malik, Danqi ChenICML 2024 · 138 citations
- Metacognitive Capabilities of LLMs: An Exploration in Mathematical Problem SolvingAniket Didolkar, Anirudh Goyal, Nan Rosemary Ke, Siyuan Guo et al.NeurIPS 2024 · 101 citations
- Can large language models explore in-context?Akshay Krishnamurthy, Keegan Harris, Dylan J. Foster, Cyril Zhang et al.NeurIPS 2024 · 95 citations
- DARG: Dynamic Evaluation of Large Language Models via Adaptive Reasoning GraphZhehao Zhang, Jiaao Chen, Diyi YangNeurIPS 2024 · 42 citations
- How Benchmark Prediction from Fewer Data Misses the MarkGuanhua Zhang, Florian E. Dorner, Moritz HardtNeurIPS 2025 · 26 citations
Builds on2
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Skill-it! A data-driven skills framework for understanding and training language modelsMayee F. Chen, Nicholas Roberts, Kush Bhatia, Jue Wang et al.NeurIPS 2023 · 143 citations
Related papers
- Can Models Learn Skill Composition from Examples?Haoyu Zhao, Simran Kaur, Dingli Yu, Anirudh Goyal et al.NeurIPS 2024 · 20 citations
- SkillAggregation: Reference-free LLM-Dependent AggregationGuangzhi Sun, Anmol Kagrecha, Potsawee Manakul, Philip C. Woodland et al.ACL 2025 · 5 citations
- MixEval: Deriving Wisdom of the Crowd from LLM Benchmark MixturesJinjie Ni, Fuzhao Xue, Xiang Yue, Yuntian Deng et al.NeurIPS 2024 · 88 citations
- Explore Theory of Mind: program-guided adversarial data generation for theory of mind reasoningMelanie Sclar, Jane Dwivedi-Yu, Maryam Fazel-Zarandi, Yulia Tsvetkov et al.ICLR 2025
- Mixture-of-Agents Enhances Large Language Model CapabilitiesJunlin Wang, Jue Wang, Ben Athiwaratkun, Ce Zhang et al.ICLR 2025
