PhyloLM: Inferring the Phylogeny of Large Language Models and Predicting their Performances in Benchmarks
Nicolas Yax, Pierre-Yves Oudeyer, Stefano Palminteri
Abstract
This paper introduces PhyloLM, a method adapting phylogenetic algorithms to Large Language Models (LLMs) to explore whether and how they relate to each other and to predict their performance characteristics. Our method calculates a phylogenetic distance metric based on the similarity of LLMs' output. The resulting metric is then used to construct dendrograms, which satisfactorily capture known relationships across a set of 111 open-source and 45 closed models. Furthermore, our phylogenetic distance predicts performance in standard benchmarks, thus demonstrating its functional validity and paving the way for a time and costeffective estimation of LLM capabilities. To sum up, by translating population genetic concepts to machine learning, we propose and validate a tool to evaluate LLM development, relationships and capabilities, even in the absence of transparent training information.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1e284509-6654-4af0-9173-e8cb7ab4de99Cited by top-tier papers3
- LLM DNA: Tracing Model Evolution via Functional RepresentationsZhaomin Wu, Haodong Zhao, Ziyang Wang, Jizhou Guo et al.ICLR 2026 · 21 citations
- Language Statistics and False Belief Reasoning: Evidence from 41 Open-Weight LMsSean Trott, Samuel M. Taylor, Cameron Robert Jones, James A. Michaelov et al.ACL 2026 · 2 citations
- Independence Tests for Language ModelsSally Zhu, Ahmed M. Ahmed, Rohith Kuditipudi, Percy LiangICML 2025
Builds on5
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- OpenWebMath: An Open Dataset of High-Quality Mathematical Web TextKeiran Paster, Marco Dos Santos, Zhangir Azerbayev, Jimmy BaICLR 2024 · 140 citations
- Multi-lingual Evaluation of Code Generation ModelsBen Athiwaratkun, Sanjay Krishna Gouda, Zijian Wang, Xiaopeng Li et al.ICLR 2023 · 28 citations
Related papers
- End-to-End Ontology Learning with Large Language ModelsAndy Lo, Albert Q. Jiang, Wenda Li, Mateja JamnikNeurIPS 2024 · 33 citations
- PhyloGen: Language Model-Enhanced Phylogenetic Inference via Graph Structure GenerationChenrui Duan, Zelin Zang, Siyuan Li, Yongjie Xu et al.NeurIPS 2024 · 8 citations
- Neural Phylogeny: Fine-Tuning Relationship Detection among Neural NetworksRunpeng Yu, Xinchao WangICLR 2025
- Spectral Signatures of Large Language ModelsZhuoying Zhang, Ishan V. Prasad, Yuanzhe Hu, Zihang Liu et al.KDD 2026
- QuanBench: Benchmarking Quantum Code Generation with Large Language ModelsXiaoyu Guo, Minggu Wang, Jianjun ZhaoASE 2025 · 5 citations
