PhyloLM: Inferring the Phylogeny of Large Language Models and Predicting their Performances in Benchmarks
Nicolas Yax, Pierre-Yves Oudeyer, Stefano Palminteri
摘要
This paper introduces PhyloLM, a method adapting phylogenetic algorithms to Large Language Models (LLMs) to explore whether and how they relate to each other and to predict their performance characteristics. Our method calculates a phylogenetic distance metric based on the similarity of LLMs' output. The resulting metric is then used to construct dendrograms, which satisfactorily capture known relationships across a set of 111 open-source and 45 closed models. Furthermore, our phylogenetic distance predicts performance in standard benchmarks, thus demonstrating its functional validity and paving the way for a time and costeffective estimation of LLM capabilities. To sum up, by translating population genetic concepts to machine learning, we propose and validate a tool to evaluate LLM development, relationships and capabilities, even in the absence of transparent training information.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- LLM DNA: Tracing Model Evolution via Functional RepresentationsZhaomin Wu, Haodong Zhao, Ziyang Wang, Jizhou Guo 等ICLR 2026 · 被引用 21 次
- Language Statistics and False Belief Reasoning: Evidence from 41 Open-Weight LMsSean Trott, Samuel M. Taylor, Cameron Robert Jones, James A. Michaelov 等ACL 2026 · 被引用 2 次
- Independence Tests for Language ModelsSally Zhu, Ahmed M. Ahmed, Rohith Kuditipudi, Percy LiangICML 2025
它引用的顶会 Paper5
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- OpenWebMath: An Open Dataset of High-Quality Mathematical Web TextKeiran Paster, Marco Dos Santos, Zhangir Azerbayev, Jimmy BaICLR 2024 · 被引用 140 次
- Multi-lingual Evaluation of Code Generation ModelsBen Athiwaratkun, Sanjay Krishna Gouda, Zijian Wang, Xiaopeng Li 等ICLR 2023 · 被引用 28 次
相关 Paper
- End-to-End Ontology Learning with Large Language ModelsAndy Lo, Albert Q. Jiang, Wenda Li, Mateja JamnikNeurIPS 2024 · 被引用 33 次
- PhyloGen: Language Model-Enhanced Phylogenetic Inference via Graph Structure GenerationChenrui Duan, Zelin Zang, Siyuan Li, Yongjie Xu 等NeurIPS 2024 · 被引用 8 次
- Neural Phylogeny: Fine-Tuning Relationship Detection among Neural NetworksRunpeng Yu, Xinchao WangICLR 2025
- Spectral Signatures of Large Language ModelsZhuoying Zhang, Ishan V. Prasad, Yuanzhe Hu, Zihang Liu 等KDD 2026
- QuanBench: Benchmarking Quantum Code Generation with Large Language ModelsXiaoyu Guo, Minggu Wang, Jianjun ZhaoASE 2025 · 被引用 5 次
