Efficient Performance Tracking: Leveraging Large Language Models for Automated Construction of Scientific Leaderboards
Furkan Sahinuç, Thy Thy Tran, Yulia Grishina, Yufang Hou, Bei Chen, Iryna Gurevych
Abstract
Scientific leaderboards are standardized ranking systems that facilitate evaluating and comparing competitive methods. Typically, a leaderboard is defined by a task, dataset, and evaluation metric (TDM) triple, allowing objective performance assessment and fostering innovation through benchmarking. However, the exponential increase in publications has made it infeasible to construct and maintain these leaderboards manually. Automatic leaderboard construction has emerged as a solution to reduce manual labor. Existing datasets for this task are based on the community-contributed leaderboards without additional curation. Our analysis shows that a large portion of these leaderboards are incomplete, and some of them contain incorrect information. In this work, we present SCILEAD, a manually-curated Scientific Leaderboard dataset that overcomes the aforementioned problems. Building on this dataset, we propose three experimental settings that simulate real-world scenarios where TDM triples are fully defined, partially defined, or undefined during leaderboard construction. While previous research has only explored the first setting, the latter two are more representative of real-world applications. To address these diverse settings, we develop a comprehensive LLM-based framework for constructing leaderboards. Our experiments and analysis reveal that various LLMs often correctly identify TDM triples while struggling to extract result values from publications. We make our code 1 and data 2 publicly available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c9e9c319-0c5a-4516-9d45-bde7fa79eba1Cited by top-tier papers4
- HalluCitation Matters: Revealing the Impact of Hallucinated References with 300 Hallucinated Papers in ACL ConferencesYusuke Sakai, Hidetaka Kamigaito, Taro WatanabeACL 2026 · 18 citations
- The Nature of NLP: Analyzing Contributions in NLP PapersAniket Pramanick, Yufang Hou, Saif M. Mohammad, Iryna GurevychACL 2025 · 9 citations
- Can LLMs Help Uncover Insights about LLMs? A Large-Scale, Evolving Literature Analysis of Frontier LLMsJungsoo Park, Junmo Kang, Gabriel Stanovsky, Alan RitterACL 2025 · 4 citations
- TaxoAlign: Scholarly Taxonomy Generation Using Language ModelsAvishek Lahiri, Yufang Hou, Debarshi Kumar SanyalEMNLP 2025
Builds on4
- Self-Adaptive In-Context Learning: An Information Compression Perspective for In-Context Example Selection and OrderingZhiyong Wu, Yaoxiang Wang, Jiacheng Ye, Lingpeng KongACL 2023 · 50 citations
- SciREX: A Challenge Dataset for Document-Level Information ExtractionSarthak Jain, Madeleine van Zuylen, Hannaneh Hajishirzi, Iz BeltagyACL 2020 · 9 citations
- AxCell: Automatic Extraction of Results from Machine Learning PapersMarcin Kardas, Piotr Czapla, Pontus Stenetorp, Sebastian Ruder et al.EMNLP 2020 · 5 citations
- A Diachronic Analysis of Paradigm Shifts in NLP Research: When, How, and Why?Aniket Pramanick, Yufang Hou, Saif M. Mohammad, Iryna GurevychEMNLP 2023 · 3 citations
Related papers
- A Position Paper on the Automatic Generation of Machine Learning LeaderboardsRoelien C. Timmer, Yufang Hou, Stephen WanEMNLP 2025
- End-to-End Argumentation Knowledge Graph ConstructionKhalid Al Khatib, Yufang Hou, Henning Wachsmuth, Charles Jochim et al.AAAI 2020 · 56 citations
- SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language ModelsXiaoxuan Wang, Ziniu Hu, Pan Lu, Yanqiao Zhu et al.ICML 2024 · 220 citations
- SciCustom: A Framework for Custom Evaluation of Scientific Capabilities in Large Language ModelsYiyang Gu, Junwei Yang, Junyu Luo, Ye Yuan et al.ACL 2026
- Evaluation Examples are not Equally Informative: How should that change NLP Leaderboards?Pedro Rodriguez, Joe Barrow, Alexander Miserlis Hoyle, John P. Lalor et al.ACL 2021
