Efficient Performance Tracking: Leveraging Large Language Models for Automated Construction of Scientific Leaderboards
Furkan Sahinuç, Thy Thy Tran, Yulia Grishina, Yufang Hou, Bei Chen, Iryna Gurevych
摘要
Scientific leaderboards are standardized ranking systems that facilitate evaluating and comparing competitive methods. Typically, a leaderboard is defined by a task, dataset, and evaluation metric (TDM) triple, allowing objective performance assessment and fostering innovation through benchmarking. However, the exponential increase in publications has made it infeasible to construct and maintain these leaderboards manually. Automatic leaderboard construction has emerged as a solution to reduce manual labor. Existing datasets for this task are based on the community-contributed leaderboards without additional curation. Our analysis shows that a large portion of these leaderboards are incomplete, and some of them contain incorrect information. In this work, we present SCILEAD, a manually-curated Scientific Leaderboard dataset that overcomes the aforementioned problems. Building on this dataset, we propose three experimental settings that simulate real-world scenarios where TDM triples are fully defined, partially defined, or undefined during leaderboard construction. While previous research has only explored the first setting, the latter two are more representative of real-world applications. To address these diverse settings, we develop a comprehensive LLM-based framework for constructing leaderboards. Our experiments and analysis reveal that various LLMs often correctly identify TDM triples while struggling to extract result values from publications. We make our code 1 and data 2 publicly available.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- HalluCitation Matters: Revealing the Impact of Hallucinated References with 300 Hallucinated Papers in ACL ConferencesYusuke Sakai, Hidetaka Kamigaito, Taro WatanabeACL 2026 · 被引用 18 次
- The Nature of NLP: Analyzing Contributions in NLP PapersAniket Pramanick, Yufang Hou, Saif M. Mohammad, Iryna GurevychACL 2025 · 被引用 9 次
- Can LLMs Help Uncover Insights about LLMs? A Large-Scale, Evolving Literature Analysis of Frontier LLMsJungsoo Park, Junmo Kang, Gabriel Stanovsky, Alan RitterACL 2025 · 被引用 4 次
- TaxoAlign: Scholarly Taxonomy Generation Using Language ModelsAvishek Lahiri, Yufang Hou, Debarshi Kumar SanyalEMNLP 2025
它引用的顶会 Paper4
- Self-Adaptive In-Context Learning: An Information Compression Perspective for In-Context Example Selection and OrderingZhiyong Wu, Yaoxiang Wang, Jiacheng Ye, Lingpeng KongACL 2023 · 被引用 50 次
- SciREX: A Challenge Dataset for Document-Level Information ExtractionSarthak Jain, Madeleine van Zuylen, Hannaneh Hajishirzi, Iz BeltagyACL 2020 · 被引用 9 次
- AxCell: Automatic Extraction of Results from Machine Learning PapersMarcin Kardas, Piotr Czapla, Pontus Stenetorp, Sebastian Ruder 等EMNLP 2020 · 被引用 5 次
- A Diachronic Analysis of Paradigm Shifts in NLP Research: When, How, and Why?Aniket Pramanick, Yufang Hou, Saif M. Mohammad, Iryna GurevychEMNLP 2023 · 被引用 3 次
相关 Paper
- A Position Paper on the Automatic Generation of Machine Learning LeaderboardsRoelien C. Timmer, Yufang Hou, Stephen WanEMNLP 2025
- End-to-End Argumentation Knowledge Graph ConstructionKhalid Al Khatib, Yufang Hou, Henning Wachsmuth, Charles Jochim 等AAAI 2020 · 被引用 56 次
- SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language ModelsXiaoxuan Wang, Ziniu Hu, Pan Lu, Yanqiao Zhu 等ICML 2024 · 被引用 220 次
- SciCustom: A Framework for Custom Evaluation of Scientific Capabilities in Large Language ModelsYiyang Gu, Junwei Yang, Junyu Luo, Ye Yuan 等ACL 2026
- Evaluation Examples are not Equally Informative: How should that change NLP Leaderboards?Pedro Rodriguez, Joe Barrow, Alexander Miserlis Hoyle, John P. Lalor 等ACL 2021
