AxCell: Automatic Extraction of Results from Machine Learning Papers
Marcin Kardas, Piotr Czapla, Pontus Stenetorp, Sebastian Ruder, Sebastian Riedel, Ross Taylor, Robert Stojnic
摘要
Tracking progress in machine learning has become increasingly difficult with the recent explosion in the number of papers. In this paper, we present AXCELL, an automatic machine learning pipeline for extracting results from papers. AXCELL uses several novel components, including a table segmentation subtask, to learn relevant structural knowledge that aids extraction. When compared with existing methods, our approach significantly improves the state of the art for results extraction. We also release a structured, annotated dataset for training models for results extraction, and a dataset for evaluating the performance of models on this task. Lastly, we show the viability of our approach enables it to be used for semi-automated results extraction in production, suggesting our improvements make this task practically viable for the first time. Code is available on GitHub. 1 Back-translation . . .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- TUTA: Tree-based Transformers for Generally Structured Table Pre-trainingZhiruo Wang, Haoyu Dong, Ran Jia, Jia Li 等KDD 2021 · 被引用 88 次
- MATE: Multi-view Attention for Table Transformer EfficiencyJulian Martin Eisenschlos, Maharshi Gor, Thomas Müller, William W. CohenEMNLP 2021 · 被引用 62 次
- DocLLM: A Layout-Aware Generative Language Model for Multimodal Document UnderstandingDongsheng Wang, Natraj Raman, Mathieu Sibue, Zhiqiang Ma 等ACL 2024 · 被引用 37 次
- Can LLMs Help Uncover Insights about LLMs? A Large-Scale, Evolving Literature Analysis of Frontier LLMsJungsoo Park, Junmo Kang, Gabriel Stanovsky, Alan RitterACL 2025 · 被引用 4 次
- SciRIFF: A Resource to Enhance Language Model Instruction-Following over Scientific LiteratureDavid Wadden, Kejian Shi, Jacob Morrison, Alan Li 等EMNLP 2025 · 被引用 2 次
它引用的顶会 Paper1
相关 Paper
- S2abEL: A Dataset for Entity Linking from Scientific TablesYuze Lou, Bailey Kuehl, Erin Bransom, Sergey Feldman 等EMNLP 2023 · 被引用 2 次
- PubTables-1M: Towards comprehensive table extraction from unstructured documentsBrandon Smock, Rohith Pesala, Robin AbrahamCVPR 2022 · 被引用 125 次
- AirQA: A Comprehensive QA Dataset for AI Research with Instance-Level EvaluationTiancheng Huang, Ruisheng Cao, Yuxin Zhang, Zhangyi Kang 等ICLR 2026 · 被引用 1 次
- Experimental Evidence Extraction System in Data Science with Hybrid Table Features and Ensemble LearningWenhao Yu, Wei Peng, Yu Shu, Qingkai Zeng 等WWW 2020 · 被引用 10 次
- Cell2Doc: ML Pipeline for Generating Documentation in Computational NotebooksTamal Mondal, Scott Barnett, Akash Lal, Jyothi VeduradaASE 2023 · 被引用 4 次
