S2ORC: The Semantic Scholar Open Research Corpus
Kyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney, Daniel S. Weld
摘要
We introduce S2ORC, 1 a large corpus of 81.1M English-language academic papers spanning many academic disciplines. The corpus consists of rich metadata, paper abstracts, resolved bibliographic references, as well as structured full text for 8.1M open access papers. Full text is annotated with automaticallydetected inline mentions of citations, figures, and tables, each linked to their corresponding paper objects. In S2ORC, we aggregate papers from hundreds of academic publishers and digital archives into a unified source, and create the largest publicly-available collection of machine-readable academic text to date. We hope this resource will facilitate research and development of tools and tasks for text mining over academic text. * denotes equal contribution 1 Instructions for access to the data and model are available at https://github.com/allenai/s2orc/ . 2 https://arxiv.org Corpus Papers w/ body text Citation contexts References to tables / figures / equations Linked to graph Academic disciplines S2ORC (PDF-parse) 8.1M full text yes S2ORC (full) multi S2ORC (LATEX-parse) 1.5M full text yes S2ORC (full) physics, math, CS PubMed Central (OA) 2.6M full text yes PubMed bio, med AAN (Radev et al., 2009) 25k full text no ACL Anthology comp ling Saier and Färber (2019) † 1.0M snippets no MAG physics, math, CS RefSeer (Huang et al., 2015) 1.0M snippets no CiteSeerX multi
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper105
- Merging Models with Fisher-Weighted AveragingMichael Matena, Colin RaffelNeurIPS 2022 · 被引用 741 次
- Data Selection for Language Models via Importance ResamplingSang Michael Xie, Shibani Santurkar, Tengyu Ma, Percy LiangNeurIPS 2023 · 被引用 383 次
- Nougat: Neural Optical Understanding for Academic DocumentsLukas Blecher, Guillem Cucurull, Thomas Scialom, Robert StojnicICLR 2024 · 被引用 243 次
- What's In My Big Data?Yanai Elazar, Akshita Bhagia, Ian Magnusson, Abhilasha Ravichander 等ICLR 2024 · 被引用 135 次
- Retrieval meets Long Context Large Language ModelsPeng Xu, Wei Ping, Xianchao Wu, Lawrence McAfee 等ICLR 2024 · 被引用 131 次
相关 Paper
- The ACL OCL Corpus: Advancing Open Science in Computational LinguisticsShaurya Rohatgi, Yanxia Qin, Benjamin Aw, Niranjana Unnithan 等EMNLP 2023 · 被引用 9 次
- SchemaPile: A Large Collection of Relational Database SchemasTill Döhmen, Radu Geacu, Madelon Hulsebos, Sebastian SchelterSIGMOD 2024 · 被引用 9 次
- What Should I Cite? A RAG Benchmark for Academic Citation PredictionLeqi Zheng, Jiajun Zhang, Canzhi Chen, Chaokun Wang 等WWW 2026 · 被引用 2 次
- MS2: Multi-Document Summarization of Medical StudiesJay DeYoung, Iz Beltagy, Madeleine van Zuylen, Bailey Kuehl 等EMNLP 2021 · 被引用 83 次
- S2abEL: A Dataset for Entity Linking from Scientific TablesYuze Lou, Bailey Kuehl, Erin Bransom, Sergey Feldman 等EMNLP 2023 · 被引用 2 次
