S2ORC: The Semantic Scholar Open Research Corpus
Kyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney, Daniel S. Weld
Abstract
We introduce S2ORC, 1 a large corpus of 81.1M English-language academic papers spanning many academic disciplines. The corpus consists of rich metadata, paper abstracts, resolved bibliographic references, as well as structured full text for 8.1M open access papers. Full text is annotated with automaticallydetected inline mentions of citations, figures, and tables, each linked to their corresponding paper objects. In S2ORC, we aggregate papers from hundreds of academic publishers and digital archives into a unified source, and create the largest publicly-available collection of machine-readable academic text to date. We hope this resource will facilitate research and development of tools and tasks for text mining over academic text. * denotes equal contribution 1 Instructions for access to the data and model are available at https://github.com/allenai/s2orc/ . 2 https://arxiv.org Corpus Papers w/ body text Citation contexts References to tables / figures / equations Linked to graph Academic disciplines S2ORC (PDF-parse) 8.1M full text yes S2ORC (full) multi S2ORC (LATEX-parse) 1.5M full text yes S2ORC (full) physics, math, CS PubMed Central (OA) 2.6M full text yes PubMed bio, med AAN (Radev et al., 2009) 25k full text no ACL Anthology comp ling Saier and Färber (2019) † 1.0M snippets no MAG physics, math, CS RefSeer (Huang et al., 2015) 1.0M snippets no CiteSeerX multi
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5bd8c49d-7b39-4a02-8f1a-c561fcabd486Cited by top-tier papers105
- Merging Models with Fisher-Weighted AveragingMichael Matena, Colin RaffelNeurIPS 2022 · 741 citations
- Data Selection for Language Models via Importance ResamplingSang Michael Xie, Shibani Santurkar, Tengyu Ma, Percy LiangNeurIPS 2023 · 383 citations
- Nougat: Neural Optical Understanding for Academic DocumentsLukas Blecher, Guillem Cucurull, Thomas Scialom, Robert StojnicICLR 2024 · 243 citations
- What's In My Big Data?Yanai Elazar, Akshita Bhagia, Ian Magnusson, Abhilasha Ravichander et al.ICLR 2024 · 135 citations
- Retrieval meets Long Context Large Language ModelsPeng Xu, Wei Ping, Xianchao Wu, Lawrence McAfee et al.ICLR 2024 · 131 citations
Related papers
- The ACL OCL Corpus: Advancing Open Science in Computational LinguisticsShaurya Rohatgi, Yanxia Qin, Benjamin Aw, Niranjana Unnithan et al.EMNLP 2023 · 9 citations
- SchemaPile: A Large Collection of Relational Database SchemasTill Döhmen, Radu Geacu, Madelon Hulsebos, Sebastian SchelterSIGMOD 2024 · 9 citations
- What Should I Cite? A RAG Benchmark for Academic Citation PredictionLeqi Zheng, Jiajun Zhang, Canzhi Chen, Chaokun Wang et al.WWW 2026 · 2 citations
- MS2: Multi-Document Summarization of Medical StudiesJay DeYoung, Iz Beltagy, Madeleine van Zuylen, Bailey Kuehl et al.EMNLP 2021 · 83 citations
- S2abEL: A Dataset for Entity Linking from Scientific TablesYuze Lou, Bailey Kuehl, Erin Bransom, Sergey Feldman et al.EMNLP 2023 · 2 citations
