SPECTER: Document-level Representation Learning using Citation-informed Transformers
Arman Cohan, Sergey Feldman, Iz Beltagy, Doug Downey, Daniel S. Weld
Abstract
Representation learning is a critical ingredient for natural language processing systems. Recent Transformer language models like BERT learn powerful textual representations, but these models are targeted towards token-and sentence-level training objectives and do not leverage information on inter-document relatedness, which limits their document-level representation power. For applications on scientific documents, such as classification and recommendation, the embeddings power strong performance on end tasks. We propose SPECTER, a new method to generate document-level embedding of scientific documents based on pretraining a Transformer language model on a powerful signal of document-level relatedness: the citation graph. Unlike existing pretrained language models, SPECTER can be easily applied to downstream applications without task-specific fine-tuning. Additionally, to encourage further research on document-level models, we introduce SCIDOCS, a new evaluation benchmark consisting of seven document-level tasks ranging from citation prediction, to document classification and recommendation. We show that SPECTER outperforms a variety of competitive baselines on the benchmark. 1 * Equal contribution 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 95da3018-242b-4def-9354-e205a1b671b1Cited by top-tier papers117
- LinkBERT: Pretraining Language Models with Document LinksMichihiro Yasunaga, Jure Leskovec, Percy LiangACL 2022 · 463 citations
- MS2: Multi-Document Summarization of Medical StudiesJay DeYoung, Iz Beltagy, Madeleine van Zuylen, Bailey Kuehl et al.EMNLP 2021 · 83 citations
- MUVERA: Multi-Vector Retrieval via Fixed Dimensional EncodingLaxman Dhulipala, Majid Hadian, Rajesh Jayaram, Jason Lee et al.NeurIPS 2024 · 56 citations
- Threddy: An Interactive System for Personalized Thread-based Exploration and Organization of Scientific LiteratureHyeonsu B. Kang, Joseph Chee Chang, Yongsung Kim, Aniket KitturUIST 2022 · 48 citations
- Neighborhood Contrastive Learning for Scientific Document Representations with Citation EmbeddingsMalte Ostendorff, Nils Rethmeier, Isabelle Augenstein, Bela Gipp et al.EMNLP 2022 · 46 citations
Related papers
- SciRepEval: A Multi-Format Benchmark for Scientific Document RepresentationsAmanpreet Singh, Mike D'Arcy, Arman Cohan, Doug Downey et al.EMNLP 2023 · 45 citations
- Span Graph Transformer for Document-Level Named Entity RecognitionHongli Mao, Xian-Ling Mao, Hanlin Tang, Yuming Shang et al.AAAI 2024 · 3 citations
- SemCSE: Semantic Contrastive Sentence Embeddings Using LLM-Generated Summaries For Scientific AbstractsMarc Felix Brinner, Sina ZarrießEMNLP 2025
- Content- and Topology-Aware Representation Learning for Scientific Multi-LiteratureKai Zhang, Kaisong Song, Yangyang Kang, Xiaozhong LiuEMNLP 2023
- SLM: Learning a Discourse Language Representation with Sentence UnshufflingHaejun Lee, Drew A. Hudson, Kangwook Lee, Christopher D. ManningEMNLP 2020 · 2 citations
