SPECTER: Document-level Representation Learning using Citation-informed Transformers
Arman Cohan, Sergey Feldman, Iz Beltagy, Doug Downey, Daniel S. Weld
摘要
Representation learning is a critical ingredient for natural language processing systems. Recent Transformer language models like BERT learn powerful textual representations, but these models are targeted towards token-and sentence-level training objectives and do not leverage information on inter-document relatedness, which limits their document-level representation power. For applications on scientific documents, such as classification and recommendation, the embeddings power strong performance on end tasks. We propose SPECTER, a new method to generate document-level embedding of scientific documents based on pretraining a Transformer language model on a powerful signal of document-level relatedness: the citation graph. Unlike existing pretrained language models, SPECTER can be easily applied to downstream applications without task-specific fine-tuning. Additionally, to encourage further research on document-level models, we introduce SCIDOCS, a new evaluation benchmark consisting of seven document-level tasks ranging from citation prediction, to document classification and recommendation. We show that SPECTER outperforms a variety of competitive baselines on the benchmark. 1 * Equal contribution 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper117
- LinkBERT: Pretraining Language Models with Document LinksMichihiro Yasunaga, Jure Leskovec, Percy LiangACL 2022 · 被引用 463 次
- MS2: Multi-Document Summarization of Medical StudiesJay DeYoung, Iz Beltagy, Madeleine van Zuylen, Bailey Kuehl 等EMNLP 2021 · 被引用 83 次
- MUVERA: Multi-Vector Retrieval via Fixed Dimensional EncodingLaxman Dhulipala, Majid Hadian, Rajesh Jayaram, Jason Lee 等NeurIPS 2024 · 被引用 56 次
- Threddy: An Interactive System for Personalized Thread-based Exploration and Organization of Scientific LiteratureHyeonsu B. Kang, Joseph Chee Chang, Yongsung Kim, Aniket KitturUIST 2022 · 被引用 48 次
- Neighborhood Contrastive Learning for Scientific Document Representations with Citation EmbeddingsMalte Ostendorff, Nils Rethmeier, Isabelle Augenstein, Bela Gipp 等EMNLP 2022 · 被引用 46 次
相关 Paper
- SciRepEval: A Multi-Format Benchmark for Scientific Document RepresentationsAmanpreet Singh, Mike D'Arcy, Arman Cohan, Doug Downey 等EMNLP 2023 · 被引用 45 次
- Span Graph Transformer for Document-Level Named Entity RecognitionHongli Mao, Xian-Ling Mao, Hanlin Tang, Yuming Shang 等AAAI 2024 · 被引用 3 次
- SemCSE: Semantic Contrastive Sentence Embeddings Using LLM-Generated Summaries For Scientific AbstractsMarc Felix Brinner, Sina ZarrießEMNLP 2025
- Content- and Topology-Aware Representation Learning for Scientific Multi-LiteratureKai Zhang, Kaisong Song, Yangyang Kang, Xiaozhong LiuEMNLP 2023
- SLM: Learning a Discourse Language Representation with Sentence UnshufflingHaejun Lee, Drew A. Hudson, Kangwook Lee, Christopher D. ManningEMNLP 2020 · 被引用 2 次
