SciRepEval: A Multi-Format Benchmark for Scientific Document Representations
Amanpreet Singh, Mike D'Arcy, Arman Cohan, Doug Downey, Sergey Feldman
Abstract
Learned representations of scientific documents can serve as valuable input features for downstream tasks without further fine-tuning. However, existing benchmarks for evaluating these representations fail to capture the diversity of relevant tasks. In response, we introduce SciRepEval, the first comprehensive benchmark for training and evaluating scientific document representations. It includes 24 challenging and realistic tasks, 8 of which are new, across four formats: classification, regression, ranking and search. We then use this benchmark to study and improve the generalization ability of scientific document representation models. We show how state-of-the-art models like SPECTER and SciNCL struggle to generalize across the task formats, and that simple multi-task training fails to improve them. However, a new approach that learns multiple embeddings per document, each tailored to a different format, can improve performance. We experiment with task-format-specific control codes and adapters and find they outperform the existing single-embedding state-of-the-art by over 2 points absolute. We release the resulting family of multi-format models, called SPECTER2, for the community to use and build on.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers25
- A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific DiscoveryYu Zhang, Xiusi Chen, Bowen Jin, Sheng Wang et al.EMNLP 2024 · 28 citations
- HLM-Cite: Hybrid Language Model Workflow for Text-based Scientific Citation PredictionQianyue Hao, Jingyang Fan, Fengli Xu, Jian Yuan et al.NeurIPS 2024 · 23 citations
- PaperWeaver: Enriching Topical Paper Alerts by Contextualizing Recommended Papers with User-collected PapersYoonjoo Lee, Hyeonsu B. Kang, Matt Latzke, Juho Kim et al.CHI 2024 · 20 citations
- Chain-of-Factors Paper-Reviewer MatchingYu Zhang, Yanzhen Shen, SeongKu Kang, Xiusi Chen et al.WWW 2025 · 12 citations
- BMRetriever: Tuning Large Language Models as Better Biomedical Text RetrieversRan Xu, Wenqi Shi, Yue Yu, Yuchen Zhuang et al.EMNLP 2024 · 8 citations
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach et al.ICLR 2022 · 1,976 citations
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu et al.NeurIPS 2020 · 1,957 citations
- LinkBERT: Pretraining Language Models with Document LinksMichihiro Yasunaga, Jure Leskovec, Percy LiangACL 2022 · 463 citations
Related papers
- SPECTER: Document-level Representation Learning using Citation-informed TransformersArman Cohan, Sergey Feldman, Iz Beltagy, Doug Downey et al.ACL 2020 · 20 citations
- SemCSE: Semantic Contrastive Sentence Embeddings Using LLM-Generated Summaries For Scientific AbstractsMarc Felix Brinner, Sina ZarrießEMNLP 2025
- MAIR: A Massive Benchmark for Evaluating Instructed RetrievalWeiwei Sun, Zhengliang Shi, Wu Long, Lingyong Yan et al.EMNLP 2024 · 1 citation
- XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual GeneralisationJunjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig et al.ICML 2020 · 1,132 citations
- AOEB: Benchmarking Agent-Oriented Multimodal EmbeddingsXin Zhang, Jiaxin Xu, mengjia zhou, Xinping Zhao et al.ICML 2026
