Neighborhood Contrastive Learning for Scientific Document Representations with Citation Embeddings
Malte Ostendorff, Nils Rethmeier, Isabelle Augenstein, Bela Gipp, Georg Rehm
摘要
Learning scientific document representations can be substantially improved through contrastive learning objectives, where the challenge lies in creating positive and negative training samples that encode the desired similarity semantics. Prior work relies on discrete citation relations to generate contrast samples. However, discrete citations enforce a hard cutoff to similarity. This is counter-intuitive to similarity-based learning and ignores that scientific papers can be very similar despite lacking a direct citation -a core problem of finding related research. Instead, we use controlled nearest neighbor sampling over citation graph embeddings for contrastive learning. This control allows us to learn continuous similarity, to sample hard-to-learn negatives and positives, and also to avoid collisions between negative and positive samples by controlling the sampling margin between them. The resulting method SciNCL outperforms the state-of-theart on the SciDocs benchmark. Furthermore, we demonstrate that it can train (or tune) language models sample-efficiently and that it can be combined with recent training-efficient methods. Perhaps surprisingly, even training a general-domain language model this way outperforms baselines pretrained in-domain. Related Work Contrastive Learning pulls representations of similar data points (positives) closer together, while representations of dissimilar documents (negatives) are pushed apart. A common contrastive objective
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- SciRepEval: A Multi-Format Benchmark for Scientific Document RepresentationsAmanpreet Singh, Mike D'Arcy, Arman Cohan, Doug Downey 等EMNLP 2023 · 被引用 45 次
- A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific DiscoveryYu Zhang, Xiusi Chen, Bowen Jin, Sheng Wang 等EMNLP 2024 · 被引用 28 次
- HLM-Cite: Hybrid Language Model Workflow for Text-based Scientific Citation PredictionQianyue Hao, Jingyang Fan, Fengli Xu, Jian Yuan 等NeurIPS 2024 · 被引用 23 次
- Chain-of-Factors Paper-Reviewer MatchingYu Zhang, Yanzhen Shen, SeongKu Kang, Xiusi Chen 等WWW 2025 · 被引用 12 次
- Topic-Guided Sampling For Data-Efficient Multi-Domain Stance DetectionErik Arakelyan, Arnav Arora, Isabelle AugensteinACL 2023 · 被引用 9 次
它引用的顶会 Paper13
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna 等NeurIPS 2020 · 被引用 7,049 次
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 被引用 2,360 次
- What Makes for Good Views for Contrastive Learning?Yonglong Tian, Chen Sun, Ben Poole, Dilip Krishnan 等NeurIPS 2020 · 被引用 1,631 次
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang 等ICLR 2021 · 被引用 1,547 次
相关 Paper
- SemCSE: Semantic Contrastive Sentence Embeddings Using LLM-Generated Summaries For Scientific AbstractsMarc Felix Brinner, Sina ZarrießEMNLP 2025
- SPECTER: Document-level Representation Learning using Citation-informed TransformersArman Cohan, Sergey Feldman, Iz Beltagy, Doug Downey 等ACL 2020 · 被引用 20 次
- Global Selection of Contrastive Batches via Optimization on Sample PermutationsVin Sachidananda, Ziyi Yang, Chenguang ZhuICML 2023 · 被引用 6 次
- Language Models as Semantic IndexersBowen Jin, Hansi Zeng, Guoyin Wang, Xiusi Chen 等ICML 2024 · 被引用 31 次
- DOGR: Leveraging Document-Oriented Contrastive Learning in Generative RetrievalPenghao Lu, Xin Dong, Yuansheng Zhou, Lei Cheng 等AAAI 2025
