BiCA: Effective Biomedical Dense Retrieval with Citation-Aware Hard Negatives
Aarush Sinha, Pavan Kumar S, Roshan Balaji, Nirav Pravinbhai Bhatt
摘要
Hard negatives are essential for training effective retrieval models. Hard-negative mining typically relies on ranking documents using cross-encoders or static embedding models based on similarity metrics such as cosine distance. Hard negative mining becomes challenging for biomedical and scientific domains due to the difficulty in distinguishing between source and hard negative documents. However, referenced documents naturally share contextual relevance with the source document but are not duplicates, making them well-suited as hard negatives. In this work, we propose BiCA: Biomedical Dense Retrieval with Citation-Aware Hard Negatives, an approach for hard-negative mining by utilizing citation links in 20,000 PubMed articles for improving a domainspecific small dense retriever. We fine-tune the GTEsmall and GTEBase models using these citation-informed negatives and observe consistent improvements in zero-shot dense retrieval using nDCG@10 for both in-domain and out-of-domain tasks on BEIR and outperform baselines on long-tailed topics in LoTTE using Success@5. Our findings highlight the potential of leveraging document link structure to generate highly informative negatives, enabling state-of-the-art performance with minimal fine-tuning and demonstrating a path towards highly data-efficient domain adaptation. Code -github.com/bisect-group/BiCA Datasetshuggingface.co/collections/bisectgroup/bica-aaai26 * Worked done as a UG student and currently at
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Improving Multi-hop Question Answering over Knowledge Graphs using Knowledge Base EmbeddingsApoorv Saxena, Aditay Tripathi, Partha P. TalukdarACL 2020 · 被引用 488 次
- LinkBERT: Pretraining Language Models with Document LinksMichihiro Yasunaga, Jure Leskovec, Percy LiangACL 2022 · 被引用 463 次
- Deep Bidirectional Language-Knowledge Graph PretrainingMichihiro Yasunaga, Antoine Bosselut, Hongyu Ren, Xikun Zhang 等NeurIPS 2022 · 被引用 294 次
- Improving Biomedical Information Retrieval with Neural RetrieversMan Luo, Arindam Mitra, Tejas Gokhale, Chitta BaralAAAI 2022 · 被引用 42 次
- SPECTER: Document-level Representation Learning using Citation-informed TransformersArman Cohan, Sergey Feldman, Iz Beltagy, Doug Downey 等ACL 2020 · 被引用 20 次
相关 Paper
- BMRetriever: Tuning Large Language Models as Better Biomedical Text RetrieversRan Xu, Wenqi Shi, Yue Yu, Yuchen Zhuang 等EMNLP 2024 · 被引用 8 次
- BERM: Training the Balanced and Extractable Representation for Matching to Improve Generalization Ability of Dense RetrievalShicheng Xu, Liang Pang, Huawei Shen, Xueqi ChengACL 2023 · 被引用 8 次
- Constructing Hard-Positive Query-Document Pairs for Dense Retrieval via Phrase RepresentativenessZhanyu Wu, Richong Zhang, Zhijie NieSIGIR 2026
- Boosting Data Utilization for Multilingual Dense RetrievalChao Huang, Fengran Mo, Yufeng Chen, Changhao Guan 等EMNLP 2025 · 被引用 2 次
- COCO-DR: Combating the Distribution Shift in Zero-Shot Dense Retrieval with Contrastive and Distributionally Robust LearningYue Yu, Chenyan Xiong, Si Sun, Chao Zhang 等EMNLP 2022 · 被引用 21 次
