HLM-Cite: Hybrid Language Model Workflow for Text-based Scientific Citation Prediction
Qianyue Hao, Jingyang Fan, Fengli Xu, Jian Yuan, Yong Li
Abstract
Citation networks are critical in modern science, and predicting which previous papers (candidates) will a new paper (query) cite is a critical problem. However, the roles of a paper's citations vary significantly, ranging from foundational knowledge basis to superficial contexts. Distinguishing these roles requires a deeper understanding of the logical relationships among papers, beyond simple edges in citation networks. The emergence of LLMs with textual reasoning capabilities offers new possibilities for discerning these relationships, but there are two major challenges. First, in practice, a new paper may select its citations from gigantic existing papers, where the texts exceed the context length of LLMs. Second, logical relationships between papers are implicit, and directly prompting an LLM to predict citations may result in surface-level textual similarities rather than the deeper logical reasoning. In this paper, we introduce the novel concept of core citation, which identifies the critical references that go beyond superficial mentions. Thereby, we elevate the citation prediction task from a simple binary classification to distinguishing core citations from both superficial citations and non-citations. To address this, we propose , a ybrid anguage odel workflow for citation prediction, which combines embedding and generative LMs. We design a curriculum finetune procedure to adapt a pretrained text embedding model to coarsely retrieve high-likelihood core citations from vast candidates and then design an LLM agentic workflow to rank the retrieved papers through one-shot reasoning, revealing the implicit relationships among papers. With the pipeline, we can scale the candidate sets to 100K papers. We evaluate HLM-Cite across 19 scientific fields, demonstrating a 17.6% performance improvement comparing SOTA methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a0bf4499-4c5f-4dd7-b8e4-3259f7cc4b24Cited by top-tier papers4
- RL of Thoughts: Navigating LLM Reasoning with Inference-time Reinforcement LearningQianyue Hao, Sibo Li, Jian Yuan, Yong LiICLR 2026 · 20 citations
- LLM-Explorer: A Plug-in Reinforcement Learning Policy Exploration Enhancement Driven by Large Language ModelsQianyue Hao, Yiwen Song, Qingmin Liao, Jian Yuan et al.NeurIPS 2025 · 6 citations
- From Newborn to Impact: Bias-Aware Citation PredictionMingfei Lu, Mengjia Wu, Jiawei Xu, Weikai Li et al.WWW 2026 · 6 citations
- Adapting Pretrained Language Models for Citation Classification via Self-Supervised Contrastive LearningTong Li, Jiachuan Wang, Yongqi Zhang, Shuangyin Li et al.KDD 2025
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
Related papers
- LitFM: A Retrieval Augmented Structure-aware Foundation Model For Citation GraphsJiasheng Zhang, Ali Maatouk, Jialin Chen, Ngoc Bui et al.KDD 2025 · 2 citations
- SelfCite: Self-Supervised Alignment for Context Attribution in Large Language ModelsYung-Sung Chuang, Benjamin Cohen-Wang, Zejiang Shen, Zhaofeng Wu et al.ICML 2025
- Information Re-Organization Improves Reasoning in Large Language ModelsXiaoxia Cheng, Zeqi Tan, Wei Xue, Weiming LuNeurIPS 2024 · 6 citations
- Selection-Inference: Exploiting Large Language Models for Interpretable Logical ReasoningAntonia Creswell, Murray Shanahan, Irina HigginsICLR 2023 · 110 citations
- Automatic Generation of Citation Texts in Scholarly Papers: A Pilot StudyXinyu Xing, Xiaosheng Fan, Xiaojun WanACL 2020 · 39 citations
