From Words to Worth: Newborn Article Impact Prediction with LLM
Penghai Zhao, Qinghua Xing, Kairan Dou, Jinyu Tian, Ying Tai, Jian Yang, Ming-Ming Cheng, Xiang Li
Abstract
As the academic landscape expands, the challenge of efficiently identifying impactful newly published articles grows increasingly vital. This paper introduces a promising approach, leveraging the capabilities of LLMs to predict the future impact of newborn articles solely based on titles and abstracts. Moving beyond traditional methods heavily reliant on external information, the proposed method employs LLM to discern the shared semantic features of highly impactful papers from a large collection of title-abstract pairs. These semantic features are further utilized to predict the proposed indicator, TNCSISP, which incorporates favorable normalization properties across value, field, and time. To facilitate parameter-efficient fine-tuning of the LLM, we have also meticulously curated a dataset containing over 12,000 entries, each annotated with titles, abstracts, and their corresponding TNCSISP values. The quantitative results, with an MAE of 0.216 and an NDCG@20 of 0.901, demonstrate that the proposed approach achieves state-of-the-art performance in predicting the impact of newborn articles when compared to several promising methods. Finally, we present a realworld application example for predicting the impact of newborn journal articles to demonstrate its noteworthy practical value. Overall, our findings challenge existing paradigms and propose a shift towards a more contentfocused prediction of academic impact, offering new insights for article impact prediction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- NAIPv2: Debiased Pairwise Learning for Efficient Paper Quality EstimationPenghai Zhao, Jinyu Tian, Qinghua Xing, Xin Zhang et al.ICLR 2026 · 6 citations
- From Newborn to Impact: Bias-Aware Citation PredictionMingfei Lu, Mengjia Wu, Jiawei Xu, Weikai Li et al.WWW 2026 · 6 citations
- WOW-Seg: A Word-free Open World Segmentation ModelDanyang Li, Tianhao Wu, Bin Lin, Zhenyuan Chen et al.ICLR 2026 · 2 citations
- Navigating Through Paper Flood: Advancing LLM-Based Paper Evaluation Through Domain-Aware Retrieval and Latent ReasoningWuqiang Zheng, Yiyan Xu, Xinyu Lin, Chongming Gao et al.AAAI 2026
Builds on3
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- DoRA: Weight-Decomposed Low-Rank AdaptationShih-Yang Liu, Chien-Yi Wang, Hongxu Yin, Pavlo Molchanov et al.ICML 2024 · 820 citations
- Large Selective Kernel Network for Remote Sensing Object DetectionYuxuan Li, Qibin Hou, Zhaohui Zheng, Ming-Ming Cheng et al.ICCV 2023 · 535 citations
Related papers
- SemCSE: Semantic Contrastive Sentence Embeddings Using LLM-Generated Summaries For Scientific AbstractsMarc Felix Brinner, Sina ZarrießEMNLP 2025
- HLM-Cite: Hybrid Language Model Workflow for Text-based Scientific Citation PredictionQianyue Hao, Jingyang Fan, Fengli Xu, Jian Yuan et al.NeurIPS 2024 · 23 citations
- Evaluating Scholarly Impact: Towards Content-Aware BibliometricsSaurav Manchanda, George KarypisEMNLP 2021 · 2 citations
- HINTS: Citation Time Series Prediction for New Publications via Dynamic Heterogeneous Information Network EmbeddingSong Jiang, Bernard Koch, Yizhou SunWWW 2021 · 45 citations
- Is It Novel and Why? Fine-Grained Patent Novelty Prediction Based on Passage RetrievalValentin Knappich, Anna Hätty, Simon Razniewski, Annemarie FriedrichSIGIR 2026 · 1 citation
