The Effect of Metadata on Scientific Literature Tagging: A Cross-Field Cross-Model Study
Yu Zhang, Bowen Jin, Qi Zhu, Yu Meng, Jiawei Han
Abstract
Due to the exponential growth of scientific publications on the Web, there is a pressing need to tag each paper with fine-grained topics so that researchers can track their interested fields of study rather than drowning in the whole literature. Scientific literature tagging is beyond a pure multi-label text classification task because papers on the Web are prevalently accompanied by metadata information such as venues, authors, and references, which may serve as additional signals to infer relevant tags. Although there have been studies making use of metadata in academic paper classification, their focus is often restricted to one or two scientific fields (e.g., computer science and biomedicine) and to one specific model. In this work, we systematically study the effect of metadata on scientific literature tagging across 19 fields. We select three representative multi-label classifiers (i.e., a bag-of-words model, a sequence-based model, and a pre-trained language model) and explore their performance change in scientific literature tagging when metadata are fed to the classifiers as additional features. We observe some ubiquitous patterns of metadata's effects across all fields (e.g., venues are consistently beneficial to paper tagging in almost all cases), as well as some unique patterns in fields other than computer science and biomedicine, which are not explored in previous studies. CCS CONCEPTS • Information systems → Digital libraries and archives; Data mining; World Wide Web.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b3ac1f42-77ae-4c66-a768-c33879b73d5cCited by top-tier papers2
- A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific DiscoveryYu Zhang, Xiusi Chen, Bowen Jin, Sheng Wang et al.EMNLP 2024 · 28 citations
- A Scalable Pretraining Framework for Link Prediction with Efficient AdaptationYu Song, Zhigang Hua, Harry Shomer, Yan Xie et al.KDD 2025 · 1 citation
Builds on9
- LightXML: Transformer with Dynamic Negative Sampling for High-Performance Extreme Multi-label Text ClassificationTing Jiang, Deqing Wang, Leilei Sun, Huayi Yang et al.AAAI 2021 · 170 citations
- Fast Multi-Resolution Transformer Fine-tuning for Extreme Multi-label Text ClassificationJiong Zhang, Wei-Cheng Chang, Hsiang-Fu Yu, Inderjit S. DhillonNeurIPS 2021 · 147 citations
- ECLARE: Extreme Classification with Label Graph CorrelationsAnshul Mittal, Noveen Sachdeva, Sheshansh Agrawal, Sumeet Agarwal et al.WWW 2021 · 71 citations
- Correlation Networks for Extreme Multi-label Text ClassificationGuangxu Xun, Kishlay Jha, Jianhui Sun, Aidong ZhangKDD 2020 · 60 citations
- Pretrained Generalized Autoregressive Model with Adaptive Probabilistic Label Clusters for Extreme Multi-label Text ClassificationHui Ye, Zhiyu Chen, Da-Han Wang, Brian D. DavisonICML 2020 · 57 citations
Related papers
- META: Metadata-Empowered Weak Supervision for Text ClassificationDheeraj Mekala, Xinyang Zhang, Jingbo ShangEMNLP 2020 · 34 citations
- Minimally Supervised Categorization of Text with MetadataYu Zhang, Yu Meng, Jiaxin Huang, Frank F. Xu et al.SIGIR 2020 · 31 citations
- Pre-Training and Prompting for Few-Shot Node Classification on Text-Attributed GraphsHuanjing Zhao, Beining Yang, Yukuo Cen, Junyu Ren et al.KDD 2024 · 12 citations
- Hierarchical Multi-Label Classification of Scientific DocumentsMobashir Sadat, Cornelia CarageaEMNLP 2022 · 16 citations
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo et al.ACL 2020 · 93 citations
