The Effect of Metadata on Scientific Literature Tagging: A Cross-Field Cross-Model Study
Yu Zhang, Bowen Jin, Qi Zhu, Yu Meng, Jiawei Han
摘要
Due to the exponential growth of scientific publications on the Web, there is a pressing need to tag each paper with fine-grained topics so that researchers can track their interested fields of study rather than drowning in the whole literature. Scientific literature tagging is beyond a pure multi-label text classification task because papers on the Web are prevalently accompanied by metadata information such as venues, authors, and references, which may serve as additional signals to infer relevant tags. Although there have been studies making use of metadata in academic paper classification, their focus is often restricted to one or two scientific fields (e.g., computer science and biomedicine) and to one specific model. In this work, we systematically study the effect of metadata on scientific literature tagging across 19 fields. We select three representative multi-label classifiers (i.e., a bag-of-words model, a sequence-based model, and a pre-trained language model) and explore their performance change in scientific literature tagging when metadata are fed to the classifiers as additional features. We observe some ubiquitous patterns of metadata's effects across all fields (e.g., venues are consistently beneficial to paper tagging in almost all cases), as well as some unique patterns in fields other than computer science and biomedicine, which are not explored in previous studies. CCS CONCEPTS • Information systems → Digital libraries and archives; Data mining; World Wide Web.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific DiscoveryYu Zhang, Xiusi Chen, Bowen Jin, Sheng Wang 等EMNLP 2024 · 被引用 28 次
- A Scalable Pretraining Framework for Link Prediction with Efficient AdaptationYu Song, Zhigang Hua, Harry Shomer, Yan Xie 等KDD 2025 · 被引用 1 次
它引用的顶会 Paper9
- LightXML: Transformer with Dynamic Negative Sampling for High-Performance Extreme Multi-label Text ClassificationTing Jiang, Deqing Wang, Leilei Sun, Huayi Yang 等AAAI 2021 · 被引用 170 次
- Fast Multi-Resolution Transformer Fine-tuning for Extreme Multi-label Text ClassificationJiong Zhang, Wei-Cheng Chang, Hsiang-Fu Yu, Inderjit S. DhillonNeurIPS 2021 · 被引用 147 次
- ECLARE: Extreme Classification with Label Graph CorrelationsAnshul Mittal, Noveen Sachdeva, Sheshansh Agrawal, Sumeet Agarwal 等WWW 2021 · 被引用 71 次
- Correlation Networks for Extreme Multi-label Text ClassificationGuangxu Xun, Kishlay Jha, Jianhui Sun, Aidong ZhangKDD 2020 · 被引用 60 次
- Pretrained Generalized Autoregressive Model with Adaptive Probabilistic Label Clusters for Extreme Multi-label Text ClassificationHui Ye, Zhiyu Chen, Da-Han Wang, Brian D. DavisonICML 2020 · 被引用 57 次
相关 Paper
- META: Metadata-Empowered Weak Supervision for Text ClassificationDheeraj Mekala, Xinyang Zhang, Jingbo ShangEMNLP 2020 · 被引用 34 次
- Minimally Supervised Categorization of Text with MetadataYu Zhang, Yu Meng, Jiaxin Huang, Frank F. Xu 等SIGIR 2020 · 被引用 31 次
- Pre-Training and Prompting for Few-Shot Node Classification on Text-Attributed GraphsHuanjing Zhao, Beining Yang, Yukuo Cen, Junyu Ren 等KDD 2024 · 被引用 12 次
- Hierarchical Multi-Label Classification of Scientific DocumentsMobashir Sadat, Cornelia CarageaEMNLP 2022 · 被引用 16 次
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo 等ACL 2020 · 被引用 93 次
