Weakly Supervised Multi-Label Classification of Full-Text Scientific Papers
Yu Zhang, Bowen Jin, Xiusi Chen, Yanzhen Shen, Yunyi Zhang, Yu Meng, Jiawei Han
Abstract
Instead of relying on human-annotated training samples to build a classifier, weakly supervised scientific paper classification aims to classify papers only using category descriptions (e.g., category names, category-indicative keywords). Existing studies on weakly supervised paper classification are less concerned with two challenges: (1) Papers should be classified into not only coarse-grained research topics but also fine-grained themes, and potentially into multiple themes, given a large and fine-grained label space; and (2) full text should be utilized to complement the paper title and abstract for classification. Moreover, instead of viewing the entire paper as a long linear sequence, one should exploit the structural information such as citation links across papers and the hierarchy of sections and paragraphs in each paper. To tackle these challenges, in this study, we propose FuTex, a framework that uses the cross-paper network structure and the in-paper hierarchy structure to classify full-text scientific papers under weak supervision. A network-aware contrastive fine-tuning module and a hierarchyaware aggregation module are designed to leverage the two types of structural signals, respectively. Experiments on two benchmark datasets demonstrate that FuTex significantly outperforms competitive baselines and is on par with fully supervised classifiers that use 1,000 to 60,000 ground-truth training samples. CCS CONCEPTS • Information systems → Data mining; • Computing methodologies → Classification and regression trees.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Evidential Mixture Machines: Deciphering Multi-Label Correlations for Active Learning SensitivityDayou Yu, Minghao Li, Weishi Shi, Qi YuNeurIPS 2024 · 3 citations
- Incubating Text Classifiers Following User Instruction with Nothing but LLMLetian Peng, Zilong Wang, Jingbo ShangEMNLP 2024
Builds on17
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- LinkBERT: Pretraining Language Models with Document LinksMichihiro Yasunaga, Jure Leskovec, Percy LiangACL 2022 · 463 citations
- S2ORC: The Semantic Scholar Open Research CorpusKyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney et al.ACL 2020 · 424 citations
- GraphFormers: GNN-nested Transformers for Representation Learning on Textual GraphJunhan Yang, Zheng Liu, Shitao Xiao, Chaozhuo Li et al.NeurIPS 2021 · 262 citations
Related papers
- Minimally-Supervised Structure-Rich Text Categorization via Learning on Text-Rich NetworksXinyang Zhang, Chenwei Zhang, Xin Luna Dong, Jingbo Shang et al.WWW 2021 · 21 citations
- META: Metadata-Empowered Weak Supervision for Text ClassificationDheeraj Mekala, Xinyang Zhang, Jingbo ShangEMNLP 2020 · 34 citations
- Scientific Paper Extractive Summarization Enhanced by Citation GraphsXiuying Chen, Mingzhe Li, Shen Gao, Rui Yan et al.EMNLP 2022 · 8 citations
- Contextualized Weak Supervision for Text ClassificationDheeraj Mekala, Jingbo ShangACL 2020 · 121 citations
- Hierarchical Multi-Label Classification of Scientific DocumentsMobashir Sadat, Cornelia CarageaEMNLP 2022 · 16 citations
