SciPrompt: Knowledge-augmented Prompting for Fine-grained Categorization of Scientific Topics
Zhiwen You, Kanyao Han, Haotian Zhu, Bertram Ludäscher, Jana Diesner
Abstract
Prompt-based fine-tuning has become an essential method for eliciting information encoded in pre-trained language models for a variety of tasks, including text classification. For multi-class classification tasks, prompt-based fine-tuning under low-resource scenarios has resulted in performance levels comparable to those of fully fine-tuning methods. Previous studies have used crafted prompt templates and verbalizers, mapping from the label terms space to the class space, to solve the classification problem as a masked language modeling task. However, cross-domain and finegrained prompt-based fine-tuning with an automatically enriched verbalizer remains unexplored, mainly due to the difficulty and costs of manually selecting domain label terms for the verbalizer, which requires humans with domain expertise. To address this challenge, we introduce SCIPROMPT, a framework designed to automatically retrieve scientific topic-related terms for low-resource text classification tasks. To this end, we select semantically correlated and domain-specific label terms within the context of scientific literature for verbalizer augmentation. Furthermore, we propose a new verbalization strategy that uses correlation scores as additional weights to enhance the prediction performance of the language model during model tuning. Our method outperforms stateof-the-art, prompt-based fine-tuning methods on scientific text classification tasks under few and zero-shot settings, especially in classifying fine-grained and emerging scientific topics 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ddebd12d-d1d3-4edb-bad7-c0420edd5c4eBuilds on13
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated PromptsTaylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace et al.EMNLP 2020 · 1,162 citations
- S2ORC: The Semantic Scholar Open Research CorpusKyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney et al.ACL 2020 · 424 citations
- Noisy Channel Language Model Prompting for Few-Shot Text ClassificationSewon Min, Mike Lewis, Hannaneh Hajishirzi, Luke ZettlemoyerACL 2022 · 237 citations
- Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMsOded Ovadia, Menachem Brief, Moshik Mishaeli, Oren ElishaEMNLP 2024 · 89 citations
- SciNLI: A Corpus for Natural Language Inference on Scientific TextMobashir Sadat, Cornelia CarageaACL 2022 · 41 citations
Related papers
- Knowledgeable Prompt-tuning: Incorporating Knowledge into Prompt Verbalizer for Text ClassificationShengding Hu, Ning Ding, Huadong Wang, Zhiyuan Liu et al.ACL 2022
- RulePrompt: Weakly Supervised Text Classification with Prompting PLMs and Self-Iterative Logical RulesMiaomiao Li, Jiaqi Zhu, Yang Wang, Yi Yang et al.WWW 2024 · 5 citations
- Prototypical Verbalizer for Prompt-based Few-shot TuningGanqu Cui, Shengding Hu, Ning Ding, Longtao Huang et al.ACL 2022
- Hierarchical Verbalizer for Few-Shot Hierarchical Text ClassificationKe Ji, Yixin Lian, Jingsheng Gao, Baoyuan WangACL 2023 · 17 citations
- MetricPrompt: Prompting Model as a Relevance Metric for Few-shot Text ClassificationHongyuan Dong, Weinan Zhang, Wanxiang CheKDD 2023 · 2 citations
