TELEClass: Taxonomy Enrichment and LLM-Enhanced Hierarchical Text Classification with Minimal Supervision
Yunyi Zhang, Ruozhen Yang, Xueqiang Xu, Rui Li, Jinfeng Xiao, Jiaming Shen, Jiawei Han
摘要
Hierarchical text classification aims to categorize each document into a set of classes in a label taxonomy, which is a fundamental web text mining task with broad applications such as web content analysis and semantic indexing. Most earlier works focus on fully or semi-supervised methods that require a large amount of human annotated data which is costly and time-consuming to acquire. To alleviate human efforts, in this paper, we work on hierarchical text classification with a minimal amount of supervision: using the sole class name of each node as the only supervision. Recently, large language models (LLM) have shown competitive performance on various tasks through zero-shot prompting, but this method performs poorly in the hierarchical setting because it is ineffective to include the large and structured label space in a prompt. On the other hand, previous weakly-supervised hierarchical text classification methods only utilize the raw taxonomy skeleton and ignore the rich information hidden in the text corpus that can serve as additional class-indicative features. To tackle the above challenges, we propose TELEClass, Taxonomy Enrichment and LLM-Enhanced weakly-supervised hierarchical text Classification, which combines the general knowledge of LLMs and task-specific features mined from an unlabeled corpus. TELEClass automatically enriches the raw taxonomy with class-indicative features for better label space understanding and utilizes novel LLM-based data annotation and generation methods specifically tailored for the hierarchical setting. Experiments show that TELEClass can significantly outperform previous baselines while achieving comparable performance to zero-shot prompting of LLMs with drastically less inference cost. CCS Concepts • Information systems → Data mining; • Computing methodologies → Natural language processing.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Into the Unknown Unknowns: Engaged Human Learning through Participation in Language Model Agent ConversationsYucheng Jiang, Yijia Shao, Dekun Ma, Sina J. Semnani 等EMNLP 2024 · 被引用 8 次
- Taxonomy-guided Semantic Indexing for Academic Paper SearchSeongKu Kang, Yunyi Zhang, Pengcheng Jiang, Dongha Lee 等EMNLP 2024 · 被引用 3 次
- DEALT: LLM-driven Diversity-Enhanced Data Augmentation for Long-Tail Text ClassificationWayne Lu, Xiaoxi CuiAAAI 2026 · 被引用 2 次
- Learning Hierarchical Knowledge in Text-Rich Networks with Taxonomy-Informed Representation LearningYunhui Liu, Yongchao Liu, Yinfeng Chen, Chuntao Hong 等KDD 2026 · 被引用 1 次
- Ensembling Prompting Strategies for Zero-Shot Hierarchical Text Classification with Large Language ModelsMingxuan Xia, Zhijie Jiang, Haobo Wang, Junbo Zhao 等EMNLP 2025 · 被引用 1 次
它引用的顶会 Paper17
- RAPTOR: Recursive Abstractive Processing for Tree-Organized RetrievalParth Sarthi, Salman Abdullah, Aditi Tuli, Shubh Khanna 等ICLR 2024 · 被引用 460 次
- Generating Training Data with Language Models: Towards Zero-Shot Language UnderstandingYu Meng, Jiaxin Huang, Yu Zhang, Jiawei HanNeurIPS 2022 · 被引用 309 次
- Text Classification Using Label Names Only: A Language Model Self-Training ApproachYu Meng, Yunyi Zhang, Jiaxin Huang, Chenyan Xiong 等EMNLP 2020 · 被引用 203 次
- Hierarchy-Aware Global Model for Hierarchical Text ClassificationJie Zhou, Chunping Ma, Dingkun Long, Guangwei Xu 等ACL 2020 · 被引用 171 次
- ZeroGen: Efficient Zero-shot Learning via Dataset GenerationJiacheng Ye, Jiahui Gao, Qintong Li, Hang Xu 等EMNLP 2022 · 被引用 96 次
相关 Paper
- Label Augmentation for Zero-Shot Hierarchical Text ClassificationLorenzo Paletto, Valerio Basile, Roberto EspositoACL 2024
- RulePrompt: Weakly Supervised Text Classification with Prompting PLMs and Self-Iterative Logical RulesMiaomiao Li, Jiaqi Zhu, Yang Wang, Yi Yang 等WWW 2024 · 被引用 5 次
- Open-world Multi-label Text Classification with Extremely Weak SupervisionXintong Li, Jinya Jiang, Ria Dharmani, Jayanth Srinivasa 等EMNLP 2024 · 被引用 3 次
- HTLM: Hyper-Text Pre-Training and Prompting of Language ModelsArmen Aghajanyan, Dmytro Okhonko, Mike Lewis, Mandar Joshi 等ICLR 2022 · 被引用 82 次
- Liberating Seen Classes: Boosting Few-Shot and Zero-Shot Text Classification via Anchor Generation and Classification ReframingHan Liu, Siyang Zhao, Xiaotong Zhang, Feng Zhang 等AAAI 2024 · 被引用 7 次
