Minimally-Supervised Structure-Rich Text Categorization via Learning on Text-Rich Networks
Xinyang Zhang, Chenwei Zhang, Xin Luna Dong, Jingbo Shang, Jiawei Han
Abstract
Text categorization is an essential task in Web content analysis. Considering the ever-evolving Web data and new emerging categories, instead of the laborious supervised setting, in this paper, we focus on the minimally-supervised setting that aims to categorize documents effectively, with a couple of seed documents annotated per category. We recognize that texts collected from the Web are often structure-rich, i.e., accompanied by various metadata. One can easily organize the corpus into a text-rich network, joining raw text documents with document attributes, high-quality phrases, label surface names as nodes, and their associations as edges. Such a network provides a holistic view of the corpus' heterogeneous data sources and enables a joint optimization for network-based analysis and deep textual model training. We therefore propose a novel framework for minimally supervised categorization by learning from the text-rich network. Specifically, we jointly train two modules with different inductive biases -a text analysis module for text understanding and a network learning module for classdiscriminative, scalable network learning. Each module generates pseudo training labels from the unlabeled document set, and both modules mutually enhance each other by co-training using pooled pseudo labels. We test our model on two real-world datasets. On the challenging e-commerce product categorization dataset with 683 categories, our experiments show that given only three seed documents per category, our framework can achieve an accuracy of about 92%, significantly outperforming all compared methods; our accuracy is only less than 2% away from the supervised BERT model trained on about 50K labeled documents. CCS CONCEPTS • Information systems → Web mining; • Computing methodologies → Natural language processing; Learning settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fbb117ce-223c-41e1-94cb-4214bfc7a466Cited by top-tier papers5
- Metadata-Induced Contrastive Learning for Zero-Shot Multi-Label Text ClassificationYu Zhang, Zhihong Shen, Chieh-Han Wu, Boya Xie et al.WWW 2022 · 34 citations
- Heterformer: Transformer-based Deep Node Representation Learning on Heterogeneous Text-Rich NetworksBowen Jin, Yu Zhang, Qi Zhu, Jiawei HanKDD 2023 · 27 citations
- KGTrust: Evaluating Trustworthiness of SIoT via Knowledge Enhanced Graph Neural NetworksZhizhi Yu, Di Jin, Cuiying Huo, Zhiqiang Wang et al.WWW 2023 · 27 citations
- Detecting Miscitation on the Scholarly Web through LLM-Augmented Text-Rich Graph LearningHuidong Wu, Haojia Xiang, Jingtong Gao, Xiangyu Zhao et al.WWW 2026
- Can Large Language Models Act as Ensembler for Multi-GNNs?Hanqi Duan, Yao Cheng, Jianxiang Yu, Yao Liu et al.EMNLP 2025
Builds on4
- Text Classification Using Label Names Only: A Language Model Self-Training ApproachYu Meng, Yunyi Zhang, Jiaxin Huang, Chenyan Xiong et al.EMNLP 2020 · 203 citations
- NetTaxo: Automated Topic Taxonomy Construction from Text-Rich NetworkJingbo Shang, Xinyang Zhang, Liyuan Liu, Sha Li et al.WWW 2020 · 66 citations
- META: Metadata-Empowered Weak Supervision for Text ClassificationDheeraj Mekala, Xinyang Zhang, Jingbo ShangEMNLP 2020 · 34 citations
- Minimally Supervised Categorization of Text with MetadataYu Zhang, Yu Meng, Jiaxin Huang, Frank F. Xu et al.SIGIR 2020 · 31 citations
Related papers
- Patton: Language Model Pretraining on Text-Rich NetworksBowen Jin, Wentao Zhang, Yu Zhang, Yu Meng et al.ACL 2023 · 14 citations
- Augmenting Low-Resource Text Classification with Graph-Grounded Pre-training and PromptingZhihao Wen, Yuan FangSIGIR 2023 · 66 citations
- Weakly Supervised Multi-Label Classification of Full-Text Scientific PapersYu Zhang, Bowen Jin, Xiusi Chen, Yanzhen Shen et al.KDD 2023 · 8 citations
- CL-WSTC: Continual Learning for Weakly Supervised Text Classification on the InternetMiaomiao Li, Jiaqi Zhu, Xin Yang, Yi Yang et al.WWW 2023 · 7 citations
- TELEClass: Taxonomy Enrichment and LLM-Enhanced Hierarchical Text Classification with Minimal SupervisionYunyi Zhang, Ruozhen Yang, Xueqiang Xu, Rui Li et al.WWW 2025 · 53 citations
