Label-Specific Feature Augmentation for Long-Tailed Multi-Label Text Classification
Pengyu Xu, Lin Xiao, Bing Liu, Sijin Lu, Liping Jing, Jian Yu
Abstract
Multi-label text classification (MLTC) involves tagging a document with its most relevant subset of labels from a label set. In real applications, labels usually follow a long-tailed distribution, where most labels (called as tail-label) only contain a small number of documents and limit the performance of MLTC. To facilitate this low-resource problem, researchers introduced a simple but effective strategy, data augmentation (DA). However, most existing DA approaches struggle in multi-label settings. The main reason is that the augmented documents for one label may inevitably influence the other co-occurring labels and further exaggerate the long-tailed problem. To mitigate this issue, we propose a new pair-level augmentation framework for MLTC, called Label-Specific Feature Augmentation (LSFA), which merely augments positive feature-label pairs for the tail-labels. LSFA contains two main parts. The first is for label-specific document representation learning in the high-level latent space, the second is for augmenting tail-label features in latent space by transferring the documents second-order statistics (intra-class semantic variations) from head labels to tail labels. At last, we design a new loss function for adjusting classifiers based on augmented datasets. The whole learning procedure can be effectively trained. Comprehensive experiments on benchmark datasets have shown that the proposed LSFA outperforms the state-of-the-art counterparts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 985d2acb-ea97-44d2-8761-6b113e8119d9Cited by top-tier papers1
Ask how each one uses itBuilds on14
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Generating Representative Samples for Few-Shot ClassificationJingyi Xu, Hieu LeCVPR 2022 · 96 citations
- FlipDA: Effective and Robust Data Augmentation for Few-Shot LearningJing Zhou, Yanan Zheng, Jie Tang, Li Jian et al.ACL 2022 · 91 citations
- Does Head Label Help for Long-Tailed Multi-Label Text ClassificationLin Xiao, Xiangliang Zhang, Liping Jing, Chi Huang et al.AAAI 2021 · 73 citations
Related papers
- Meta-LMTC: Meta-Learning for Large-Scale Multi-Label Text ClassificationRan Wang, Xi'ao Su, Siyu Long, Xinyu Dai et al.EMNLP 2021 · 10 citations
- DEALT: LLM-driven Diversity-Enhanced Data Augmentation for Long-Tail Text ClassificationWayne Lu, Xiaoxi CuiAAAI 2026 · 2 citations
- A Model of Two Tales: Dual Transfer Learning Framework for Improved Long-tail Item RecommendationYin Zhang, Derek Zhiyuan Cheng, Tiansheng Yao, Xinyang Yi et al.WWW 2021 · 124 citations
- Incorporating Label Embedding and Feature Augmentation for Multi-Dimensional ClassificationHaobo Wang, Chen Chen, Weiwei Liu, Ke Chen et al.AAAI 2020 · 25 citations
- LLM-AutoDA: Large Language Model-Driven Automatic Data Augmentation for Long-tailed ProblemsPengkun Wang, Zhe Zhao, Haibin Wen, Fanfu Wang et al.NeurIPS 2024 · 26 citations
