Label-Specific Feature Augmentation for Long-Tailed Multi-Label Text Classification
Pengyu Xu, Lin Xiao, Bing Liu, Sijin Lu, Liping Jing, Jian Yu
摘要
Multi-label text classification (MLTC) involves tagging a document with its most relevant subset of labels from a label set. In real applications, labels usually follow a long-tailed distribution, where most labels (called as tail-label) only contain a small number of documents and limit the performance of MLTC. To facilitate this low-resource problem, researchers introduced a simple but effective strategy, data augmentation (DA). However, most existing DA approaches struggle in multi-label settings. The main reason is that the augmented documents for one label may inevitably influence the other co-occurring labels and further exaggerate the long-tailed problem. To mitigate this issue, we propose a new pair-level augmentation framework for MLTC, called Label-Specific Feature Augmentation (LSFA), which merely augments positive feature-label pairs for the tail-labels. LSFA contains two main parts. The first is for label-specific document representation learning in the high-level latent space, the second is for augmenting tail-label features in latent space by transferring the documents second-order statistics (intra-class semantic variations) from head labels to tail labels. At last, we design a new loss function for adjusting classifiers based on augmented datasets. The whole learning procedure can be effectively trained. Comprehensive experiments on benchmark datasets have shown that the proposed LSFA outperforms the state-of-the-art counterparts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper14
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna 等NeurIPS 2020 · 被引用 7,049 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Generating Representative Samples for Few-Shot ClassificationJingyi Xu, Hieu LeCVPR 2022 · 被引用 96 次
- FlipDA: Effective and Robust Data Augmentation for Few-Shot LearningJing Zhou, Yanan Zheng, Jie Tang, Li Jian 等ACL 2022 · 被引用 91 次
- Does Head Label Help for Long-Tailed Multi-Label Text ClassificationLin Xiao, Xiangliang Zhang, Liping Jing, Chi Huang 等AAAI 2021 · 被引用 73 次
相关 Paper
- Meta-LMTC: Meta-Learning for Large-Scale Multi-Label Text ClassificationRan Wang, Xi'ao Su, Siyu Long, Xinyu Dai 等EMNLP 2021 · 被引用 10 次
- DEALT: LLM-driven Diversity-Enhanced Data Augmentation for Long-Tail Text ClassificationWayne Lu, Xiaoxi CuiAAAI 2026 · 被引用 2 次
- A Model of Two Tales: Dual Transfer Learning Framework for Improved Long-tail Item RecommendationYin Zhang, Derek Zhiyuan Cheng, Tiansheng Yao, Xinyang Yi 等WWW 2021 · 被引用 124 次
- Incorporating Label Embedding and Feature Augmentation for Multi-Dimensional ClassificationHaobo Wang, Chen Chen, Weiwei Liu, Ke Chen 等AAAI 2020 · 被引用 25 次
- LLM-AutoDA: Large Language Model-Driven Automatic Data Augmentation for Long-tailed ProblemsPengkun Wang, Zhe Zhao, Haibin Wen, Fanfu Wang 等NeurIPS 2024 · 被引用 26 次
