Re-embedding Difficult Samples via Mutual Information Constrained Semantically Oversampling for Imbalanced Text Classification
Jiachen Tian, Shizhan Chen, Xiaowang Zhang, Zhiyong Feng, Deyi Xiong, Shaojuan Wu, Chunliu Dou
Abstract
Difficult samples of the minority class in imbalanced text classification are usually hard to be classified as they are embedded into an overlapping semantic region with the majority class. In this paper, we propose a Mutual Information constrained Semantically Oversampling framework (MISO) that can generate anchor instances to help the backbone network determine the re-embedding position of a non-overlapping representation for each difficult sample. MISO consists of (1) a semantic fusion module that learns entangled semantics among difficult and majority samples with an adaptive multi-head attention mechanism, (2) a mutual information loss that forces our model to learn new representations of entangled semantics in the non-overlapping region of the minority class, and (3) a coupled adversarial encoder-decoder that fine-tunes disentangled semantic representations to remain their correlations with the minority class, and then using these disentangled semantic representations to generate anchor instances for each difficult sample. Experiments on a variety of imbalanced text classification tasks demonstrate that anchor instances help classifiers achieve significant improvements over strong baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Counterfactual-Enhanced Information Bottleneck for Aspect-Based Sentiment AnalysisMingshan Chang, Min Yang, Qingshan Jiang, Ruifeng XuAAAI 2024 · 12 citations
- Reducing Sentiment Bias in Pre-trained Sentiment Classification via Adaptive Gumbel AttackJiachen Tian, Shizhan Chen, Xiaowang Zhang, Xin Wang et al.AAAI 2023 · 5 citations
- Text Grafting: Near-Distribution Weak Supervision for Minority Classes in Text ClassificationLetian Peng, Yi Gu, Chengyu Dong, Zihan Wang et al.EMNLP 2024
Builds on5
- Decoupling Representation and Classifier for Long-Tailed RecognitionBingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan et al.ICLR 2020 · 1,496 citations
- Contrastive Learning with Hard Negative SamplesJoshua David Robinson, Ching-Yao Chuang, Suvrit Sra, Stefanie JegelkaICLR 2021 · 999 citations
- Dice Loss for Data-imbalanced NLP TasksXiaoya Li, Xiaofei Sun, Yuxian Meng, Junjun Liang et al.ACL 2020 · 575 citations
- MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text ClassificationJiaao Chen, Zichao Yang, Diyi YangACL 2020 · 340 citations
- RiSAWOZ: A Large-Scale Multi-Domain Wizard-of-Oz Dataset with Rich Semantic Annotations for Task-Oriented Dialogue ModelingJun Quan, Shian Zhang, Qian Cao, Zizhong Li et al.EMNLP 2020 · 41 citations
Related papers
- Efficient Augmentation for Imbalanced Deep LearningDamien A. Dablain, Colin Bellinger, Bartosz Krawczyk, Nitesh V. ChawlaICDE 2023 · 19 citations
- MIANet: Aggregating Unbiased Instance and General Information for Few-Shot Semantic SegmentationYong Yang, Qiong Chen, Yuan Feng, Tianlin HuangCVPR 2023
- Improving the Accuracy of Learning Example Weights for Imbalance ClassificationYuqi Liu, Bin Cao, Jing FanICLR 2022 · 10 citations
- Difficulty-Based Sampling for Debiased Contrastive Representation LearningTaeuk Jang, Xiaoqian WangCVPR 2023
- A Unified Loss for Handling Inter-Class and Intra-Class Imbalance in Medical Image SegmentationFeilong Xu, Feiyang Yang, Xiongfei Li, Xiaoli ZhangAAAI 2025 · 6 citations
