Re-embedding Difficult Samples via Mutual Information Constrained Semantically Oversampling for Imbalanced Text Classification
Jiachen Tian, Shizhan Chen, Xiaowang Zhang, Zhiyong Feng, Deyi Xiong, Shaojuan Wu, Chunliu Dou
摘要
Difficult samples of the minority class in imbalanced text classification are usually hard to be classified as they are embedded into an overlapping semantic region with the majority class. In this paper, we propose a Mutual Information constrained Semantically Oversampling framework (MISO) that can generate anchor instances to help the backbone network determine the re-embedding position of a non-overlapping representation for each difficult sample. MISO consists of (1) a semantic fusion module that learns entangled semantics among difficult and majority samples with an adaptive multi-head attention mechanism, (2) a mutual information loss that forces our model to learn new representations of entangled semantics in the non-overlapping region of the minority class, and (3) a coupled adversarial encoder-decoder that fine-tunes disentangled semantic representations to remain their correlations with the minority class, and then using these disentangled semantic representations to generate anchor instances for each difficult sample. Experiments on a variety of imbalanced text classification tasks demonstrate that anchor instances help classifiers achieve significant improvements over strong baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Counterfactual-Enhanced Information Bottleneck for Aspect-Based Sentiment AnalysisMingshan Chang, Min Yang, Qingshan Jiang, Ruifeng XuAAAI 2024 · 被引用 12 次
- Reducing Sentiment Bias in Pre-trained Sentiment Classification via Adaptive Gumbel AttackJiachen Tian, Shizhan Chen, Xiaowang Zhang, Xin Wang 等AAAI 2023 · 被引用 5 次
- Text Grafting: Near-Distribution Weak Supervision for Minority Classes in Text ClassificationLetian Peng, Yi Gu, Chengyu Dong, Zihan Wang 等EMNLP 2024
它引用的顶会 Paper5
- Decoupling Representation and Classifier for Long-Tailed RecognitionBingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan 等ICLR 2020 · 被引用 1,496 次
- Contrastive Learning with Hard Negative SamplesJoshua David Robinson, Ching-Yao Chuang, Suvrit Sra, Stefanie JegelkaICLR 2021 · 被引用 999 次
- Dice Loss for Data-imbalanced NLP TasksXiaoya Li, Xiaofei Sun, Yuxian Meng, Junjun Liang 等ACL 2020 · 被引用 575 次
- MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text ClassificationJiaao Chen, Zichao Yang, Diyi YangACL 2020 · 被引用 340 次
- RiSAWOZ: A Large-Scale Multi-Domain Wizard-of-Oz Dataset with Rich Semantic Annotations for Task-Oriented Dialogue ModelingJun Quan, Shian Zhang, Qian Cao, Zizhong Li 等EMNLP 2020 · 被引用 41 次
相关 Paper
- Efficient Augmentation for Imbalanced Deep LearningDamien A. Dablain, Colin Bellinger, Bartosz Krawczyk, Nitesh V. ChawlaICDE 2023 · 被引用 19 次
- MIANet: Aggregating Unbiased Instance and General Information for Few-Shot Semantic SegmentationYong Yang, Qiong Chen, Yuan Feng, Tianlin HuangCVPR 2023
- Improving the Accuracy of Learning Example Weights for Imbalance ClassificationYuqi Liu, Bin Cao, Jing FanICLR 2022 · 被引用 10 次
- Difficulty-Based Sampling for Debiased Contrastive Representation LearningTaeuk Jang, Xiaoqian WangCVPR 2023
- A Unified Loss for Handling Inter-Class and Intra-Class Imbalance in Medical Image SegmentationFeilong Xu, Feiyang Yang, Xiongfei Li, Xiaoli ZhangAAAI 2025 · 被引用 6 次
