Simple-Sampling and Hard-Mixup with Prototypes to Rebalance Contrastive Learning for Text Classification
Mengyu Li, Yonghao Liu, Fausto Giunchiglia, Ximing Li, Xiaoyue Feng, Renchu Guan
摘要
Text classification is a crucial and fundamental task in web content mining. Compared with the previous learning paradigm of pre-training and fine-tuning by cross entropy loss, the recently proposed supervised contrastive learning approach has received tremendous attention due to its powerful feature learning capability and robustness. Although several studies have incorporated this technique for text classification, some limitations remain. First, many text datasets are imbalanced, and the learning mechanism of supervised contrastive learning is sensitive to data imbalance, which may harm the model's performance. Moreover, these models leverage separate classification branches with cross entropy and supervised contrastive learning branches without explicit mutual guidance. To this end, we propose a novel model named SharpReCL for imbalanced text classification tasks. First, we obtain the prototype vector of each class in the balanced classification branch to act as a representation of each class. Then, by further explicitly leveraging the prototype vectors, we construct a proper and sufficient target sample set with the same size for each class to perform the supervised contrastive learning procedure. The empirical results show the effectiveness of our model, which even outperforms popular large language models across several datasets. Our code is available here. CCS Concepts • Computing methodologies → Natural language processing; • Information systems → Document representation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Dual-level Mixup for Graph Few-shot Learning with Fewer TasksYonghao Liu, Mengyu Li, Fausto Giunchiglia, Lan Huang 等WWW 2025 · 被引用 8 次
- A Simple Graph Contrastive Learning Framework for Short Text ClassificationYonghao Liu, Fausto Giunchiglia, Lan Huang, Ximing Li 等AAAI 2025 · 被引用 6 次
- Boosting Short Text Classification with Multi-Source Information Exploration and Dual-Level Contrastive LearningYonghao Liu, Mengyu Li, Wei Pang, Fausto Giunchiglia 等AAAI 2025 · 被引用 6 次
- TRACE: Discovering Task-Specific Parameter via Adaptation-Aware Probing for Continual Fine-TuningXiaosong Han, Ke Chen, Xindi Dai, Di Liang 等KDD 2026 · 被引用 1 次
- Enhancing Unsupervised Graph Few-shot Learning via Set Functions and Optimal TransportYonghao Liu, Fausto Giunchiglia, Ximing Li, Lan Huang 等KDD 2025
它引用的顶会 Paper19
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- Contrastive Learning with Hard Negative SamplesJoshua David Robinson, Ching-Yao Chuang, Suvrit Sra, Stefanie JegelkaICLR 2021 · 被引用 999 次
相关 Paper
- DualCL: Principled Supervised Contrastive Learning as Mutual Information Maximization for Text ClassificationJunfan Chen, Richong Zhang, Yaowei Zheng, Qianben Chen 等WWW 2024 · 被引用 4 次
- Robust Representation Learning with Reliable Pseudo-labels Generation via Self-Adaptive Optimal Transport for Short Text ClusteringXiaolin Zheng, Mengling Hu, Weiming Liu, Chaochao Chen 等ACL 2023 · 被引用 12 次
- Subclass-balancing Contrastive Learning for Long-tailed RecognitionChengkai Hou, Jieyu Zhang, Haonan Wang, Tianyi ZhouICCV 2023 · 被引用 50 次
- Exploring Balanced Feature Spaces for Representation LearningBingyi Kang, Yu Li, Sa Xie, Zehuan Yuan 等ICLR 2021 · 被引用 296 次
- Unified Contrastive Learning in Image-Text-Label SpaceJianwei Yang, Chunyuan Li, Pengchuan Zhang, Bin Xiao 等CVPR 2022 · 被引用 182 次
