OoMMix: Out-of-manifold Regularization in Contextual Embedding Space for Text Classification
Seonghyeon Lee, Dongha Lee, Hwanjo Yu
摘要
Recent studies on neural networks with pretrained weights (i.e., BERT) have mainly focused on a low-dimensional subspace, where the embedding vectors computed from input words (or their contexts) are located. In this work, we propose a new approach, called OoMMix, to finding and regularizing the remainder of the space, referred to as out-ofmanifold, which cannot be accessed through the words. Specifically, we synthesize the outof-manifold embeddings based on two embeddings obtained from actually-observed words, to utilize them for fine-tuning the network. A discriminator is trained to detect whether an input embedding is located inside the manifold or not, and simultaneously, a generator is optimized to produce new embeddings that can be easily identified as out-of-manifold by the discriminator. These two modules successfully collaborate in a unified and end-to-end manner for regularizing the out-of-manifold. Our extensive evaluation on various text classification benchmarks demonstrates the effectiveness of our approach, as well as its good compatibility with existing data augmentation techniques which aim to enhance the manifold.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong 等NeurIPS 2020 · 被引用 2,774 次
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text ClassificationJiaao Chen, Zichao Yang, Diyi YangACL 2020 · 被引用 340 次
- Nonlinear Mixup: Out-Of-Manifold Data Augmentation for Text ClassificationHongyu GuoAAAI 2020 · 被引用 124 次
- Isotropy in the Contextual Embedding Space: Clusters and ManifoldsXingyu Cai, Jiaji Huang, Yuchen Bian, Kenneth ChurchICLR 2021 · 被引用 50 次
相关 Paper
- Calibrated Language Model Fine-Tuning for In- and Out-of-Distribution DataLingkai Kong, Haoming Jiang, Yuchen Zhuang, Jie Lyu 等EMNLP 2020 · 被引用 47 次
- MASKER: Masked Keyword Regularization for Reliable Text ClassificationSeung Jun Moon, Sangwoo Mo, Kimin Lee, Jaeho Lee 等AAAI 2021 · 被引用 39 次
- TextManiA: Enriching Visual Feature by Text-driven Manifold AugmentationMoon Ye-Bin, Jisoo Kim, Hongyeob Kim, Kilho Son 等ICCV 2023 · 被引用 14 次
- Towards Robustness of Deep Neural Networks via RegularizationYao Li, Martin Renqiang Min, Thomas C. M. Lee, Wenchao Yu 等ICCV 2021 · 被引用 8 次
- Not All LLM-Generated Data Are Equal: Rethinking Data Weighting in Text ClassificationHsun-Yu Kuo, Yin-Hsiang Liao, Yu-Chieh Chao, Wei-Yun Ma 等ICLR 2025
