Coarse2Fine: Fine-grained Text Classification on Coarsely-grained Annotated Data
Dheeraj Mekala, Varun Gangal, Jingbo Shang
摘要
Existing text classification methods mainly focus on a fixed label set, whereas many realworld applications require extending to new fine-grained classes as the number of samples per label increases. To accommodate such requirements, we introduce a new problem called coarse-to-fine grained classification, which aims to perform fine-grained classification on coarsely annotated data. Instead of asking for new fine-grained human annotations, we opt to leverage label surface names as the only human guidance and weave in rich pretrained generative language models into the iterative weak supervision strategy. Specifically, we first propose a label-conditioned finetuning formulation to attune these generators for our task. Furthermore, we devise a regularization objective based on the coarse-fine label constraints derived from our problem setting, giving us even further improvements over the prior formulation. Our framework uses the fine-tuned generative models to sample pseudo-training data for training the classifier, and bootstraps on real unlabeled data for model refinement. Extensive experiments and case studies on two real-world datasets demonstrate superior performance over SOTA zeroshot classification baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Fine-grained Category Discovery under Coarse-grained supervision with Hierarchical Weighted Self-contrastive LearningWenbin An, Feng Tian, Ping Chen, Siliang Tang 等EMNLP 2022 · 被引用 13 次
- Leveraging QA Datasets to Improve Generative Data AugmentationDheeraj Mekala, Tu Vu, Timo Schick, Jingbo ShangEMNLP 2022 · 被引用 8 次
- Label-Aware Hyperbolic Embeddings for Fine-grained Emotion ClassificationChih-Yao Chen, Tun-Min Hung, Yi-Li Hsu, Lun-Wei KuACL 2023 · 被引用 6 次
- DNA: Denoised Neighborhood Aggregation for Fine-grained Category DiscoveryWenbin An, Feng Tian, Wenkai Shi, Yan Chen 等EMNLP 2023 · 被引用 3 次
- A Generic Method for Fine-grained Category Discovery in Natural Language TextsChang Tian, Matthew B. Blaschko, Wenpeng Yin, Mingzhe Xing 等EMNLP 2024 · 被引用 2 次
它引用的顶会 Paper5
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Text Classification Using Label Names Only: A Language Model Self-Training ApproachYu Meng, Yunyi Zhang, Jiaxin Huang, Chenyan Xiong 等EMNLP 2020 · 被引用 203 次
- Contextualized Weak Supervision for Text ClassificationDheeraj Mekala, Jingbo ShangACL 2020 · 被引用 121 次
- Likelihood Ratios and Generative Classifiers for Unsupervised Out-of-Domain Detection in Task Oriented DialogVarun Gangal, Abhinav Arora, Arash Einolghozati, Sonal GuptaAAAI 2020 · 被引用 59 次
- META: Metadata-Empowered Weak Supervision for Text ClassificationDheeraj Mekala, Xinyang Zhang, Jingbo ShangEMNLP 2020 · 被引用 34 次
相关 Paper
- PIEClass: Weakly-Supervised Text Classification with Prompting and Noise-Robust Iterative Ensemble TrainingYunyi Zhang, Minhao Jiang, Yu Meng, Yu Zhang 等EMNLP 2023 · 被引用 16 次
- Liberating Seen Classes: Boosting Few-Shot and Zero-Shot Text Classification via Anchor Generation and Classification ReframingHan Liu, Siyang Zhao, Xiaotong Zhang, Feng Zhang 等AAAI 2024 · 被引用 7 次
- The Benefits of Label-Description Training for Zero-Shot Text ClassificationLingyu Gao, Debanjan Ghosh, Kevin GimpelEMNLP 2023 · 被引用 6 次
- RulePrompt: Weakly Supervised Text Classification with Prompting PLMs and Self-Iterative Logical RulesMiaomiao Li, Jiaqi Zhu, Yang Wang, Yi Yang 等WWW 2024 · 被引用 5 次
- Zero-Shot Text Classification with Self-TrainingAriel Gera, Alon Halfon, Eyal Shnarch, Yotam Perlitz 等EMNLP 2022 · 被引用 48 次
