Multi-Label Intent Detection via Contrastive Task Specialization of Sentence Encoders
Ivan Vulic, Iñigo Casanueva, Georgios Spithourakis, Avishek Mondal, Tsung-Hsien Wen, Pawel Budzianowski
摘要
Deploying task-oriented dialog (TOD) systems for new domains and tasks requires natural language understanding models that are 1) resource-efficient and work under low-data regimes; 2) adaptable, efficient, and quickto-train; 3) expressive and can handle complex TOD scenarios with multiple user intents in a single utterance. Motivated by these requirements, we introduce a novel framework for multi-label intent detection (mID): MULTI-CONVFIT (Multi-Label Intent Detection via Contrastive Conversational Fine-Tuning). While previous work on efficient single-label intent detection learns a classifier on top of a fixed sentence encoder (SE), we propose to 1) transform general-purpose SEs into task-specialized SEs via contrastive fine-tuning on annotated multi-label data, 2) where task specialization knowledge can be stored into lightweight adapter modules without updating the original parameters of the input SE, and then 3) we build improved mID classifiers stacked on top of fixed specialized SEs. Our main results indicate that MULTI-CONVFIT yields effective mID models, with large gains over non-specialized SEs reported across a spectrum of different mID datasets, both in low-data and high-data regimes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Divide-Solve-Combine: An Interpretable and Accurate Prompting Framework for Zero-shot Multi-Intent DetectionLibo Qin, Qiguang Chen, Jingxuan Zhou, Jin Wang 等AAAI 2025 · 被引用 9 次
- DualCL: Principled Supervised Contrastive Learning as Mutual Information Maximization for Text ClassificationJunfan Chen, Richong Zhang, Yaowei Zheng, Qianben Chen 等WWW 2024 · 被引用 4 次
它引用的顶会 Paper16
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao 等NeurIPS 2020 · 被引用 2,727 次
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu 等NeurIPS 2020 · 被引用 1,957 次
- Towards a Unified View of Parameter-Efficient Transfer LearningJunxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick 等ICLR 2022 · 被引用 1,182 次
- Contrastive Learning with Hard Negative SamplesJoshua David Robinson, Ching-Yao Chuang, Suvrit Sra, Stefanie JegelkaICLR 2021 · 被引用 999 次
相关 Paper
- ConvFiT: Conversational Fine-Tuning of Pretrained Language ModelsIvan Vulic, Pei-Hao Su, Samuel Coope, Daniela Gerz 等EMNLP 2021 · 被引用 30 次
- TOD-BERT: Pre-trained Natural Language Understanding for Task-Oriented DialogueChien-Sheng Wu, Steven C. H. Hoi, Richard Socher, Caiming XiongEMNLP 2020 · 被引用 210 次
- Transfer-Free Data-Efficient Multilingual Slot LabelingEvgeniia Razumovskaia, Ivan Vulic, Anna KorhonenEMNLP 2023 · 被引用 1 次
- Pre-training Intent-Aware Encoders for Zero- and Few-Shot Intent ClassificationMujeen Sung, James Gung, Elman Mansimov, Nikolaos Pappas 等EMNLP 2023 · 被引用 1 次
- Unified Dialog Model Pre-training for Task-Oriented Dialog Understanding and GenerationWanwei He, Yinpei Dai, Min Yang, Jian Sun 等SIGIR 2022 · 被引用 41 次
