Multi-Label Intent Detection via Contrastive Task Specialization of Sentence Encoders
Ivan Vulic, Iñigo Casanueva, Georgios Spithourakis, Avishek Mondal, Tsung-Hsien Wen, Pawel Budzianowski
Abstract
Deploying task-oriented dialog (TOD) systems for new domains and tasks requires natural language understanding models that are 1) resource-efficient and work under low-data regimes; 2) adaptable, efficient, and quickto-train; 3) expressive and can handle complex TOD scenarios with multiple user intents in a single utterance. Motivated by these requirements, we introduce a novel framework for multi-label intent detection (mID): MULTI-CONVFIT (Multi-Label Intent Detection via Contrastive Conversational Fine-Tuning). While previous work on efficient single-label intent detection learns a classifier on top of a fixed sentence encoder (SE), we propose to 1) transform general-purpose SEs into task-specialized SEs via contrastive fine-tuning on annotated multi-label data, 2) where task specialization knowledge can be stored into lightweight adapter modules without updating the original parameters of the input SE, and then 3) we build improved mID classifiers stacked on top of fixed specialized SEs. Our main results indicate that MULTI-CONVFIT yields effective mID models, with large gains over non-specialized SEs reported across a spectrum of different mID datasets, both in low-data and high-data regimes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Divide-Solve-Combine: An Interpretable and Accurate Prompting Framework for Zero-shot Multi-Intent DetectionLibo Qin, Qiguang Chen, Jingxuan Zhou, Jin Wang et al.AAAI 2025 · 9 citations
- DualCL: Principled Supervised Contrastive Learning as Mutual Information Maximization for Text ClassificationJunfan Chen, Richong Zhang, Yaowei Zheng, Qianben Chen et al.WWW 2024 · 4 citations
Builds on16
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao et al.NeurIPS 2020 · 2,727 citations
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu et al.NeurIPS 2020 · 1,957 citations
- Towards a Unified View of Parameter-Efficient Transfer LearningJunxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick et al.ICLR 2022 · 1,182 citations
- Contrastive Learning with Hard Negative SamplesJoshua David Robinson, Ching-Yao Chuang, Suvrit Sra, Stefanie JegelkaICLR 2021 · 999 citations
Related papers
- ConvFiT: Conversational Fine-Tuning of Pretrained Language ModelsIvan Vulic, Pei-Hao Su, Samuel Coope, Daniela Gerz et al.EMNLP 2021 · 30 citations
- TOD-BERT: Pre-trained Natural Language Understanding for Task-Oriented DialogueChien-Sheng Wu, Steven C. H. Hoi, Richard Socher, Caiming XiongEMNLP 2020 · 210 citations
- Transfer-Free Data-Efficient Multilingual Slot LabelingEvgeniia Razumovskaia, Ivan Vulic, Anna KorhonenEMNLP 2023 · 1 citation
- Pre-training Intent-Aware Encoders for Zero- and Few-Shot Intent ClassificationMujeen Sung, James Gung, Elman Mansimov, Nikolaos Pappas et al.EMNLP 2023 · 1 citation
- Unified Dialog Model Pre-training for Task-Oriented Dialog Understanding and GenerationWanwei He, Yinpei Dai, Min Yang, Jian Sun et al.SIGIR 2022 · 41 citations
