ConvFiT: Conversational Fine-Tuning of Pretrained Language Models
Ivan Vulic, Pei-Hao Su, Samuel Coope, Daniela Gerz, Pawel Budzianowski, Iñigo Casanueva, Nikola Mrksic, Tsung-Hsien Wen
Abstract
Transformer-based language models (LMs) pretrained on large text collections are proven to store a wealth of semantic knowledge. However, 1) they are not effective as sentence encoders when used off-the-shelf, and 2) thus typically lag behind conversationally pretrained (e.g., via response selection) encoders on conversational tasks such as intent detection (ID). In this work, we propose CON-VFIT, a simple and efficient two-stage procedure which turns any pretrained LM into a universal conversational encoder (after Stage 1 CONVFIT-ing) and task-specialised sentence encoder (after Stage 2). We demonstrate that 1) full-blown conversational pretraining is not required, and that LMs can be quickly transformed into effective conversational encoders with much smaller amounts of unannotated data; 2) pretrained LMs can be fine-tuned into task-specialised sentence encoders, optimised for the fine-grained semantics of a particular task. Consequently, such specialised sentence encoders allow for treating ID as a simple semantic similarity task based on interpretable nearest neighbours retrieval. We validate the robustness and versatility of the CON-VFIT framework with such similarity-based inference on the standard ID evaluation sets: CONVFIT-ed LMs achieve state-of-the-art ID performance across the board, with particular gains in the most challenging, few-shot setups. Stage 1 loss (c, r) = (context, response) Input LM Input LM Pooling Pooling c r ConvFiT: Stage 1 (Behavioral) fine-tuning on Reddit data ConvFiT: Stage 2 Task-based fine-tuning on intent (task) data
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a4c3dfb6-25e0-4e45-84bc-25a7339314f4Cited by top-tier papers6
- New Intent Discovery with Pre-training and Contrastive LearningYuwei Zhang, Haode Zhang, Li-Ming Zhan, Xiao-Ming Wu et al.ACL 2022 · 55 citations
- Unified Dialog Model Pre-training for Task-Oriented Dialog Understanding and GenerationWanwei He, Yinpei Dai, Min Yang, Jian Sun et al.SIGIR 2022 · 41 citations
- Thought-Augmented Planning for LLM-Powered Interactive Recommender AgentHaocheng Yu, Yaxiong Wu, Hao Wang, Wei Guo et al.KDD 2026 · 16 citations
- Multi-Label Intent Detection via Contrastive Task Specialization of Sentence EncodersIvan Vulic, Iñigo Casanueva, Georgios Spithourakis, Avishek Mondal et al.EMNLP 2022 · 6 citations
- Winnie: Task-Oriented Dialog System with Structure-Aware Contrastive Learning and Enhanced Policy PlanningKaizhi Gao, Tianyu Wang, Zhongjing Ma, Suli ZouAAAI 2024 · 2 citations
Builds on17
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat et al.ICML 2020 · 2,937 citations
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 1,305 citations
- Generalization through Memorization: Nearest Neighbor Language ModelsUrvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer et al.ICLR 2020 · 1,038 citations
- Contrastive Learning with Hard Negative SamplesJoshua David Robinson, Ching-Yao Chuang, Suvrit Sra, Stefanie JegelkaICLR 2021 · 999 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
Related papers
- LexFit: Lexical Fine-Tuning of Pretrained Language ModelsIvan Vulic, Edoardo Maria Ponti, Anna Korhonen, Goran GlavasACL 2021
- Condenser: a Pre-training Architecture for Dense RetrievalLuyu Gao, Jamie CallanEMNLP 2021
- Fast, Effective, and Self-Supervised: Transforming Masked Language Models into Universal Lexical and Sentence EncodersFangyu Liu, Ivan Vulic, Anna Korhonen, Nigel CollierEMNLP 2021 · 85 citations
- Unified Multi-modal Pre-training for Few-shot Sentiment Analysis with Prompt-based LearningYang Yu, Dong Zhang, Shoushan LiACM MM 2022 · 44 citations
- ConSERT: A Contrastive Framework for Self-Supervised Sentence Representation TransferYuanmeng Yan, Rumei Li, Sirui Wang, Fuzheng Zhang et al.ACL 2021
