ConvFiT: Conversational Fine-Tuning of Pretrained Language Models
Ivan Vulic, Pei-Hao Su, Samuel Coope, Daniela Gerz, Pawel Budzianowski, Iñigo Casanueva, Nikola Mrksic, Tsung-Hsien Wen
摘要
Transformer-based language models (LMs) pretrained on large text collections are proven to store a wealth of semantic knowledge. However, 1) they are not effective as sentence encoders when used off-the-shelf, and 2) thus typically lag behind conversationally pretrained (e.g., via response selection) encoders on conversational tasks such as intent detection (ID). In this work, we propose CON-VFIT, a simple and efficient two-stage procedure which turns any pretrained LM into a universal conversational encoder (after Stage 1 CONVFIT-ing) and task-specialised sentence encoder (after Stage 2). We demonstrate that 1) full-blown conversational pretraining is not required, and that LMs can be quickly transformed into effective conversational encoders with much smaller amounts of unannotated data; 2) pretrained LMs can be fine-tuned into task-specialised sentence encoders, optimised for the fine-grained semantics of a particular task. Consequently, such specialised sentence encoders allow for treating ID as a simple semantic similarity task based on interpretable nearest neighbours retrieval. We validate the robustness and versatility of the CON-VFIT framework with such similarity-based inference on the standard ID evaluation sets: CONVFIT-ed LMs achieve state-of-the-art ID performance across the board, with particular gains in the most challenging, few-shot setups. Stage 1 loss (c, r) = (context, response) Input LM Input LM Pooling Pooling c r ConvFiT: Stage 1 (Behavioral) fine-tuning on Reddit data ConvFiT: Stage 2 Task-based fine-tuning on intent (task) data
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- New Intent Discovery with Pre-training and Contrastive LearningYuwei Zhang, Haode Zhang, Li-Ming Zhan, Xiao-Ming Wu 等ACL 2022 · 被引用 55 次
- Unified Dialog Model Pre-training for Task-Oriented Dialog Understanding and GenerationWanwei He, Yinpei Dai, Min Yang, Jian Sun 等SIGIR 2022 · 被引用 41 次
- Thought-Augmented Planning for LLM-Powered Interactive Recommender AgentHaocheng Yu, Yaxiong Wu, Hao Wang, Wei Guo 等KDD 2026 · 被引用 16 次
- Multi-Label Intent Detection via Contrastive Task Specialization of Sentence EncodersIvan Vulic, Iñigo Casanueva, Georgios Spithourakis, Avishek Mondal 等EMNLP 2022 · 被引用 6 次
- Winnie: Task-Oriented Dialog System with Structure-Aware Contrastive Learning and Enhanced Policy PlanningKaizhi Gao, Tianyu Wang, Zhongjing Ma, Suli ZouAAAI 2024 · 被引用 2 次
它引用的顶会 Paper17
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat 等ICML 2020 · 被引用 2,937 次
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 被引用 1,305 次
- Generalization through Memorization: Nearest Neighbor Language ModelsUrvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer 等ICLR 2020 · 被引用 1,038 次
- Contrastive Learning with Hard Negative SamplesJoshua David Robinson, Ching-Yao Chuang, Suvrit Sra, Stefanie JegelkaICLR 2021 · 被引用 999 次
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
相关 Paper
- LexFit: Lexical Fine-Tuning of Pretrained Language ModelsIvan Vulic, Edoardo Maria Ponti, Anna Korhonen, Goran GlavasACL 2021
- Condenser: a Pre-training Architecture for Dense RetrievalLuyu Gao, Jamie CallanEMNLP 2021
- Fast, Effective, and Self-Supervised: Transforming Masked Language Models into Universal Lexical and Sentence EncodersFangyu Liu, Ivan Vulic, Anna Korhonen, Nigel CollierEMNLP 2021 · 被引用 85 次
- Unified Multi-modal Pre-training for Few-shot Sentiment Analysis with Prompt-based LearningYang Yu, Dong Zhang, Shoushan LiACM MM 2022 · 被引用 44 次
- ConSERT: A Contrastive Framework for Self-Supervised Sentence Representation TransferYuanmeng Yan, Rumei Li, Sirui Wang, Fuzheng Zhang 等ACL 2021
