Tuning Less, Prompting More: In-Context Preference Learning Pipeline for Natural Language Transformation
Shuyun Yang, Yan Zhang, Zhengmao Ye, Lei Duan, Mingjie Tang
摘要
Natural language transformation (NLT) tasks, such as machine translation (MT) and text style transfer (TST), require models to generate accurate and contextually appropriate outputs. However, existing approaches face significant challenges, including the computational costs of leveraging large pre-trained models and the limited generalization ability of finetuned smaller models. In this paper, we propose a novel framework that combines the flexibility of prompting with the cost-effectiveness of fine-tuning. Our method enhances smaller models by integrating In-Context Examples (ICE) from retrieval, enabling the model to better capture contextual information and align with userlevel preferences. We further improve performance through hierarchical contrastive learning and dynamic preference inference mechanisms. Experimental results demonstrate that our approach outperforms existing methods, such as Supervised Fine Tuning (SFT), Direct Preference Optimization (DPO), and Contrastive Preference Optimization (CPO), across both MT and TST tasks, providing a more efficient solution for resource-constrained environments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine TranslationHaoran Xu, Amr Sharaf, Yunmo Chen, Weiting Tan 等ICML 2024 · 被引用 447 次
- A Paradigm Shift in Machine Translation: Boosting Translation Performance of Large Language ModelsHaoran Xu, Young Jin Kim, Amr Sharaf, Hany Hassan AwadallaICLR 2024 · 被引用 122 次
- COMET: A Neural Framework for MT EvaluationRicardo Rei, Craig Stewart, Ana C. Farinha, Alon LavieEMNLP 2020 · 被引用 6 次
- Robust Prompt Optimization for Large Language Models Against Distribution ShiftsMoxin Li, Wenjie Wang, Fuli Feng, Yixin Cao 等EMNLP 2023 · 被引用 5 次
相关 Paper
- Two-stage LLM Fine-tuning with Less Specialization and More GeneralizationYihan Wang, Si Si, Daliang Li, Michal Lukasik 等ICLR 2024 · 被引用 45 次
- RLPrompt: Optimizing Discrete Text Prompts with Reinforcement LearningMingkai Deng, Jianyu Wang, Cheng-Ping Hsieh, Yihan Wang 等EMNLP 2022 · 被引用 141 次
- Adaptive Prompt Routing for Arbitrary Text Style Transfer with Pre-trained Language ModelsQingyi Liu, Jinghui Qin, Wenxuan Ye, Hao Mou 等AAAI 2024 · 被引用 9 次
- Teaching Large Language Models to Translate with ComparisonJiali Zeng, Fandong Meng, Yongjing Yin, Jie ZhouAAAI 2024 · 被引用 74 次
- Text Style Transfer with Contrastive Transfer Pattern MiningJingxuan Han, Quan Wang, Licheng Zhang, Weidong Chen 等ACL 2023 · 被引用 6 次
