Paraphrase Types Elicit Prompt Engineering Capabilities
Jan Philip Wahle, Terry Ruas, Yang Xu, Bela Gipp
Abstract
Much of the success of modern language models depends on finding a suitable prompt to instruct the model. Until now, it has been largely unknown how variations in the linguistic expression of prompts affect these models. This study systematically and empirically evaluates which linguistic features influence models through paraphrase types, i.e., different linguistic changes at particular positions. We measure behavioral changes for five models across 120 tasks and six families of paraphrases (i.e., morphology, syntax, lexicon, lexico-syntax, discourse, and others). We also control for other prompt engineering factors (e.g., prompt length, lexical diversity, and proximity to training data). Our results show a potential for language models to improve tasks when their prompts are adapted in specific paraphrase types (e.g., 6.7% median gain in Mixtral 8x7B; 5.5% in LLaMA 3 8B; cf. Figure 1 ). In particular, changes in morphology and lexicon, i.e., the vocabulary used, showed promise in improving prompts. These findings contribute to developing more robust language models capable of handling variability in linguistic expression. Code: https://github.com/jpwahle/ emnlp24-prompt-paraphrase .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- HalluCitation Matters: Revealing the Impact of Hallucinated References with 300 Hallucinated Papers in ACL ConferencesYusuke Sakai, Hidetaka Kamigaito, Taro WatanabeACL 2026 · 18 citations
- SPaRC: A Spatial Pathfinding Reasoning ChallengeLars Benedikt Kaesberg, Jan Philip Wahle, Terry Ruas, Bela GippEMNLP 2025
- How LLMs Comprehend Temporal Meaning in Narratives: A Case Study in Cognitive Evaluation of LLMsKarin de Langis, Jong Inn Park, Andreas Schramm, Bin Hu et al.ACL 2025
- RCScore: Quantifying Response Consistency in Large Language ModelsDongjun Jang, Youngchae Ahn, Hyopil ShinEMNLP 2025
Builds on13
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein et al.ICML 2021 · 1,843 citations
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated PromptsTaylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace et al.EMNLP 2020 · 1,162 citations
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu et al.ACL 2023 · 540 citations
- Large Language Models are Human-Level Prompt EngineersYongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster et al.ICLR 2023 · 297 citations
- Differentiable Prompt Makes Pre-trained Language Models Better Few-shot LearnersNingyu Zhang, Luoqiu Li, Xiang Chen, Shumin Deng et al.ICLR 2022 · 205 citations
Related papers
- Extracting Linguistic Information from Large Language Models: Syntactic Relations and Derivational KnowledgeTsedeniya Kinfe Temesgen, Marion Di Marco, Alexander FraserEMNLP 2025 · 2 citations
- Paraphrase Types for Generation and DetectionJan Philip Wahle, Bela Gipp, Terry RuasEMNLP 2023 · 7 citations
- Training Prompt Matters: State-Adaptive Optimization for Robust Fine-TuningWenhang Shi, Yiren Chen, Shuqing Bian, Zhe Zhao et al.ICML 2026
- It's Not What You Say, It's How You Say It: Evaluating LLM Responses to Expressions of BeliefKevin Du, Clara Kümpel, Michelle Wastl, Alex WarstadtACL 2026
- Effects of diversity incentives on sample diversity and downstream model performance in LLM-based text augmentationJán Cegin, Branislav Pecher, Jakub Simko, Ivan Srba et al.ACL 2024 · 5 citations
