What's the Difference? Supporting Users in Identifying the Effects of Prompt and Model Changes Through Token Patterns
Michael A. Hedderich, Anyi Wang, Raoyuan Zhao, Florian Eichin, Jonas Fischer, Barbara Plank
摘要
Prompt engineering for large language models is challenging, as even small prompt perturbations or model changes can significantly impact the generated output texts. Existing evaluation methods of LLM outputs, either automated metrics or human evaluation, have limitations, such as providing limited insights or being labor-intensive. We propose Spotlight, a new approach that combines both automation and human analysis. Based on data mining techniques, we automatically distinguish between random (decoding) variations and systematic differences in language model outputs. This process provides token patterns that describe the systematic differences and guide the user in manually analyzing the effects of their prompts and changes in models efficiently. We create three benchmarks to quantitatively test the reliability of token pattern extraction methods and demonstrate that our approach provides new insights into established prompt data. From a human-centric perspective, through demonstration studies and a user study, we show that our token pattern approach helps users understand the systematic differences of language model outputs. We are further able to discover relevant differences caused by prompt and model changes (e.g. related to gender or culture), thus supporting the prompt engineering process and human-centric model behavior research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Why Johnny Can't Prompt: How Non-AI Experts Try (and Fail) to Design LLM PromptsJ. D. Zamfirescu-Pereira, Richmond Y. Wong, Bjoern Hartmann, Qian YangCHI 2023 · 被引用 892 次
- TaleBrush: Sketching Stories with Generative Pretrained Language ModelsJohn Joon Young Chung, Wooseok Kim, Kang Min Yoo, Hwaran Lee 等CHI 2022 · 被引用 202 次
- RAIN: Your Language Models Can Align Themselves without FinetuningYuhui Li, Fangyun Wei, Jinjing Zhao, Chao Zhang 等ICLR 2024 · 被引用 171 次
- Adaptive Testing and Debugging of NLP ModelsMarco Túlio Ribeiro, Scott M. LundbergACL 2022 · 被引用 99 次
相关 Paper
- Towards Dataset-Scale and Feature-Oriented Evaluation of Text Summarization in Large Language Model PromptsSam Yu-Te Lee, Aryaman Bahukhandi, Dongyu Liu, Kwan-Liu MaIEEE VIS 2024 · 被引用 18 次
- Statistical Hypothesis Testing for Auditing Robustness in Language ModelsPaulius Rauba, Qiyao Wei, Mihaela van der SchaarICML 2025
- PrExMe! Large Scale Prompt Exploration of Open Source LLMs for Machine Translation and Summarization EvaluationChristoph Leiter, Steffen EgerEMNLP 2024 · 被引用 5 次
- Teach Better or Show Smarter? On Instructions and Exemplars in Automatic Prompt OptimizationXingchen Wan, Ruoxi Sun, Hootan Nakhost, Sercan Ö. ArikNeurIPS 2024 · 被引用 35 次
- Improving Diversity of Demographic Representation in Large Language Models via Collective-Critiques and Self-VotingPreethi Lahoti, Nicholas Blumm, Xiao Ma, Raghavendra Kotikalapudi 等EMNLP 2023 · 被引用 14 次
