What Makes a Good Natural Language Prompt?
Do Xuan Long, Duy Dinh, Ngoc-Hai Nguyen, Kenji Kawaguchi, Nancy F. Chen, Shafiq Joty, Min-Yen Kan
摘要
As large language models (LLMs) have progressed towards more human-like and human-AI communications prevalent, prompting has emerged as a decisive component. However, there is limited conceptual consensus on what exactly quantifies natural language prompts. We attempt to address this question by conducting a meta-analysis surveying 150+ promptingrelated papers from leading NLP and AI conferences (2022-2025), and blogs. We propose a property-and human-centric framework for evaluating prompt quality, encompassing 21 properties categorized into six dimensions. We then examine how existing studies assess their impact on LLMs, revealing their imbalanced support across models and tasks, and substantial research gaps. Further, we analyze correlations among properties in high-quality natural language prompts, deriving prompting recommendations. We then empirically explore multiproperty prompt enhancements in reasoning tasks, observing that single-property enhancements often have the greatest impact. Finally, we discover that instruction-tuning on propertyenhanced prompts can result in better reasoning models. Our findings establish a foundation for property-centric prompt evaluation and optimization, bridging the gaps between human-AI communication and opening new prompting research directions 1 . * Equal contribution. Works done during the internship at WING, NUS. 1 Our codes and data will be made publicly available at here. Prompt quality evaluation We begin our study by conducting a comprehensive survey of over 150 papers and blogs. Our methodology is straightforward: we first examine papers published in ACL, EMNLP, NAACL from ACL Anthology 2 , and ICLR, and NeurIPS on 2 https:
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- VISTA: A Test-Time Self-Improving Video Generation AgentDo Xuan Long, Xingchen Wan, Hootan Nakhost, Chen-Yu Lee 等CVPR 2026 · 被引用 30 次
- "I Just Need GPT to Refine My Prompts": Rethinking Onboarding and Help-Seeking with Generative 3D Modelling ToolsKanak Gautam, Poorvi Bhatia, Parmit K. ChilanaCHI 2026 · 被引用 1 次
- UniAPO: Unified Multimodal Automated Prompt OptimizationQipeng Zhu, Yanzhe Chen, Huasong Zhong, Jie Chen 等AAAI 2026
- Hey, ChatGPT, Look at My Work: Using Conversational AI in Requirements Engineering EducationSahar Badihi, Michael Tegegn, Evelien Riddell, Krzysztof Czarnecki 等ICSE 2026
它引用的顶会 Paper72
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Large Language Models Can Be Easily Distracted by Irrelevant ContextFreda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales 等ICML 2023 · 被引用 970 次
相关 Paper
- A Sentiment Consolidation Framework for Meta-Review GenerationMiao Li, Jey Han Lau, Eduard H. HovyACL 2024 · 被引用 3 次
- When Prompt Engineering Meets Software Engineering: CNL-P as Natural and Robust "APIs" for Human-AI InteractionZhenchang Xing, Yang Liu, Zhuo Cheng, Qing Huang 等ICLR 2025
- GLaPE: Gold Label-agnostic Prompt Evaluation for Large Language ModelsXuanchang Zhang, Zhuosheng Zhang, Hai ZhaoEMNLP 2024 · 被引用 3 次
- ZERA: Zero-init Instruction Evolving Refinement Agent - From Zero Instructions to Structured Prompts via Principle-based OptimizationSeungyoun Yi, Minsoo Khang, Sungrae ParkEMNLP 2025
- CogBench: a large language model walks into a psychology labJulian Coda-Forno, Marcel Binz, Jane X. Wang, Eric SchulzICML 2024 · 被引用 60 次
