What Makes a Good Natural Language Prompt?
Do Xuan Long, Duy Dinh, Ngoc-Hai Nguyen, Kenji Kawaguchi, Nancy F. Chen, Shafiq Joty, Min-Yen Kan
Abstract
As large language models (LLMs) have progressed towards more human-like and human-AI communications prevalent, prompting has emerged as a decisive component. However, there is limited conceptual consensus on what exactly quantifies natural language prompts. We attempt to address this question by conducting a meta-analysis surveying 150+ promptingrelated papers from leading NLP and AI conferences (2022-2025), and blogs. We propose a property-and human-centric framework for evaluating prompt quality, encompassing 21 properties categorized into six dimensions. We then examine how existing studies assess their impact on LLMs, revealing their imbalanced support across models and tasks, and substantial research gaps. Further, we analyze correlations among properties in high-quality natural language prompts, deriving prompting recommendations. We then empirically explore multiproperty prompt enhancements in reasoning tasks, observing that single-property enhancements often have the greatest impact. Finally, we discover that instruction-tuning on propertyenhanced prompts can result in better reasoning models. Our findings establish a foundation for property-centric prompt evaluation and optimization, bridging the gaps between human-AI communication and opening new prompting research directions 1 . * Equal contribution. Works done during the internship at WING, NUS. 1 Our codes and data will be made publicly available at here. Prompt quality evaluation We begin our study by conducting a comprehensive survey of over 150 papers and blogs. Our methodology is straightforward: we first examine papers published in ACL, EMNLP, NAACL from ACL Anthology 2 , and ICLR, and NeurIPS on 2 https:
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 82d5c39c-e694-4069-b2e5-30d0bb619a46Cited by top-tier papers4
- VISTA: A Test-Time Self-Improving Video Generation AgentDo Xuan Long, Xingchen Wan, Hootan Nakhost, Chen-Yu Lee et al.CVPR 2026 · 30 citations
- "I Just Need GPT to Refine My Prompts": Rethinking Onboarding and Help-Seeking with Generative 3D Modelling ToolsKanak Gautam, Poorvi Bhatia, Parmit K. ChilanaCHI 2026 · 1 citation
- UniAPO: Unified Multimodal Automated Prompt OptimizationQipeng Zhu, Yanzhe Chen, Huasong Zhong, Jie Chen et al.AAAI 2026
- Hey, ChatGPT, Look at My Work: Using Conversational AI in Requirements Engineering EducationSahar Badihi, Michael Tegegn, Evelien Riddell, Krzysztof Czarnecki et al.ICSE 2026
Builds on72
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Large Language Models Can Be Easily Distracted by Irrelevant ContextFreda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales et al.ICML 2023 · 970 citations
Related papers
- A Sentiment Consolidation Framework for Meta-Review GenerationMiao Li, Jey Han Lau, Eduard H. HovyACL 2024 · 3 citations
- When Prompt Engineering Meets Software Engineering: CNL-P as Natural and Robust "APIs" for Human-AI InteractionZhenchang Xing, Yang Liu, Zhuo Cheng, Qing Huang et al.ICLR 2025
- GLaPE: Gold Label-agnostic Prompt Evaluation for Large Language ModelsXuanchang Zhang, Zhuosheng Zhang, Hai ZhaoEMNLP 2024 · 3 citations
- ZERA: Zero-init Instruction Evolving Refinement Agent - From Zero Instructions to Structured Prompts via Principle-based OptimizationSeungyoun Yi, Minsoo Khang, Sungrae ParkEMNLP 2025
- CogBench: a large language model walks into a psychology labJulian Coda-Forno, Marcel Binz, Jane X. Wang, Eric SchulzICML 2024 · 60 citations
