Prompting in the Dark: Assessing Human Performance in Prompt Engineering for Data Labeling When Gold Labels Are Absent
Zeyu He, Saniya Naphade, Ting-Hao 'Kenneth' Huang
Abstract
Millions of users prompt large language models (LLMs) for various tasks, but how good are people at prompt engineering? Do users actually get closer to their desired outcome over multiple iterations of their prompts? These questions are crucial when no gold-standard labels are available to measure progress. This paper investigates a scenario in LLM-powered data labeling, "prompting in the dark," where users iteratively prompt LLMs to label data without using manually-labeled benchmarks. We developed PromptingSheet, a Google Sheets add-on that enables users to compose, revise, and iteratively label data through spreadsheets. Through a study with 20 participants, we found that prompting in the dark was highly unreliable -- only 9 participants improved labeling accuracy after four or more iterations. Automated prompt optimization tools like DSPy also struggled when few gold labels were available. Our findings highlight the importance of gold labels and the needs, as well as the risks, of automated support in human prompt engineering, providing insights for future tool design.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 41a03f75-643f-4957-97c1-39dec770653eCited by top-tier papers2
- Measurement as Bricolage: Examining How Data Scientists Construct Target Variables for Predictive Modeling TasksLuke Guerdan, Devansh Saxena, Stevie Chancellor, Zhiwei Steven Wu et al.CSCW 2025 · 3 citations
- Data-Prompt Co-Evolution: Growing Test Sets to Refine LLM BehaviorMinjae Lee, Minsuk KahngCHI 2026 · 1 citation
Builds on21
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 1,792 citations
- To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-makingZana Buçinca, Maja Barbara Malaya, Krzysztof Z. GajosCSCW 2021 · 962 citations
- Why Johnny Can't Prompt: How Non-AI Experts Try (and Fail) to Design LLM PromptsJ. D. Zamfirescu-Pereira, Richmond Y. Wong, Bjoern Hartmann, Qian YangCHI 2023 · 892 citations
- Explanations Can Reduce Overreliance on AI Systems During Decision-MakingHelena Vasconcelos, Matthew Jörke, Madeleine Grunde-McLaughlin, Tobias Gerstenberg et al.CSCW 2023 · 362 citations
- Can Large Language Models Be an Alternative to Human Evaluations?David Cheng-Han Chiang, Hung-yi LeeACL 2023 · 254 citations
Related papers
- GLaPE: Gold Label-agnostic Prompt Evaluation for Large Language ModelsXuanchang Zhang, Zhuosheng Zhang, Hai ZhaoEMNLP 2024 · 3 citations
- Large Language Models as OptimizersChengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu et al.ICLR 2024 · 817 citations
- CoPrompt: Supporting Prompt Sharing and Referring in Collaborative Natural Language ProgrammingLi Feng, Ryan Yen, Yuzhe You, Mingming Fan et al.CHI 2024 · 28 citations
- Automatic Prompt Optimization with "Gradient Descent" and Beam SearchReid Pryzant, Dan Iter, Jerry Li, Yin Tat Lee et al.EMNLP 2023 · 137 citations
- PRompt Optimization in Multi-Step Tasks (PROMST): Integrating Human Feedback and Heuristic-based SamplingYongchao Chen, Jacob Arkin, Yilun Hao, Yang Zhang et al.EMNLP 2024 · 6 citations
