Designing Staged Evaluation Workflows for LLMs: Integrating Domain Experts, Lay Users, and Model-Generated Evaluation Criteria
Annalisa Szymanski, Simret Araya Gebreegziabher, Oghenemaro Anuyah, Ronald A. Metoyer, Toby Jia-Jun Li
摘要
Large Language Models (LLMs) are increasingly utilized for domainspecific tasks, yet evaluating their outputs remains challenging. A common strategy is to apply evaluation criteria to assess alignment with domain-specific standards, yet little is understood about how criteria differ across sources or where each type is most useful in the evaluation process. This study investigates criteria developed by domain experts, lay users, and LLMs to identify their complementary roles within an evaluation workflow. Results show that experts produce fact-based criteria with long-term value, lay users emphasize usability with a shorter-term focus, and LLMs target procedural checks for immediate task requirements. We also examine how criteria evolve between a priori and a posteriori phases, noting drift across stages as well as convergence in the a posteriori phase. Based on our observations, we propose design guidelines for a staged evaluation workflow combining the complementary strengths of these sources to balance quality, cost, and scalability.
• Human-centered computing → Empirical studies in HCI; Empirical studies in HCI.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Supporting Co-Adaptive Machine Teaching through Human Concept Learning and Cognitive TheoriesSimret Araya Gebreegziabher, Yukun Yang, Elena L. Glassman, Toby Jia-Jun LiCHI 2025 · 被引用 5 次
- Measurement as Bricolage: Examining How Data Scientists Construct Target Variables for Predictive Modeling TasksLuke Guerdan, Devansh Saxena, Stevie Chancellor, Zhiwei Steven Wu 等CSCW 2025 · 被引用 3 次
- "Do I Trust the AI?" Towards Trustworthy AI-Assisted Diagnosis: Understanding User Perception in LLM-Supported Clinical ReasoningYuansong Xu, Yichao Zhu, Haokai Wang, Yuchen Wu 等CHI 2026 · 被引用 1 次
- BloomIntent: Automating Search Evaluation with LLM-Generated Fine-Grained User IntentsYoonseo Choi, Eunhye Kim, Hyunwoo Kim, Donghyun Park 等UIST 2025 · 被引用 1 次
- Evalet: Evaluating Large Language Models through Functional FragmentationTae Soo Kim, Heechan Lee, Yoonjoo Lee, Joseph Seering 等CHI 2026 · 被引用 1 次
它引用的顶会 Paper26
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos 等ICML 2024 · 被引用 1,212 次
- Why Johnny Can't Prompt: How Non-AI Experts Try (and Fail) to Design LLM PromptsJ. D. Zamfirescu-Pereira, Richmond Y. Wong, Bjoern Hartmann, Qian YangCHI 2023 · 被引用 892 次
- Design Guidelines for Prompt Engineering Text-to-Image Generative ModelsVivian Liu, Lydia B. ChiltonCHI 2022 · 被引用 586 次
- Self-Play Fine-Tuning Converts Weak Language Models to Strong Language ModelsZixiang Chen, Yihe Deng, Huizhuo Yuan, Kaixuan Ji 等ICML 2024 · 被引用 527 次
- Evaluating Large Language Models in Generating Synthetic HCI Research Data: a Case StudyPerttu Hämäläinen, Mikke Tavast, Anton KunnariCHI 2023 · 被引用 244 次
相关 Paper
- EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined CriteriaTae Soo Kim, Yoonjoo Lee, Jamin Shin, Young-Ho Kim 等CHI 2024 · 被引用 81 次
- Large Language Models in Qualitative Research: Uses, Tensions, and IntentionsHope Schroeder, Marianne Aubin Le Quéré, Casey Randazzo, David Mimno 等CHI 2025 · 被引用 40 次
- Understanding the LLM-ification of CHI: Unpacking the Impact of LLMs at CHI through a Systematic Literature ReviewRock Yuren Pang, Hope Schroeder, Kynnedy Simone Smith, Solon Barocas 等CHI 2025 · 被引用 51 次
- HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria DecompositionYuxuan Liu, Tianchi Yang, Shaohan Huang, Zihan Zhang 等ACL 2024 · 被引用 1 次
- Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human PreferencesShreya Shankar, J. D. Zamfirescu-Pereira, Bjoern Hartmann, Aditya G. Parameswaran 等UIST 2024 · 被引用 143 次
