ChainForge: A Visual Toolkit for Prompt Engineering and LLM Hypothesis Testing
Ian Arawjo, Chelse Swoopes, Priyan Vaithilingam, Martin Wattenberg, Elena L. Glassman
Abstract
Evaluating outputs of large language models (LLMs) is challenging, requiring making—and making sense of—many responses. Yet tools that go beyond basic prompting tend to require knowledge of programming APIs, focus on narrow domains, or are closed-source. We present ChainForge, an open-source visual toolkit for prompt engineering and on-demand hypothesis testing of text generation LLMs. ChainForge provides a graphical interface for comparison of responses across models and prompt variations. Our system was designed to support three tasks: model selection, prompt template design, and hypothesis testing (e.g., auditing). We released ChainForge early in its development and iterated on its design with academics and online users. Through in-lab and interview studies, we find that a range of people could use ChainForge to investigate hypotheses that matter to them, including in real-world settings. We identify three modes of prompt engineering and LLM hypothesis testing: opportunistic exploration, limited evaluation, and iterative refinement.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fa993874-59bd-486d-9979-021ccde134f3Cited by top-tier papers58
- Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human PreferencesShreya Shankar, J. D. Zamfirescu-Pereira, Bjoern Hartmann, Aditya G. Parameswaran et al.UIST 2024 · 143 citations
- DirectGPT: A Direct Manipulation Interface to Interact with Large Language ModelsDamien Masson, Sylvain Malacria, Géry Casiez, Daniel VogelCHI 2024 · 104 citations
- TeachTune: Reviewing Pedagogical Agents Against Diverse Student Profiles with Simulated StudentsHyoungwook Jin, Minju Yoo, Jeongeon Park, Yokyung Lee et al.CHI 2025 · 58 citations
- Supporting Sensemaking of Large Language Model Outputs at ScaleKaty Ilonka Gero, Chelse Swoopes, Ziwei Gu, Jonathan K. Kummerfeld et al.CHI 2024 · 52 citations
- Understanding the LLM-ification of CHI: Unpacking the Impact of LLMs at CHI through a Systematic Literature ReviewRock Yuren Pang, Hope Schroeder, Kynnedy Simone Smith, Solon Barocas et al.CHI 2025 · 51 citations
Builds on13
- Why Johnny Can't Prompt: How Non-AI Experts Try (and Fail) to Design LLM PromptsJ. D. Zamfirescu-Pereira, Richmond Y. Wong, Bjoern Hartmann, Qian YangCHI 2023 · 892 citations
- AI Chains: Transparent and Controllable Human-AI Interaction by Chaining Large Language Model PromptsTongshuang Wu, Michael Terry, Carrie Jun CaiCHI 2022 · 465 citations
- Grounded Copilot: How Programmers Interact with Code-Generating ModelsShraddha Barke, Michael B. James, Nadia PolikarpovaOOPSLA 2023 · 408 citations
- Large Language Models are Human-Level Prompt EngineersYongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster et al.ICLR 2023 · 297 citations
- Interactive and Visual Prompt Engineering for Ad-hoc Task Adaptation with Large Language ModelsHendrik Strobelt, Albert Webson, Victor Sanh, Benjamin Hoover et al.IEEE VIS 2022 · 191 citations
Related papers
- ChainBuddy: An AI-assisted Agent System for Generating LLM PipelinesJingyue Zhang, Ian ArawjoCHI 2025 · 13 citations
- LLM Comparator: Interactive Analysis of Side-by-Side Evaluation of Large Language ModelsMinsuk Kahng, Ian Tenney, Mahima Pushkarna, Michael Xieyang Liu et al.IEEE VIS 2024 · 23 citations
- EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined CriteriaTae Soo Kim, Yoonjoo Lee, Jamin Shin, Young-Ho Kim et al.CHI 2024 · 81 citations
- From Assistant to Independent Developer — Are GPTs Ready for Software Development?Dezhi Ran, Yuan Cao, Mengzhou Wu, Simin Chen et al.ICLR 2026 · 4 citations
- CoPrompt: Supporting Prompt Sharing and Referring in Collaborative Natural Language ProgrammingLi Feng, Ryan Yen, Yuzhe You, Mingming Fan et al.CHI 2024 · 28 citations
