ChainBuddy: An AI-assisted Agent System for Generating LLM Pipelines
Jingyue Zhang, Ian Arawjo
Abstract
As large language models (LLMs) advance, their potential applications have grown significantly. However, it remains difficult to evaluate LLM behavior on user-defined tasks and craft effective pipelines to do so. Many users struggle with where to start, often referred to as the "blank page problem." ChainBuddy, an AI workflow generation assistant built into the ChainForge platform, aims to tackle this issue. From a single prompt or chat, ChainBuddy generates a starter evaluative LLM pipeline in ChainForge aligned to the user’s requirements. ChainBuddy offers a straightforward and user-friendly way to plan and evaluate LLM behavior and make the process less daunting and more accessible across a wide range of possible tasks and use cases. We report a within-subjects user study comparing ChainBuddy to the baseline interface. We find that when using AI assistance, participants with a variety of technical expertise reported a less demanding workload, felt more confident, and produced higher quality pipelines evaluating LLM behavior. However, we also uncover a mismatch between subjective and objective ratings of performance: participants rated their successfulness similarly across conditions, while independent experts rated participant workflows significantly higher with AI assistance. Drawing connections to the Dunning–Kruger effect, we discuss implications for the future design of workflow generation assistants regarding the risk of over-reliance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Creativity in LLM-based Multi-Agent Systems: A SurveyYi-Cheng Lin, Kang-Chieh Chen, Zhe-Yan Li, Tzu-Heng Wu et al.EMNLP 2025 · 3 citations
- Reporting and Reviewing LLM-Integrated Systems in HCI: Challenges and ConsiderationsKarla Felix Navarro, Eugene Syriani, Ian ArawjoCHI 2026 · 1 citation
- Evalet: Evaluating Large Language Models through Functional FragmentationTae Soo Kim, Heechan Lee, Yoonjoo Lee, Joseph Seering et al.CHI 2026 · 1 citation
- Botender: Supporting Communities in Collaboratively Designing AI Agents through Case-Based ProvocationsTzu-Sheng Kuo, Sophia Liu, Quan Ze Chen, Joseph Seering et al.CHI 2026 · 1 citation
- Steering Semantic Data Processing With DocWranglerShreya Shankar, Bhavya Chopra, Mawil Hasan, Stephen Lee et al.UIST 2025
Builds on16
- To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-makingZana Buçinca, Maja Barbara Malaya, Krzysztof Z. GajosCSCW 2021 · 962 citations
- Why Johnny Can't Prompt: How Non-AI Experts Try (and Fail) to Design LLM PromptsJ. D. Zamfirescu-Pereira, Richmond Y. Wong, Bjoern Hartmann, Qian YangCHI 2023 · 892 citations
- Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team PerformanceGagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok et al.CHI 2021 · 713 citations
- AI Chains: Transparent and Controllable Human-AI Interaction by Chaining Large Language Model PromptsTongshuang Wu, Michael Terry, Carrie Jun CaiCHI 2022 · 465 citations
- Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language ModelsLei Wang, Wanyu Xu, Yihuai Lan, Zhiqiang Hu et al.ACL 2023 · 249 citations
Related papers
- ChainForge: A Visual Toolkit for Prompt Engineering and LLM Hypothesis TestingIan Arawjo, Chelse Swoopes, Priyan Vaithilingam, Martin Wattenberg et al.CHI 2024 · 141 citations
- Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human PreferencesShreya Shankar, J. D. Zamfirescu-Pereira, Bjoern Hartmann, Aditya G. Parameswaran et al.UIST 2024 · 143 citations
- Beyond Factuality: A Comprehensive Evaluation of Large Language Models as Knowledge GeneratorsLiang Chen, Yang Deng, Yatao Bian, Zeyu Qin et al.EMNLP 2023 · 21 citations
- The Bitter Lesson of Diffusion Language Models for Agentic Workflows: A Comprehensive Reality CheckQingyu Lu, Liang Ding, Kanjian Zhang, Jinxia Zhang et al.ACL 2026 · 1 citation
- Bridging Human and LLM Judgments: Understanding and Narrowing the GapFelipe Maia Polo, Xinhe Wang, Mikhail Yurochkin, Gongjun Xu et al.NeurIPS 2025 · 7 citations
