Surfacing Governing Principles for Chatbots: A Workbench and Comparative Study
Maria Antonietta Grasso, Jisun Park, Jutta Katharina Willamowski, Laurent Besacier, Jos Rozen
Abstract
Trust in Large Language Model chatbots depends not only on what these systems do but also on how their behavior is governed and communicated. We present Trust Mediator, a workbench that supports service owners in authoring and assessing principle sets for LLM-driven chatbots through persona-based exploration and structured scaffolds. To examine this workflow, we use three analytic lenses—specificity, coverage, and coherence—to characterize the principles produced. In an exploratory between-subjects study, we compared manual and assisted principle authoring. Participants in both conditions viewed principles as useful for governing and assessing chatbot behavior. Assisted authoring was generally perceived as more supportive and tended to broaden coverage. Manual authoring required more effort but yielded principles that were significantly more specific. These findings highlight complementary strengths of assisted and manual pathways, illustrating the value of treating principle sets as design objects within governance workflows. Beyond their analytic role in this study, the lenses also suggest opportunities for supporting the construction and inspection of principle sets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on9
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Co-Designing Checklists to Understand Organizational Challenges and Opportunities around Fairness in AIMichael A. Madaio, Luke Stark, Jennifer Wortman Vaughan, Hanna M. WallachCHI 2020 · 428 citations
- Red Teaming Language Models with Language ModelsEthan Perez, Saffron Huang, H. Francis Song, Trevor Cai et al.EMNLP 2022 · 239 citations
- EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined CriteriaTae Soo Kim, Yoonjoo Lee, Jamin Shin, Young-Ho Kim et al.CHI 2024 · 81 citations
- The U in Crypto Stands for Usable: An Empirical Study of User Experience with Mobile Cryptocurrency WalletsArtemij Voskobojnikov, Oliver Wiese, Masoud Mehrabi Koushki, Volker Roth et al.CHI 2021 · 61 citations
Related papers
- The Typing Cure: Experiences with Large Language Model Chatbots for Mental Health SupportInhwa Song, Sachin R. Pendse, Neha Kumar, Munmun De ChoudhuryCSCW 2025 · 41 citations
- Why Johnny Can't Prompt: How Non-AI Experts Try (and Fail) to Design LLM PromptsJ. D. Zamfirescu-Pereira, Richmond Y. Wong, Bjoern Hartmann, Qian YangCHI 2023 · 892 citations
- TACO: Trust Assessment of Large Language Models in Coding Assistance TasksShihao Weng, Yang Feng, Jincheng Li, Yining Yin et al.ICSE 2026
- Supporting Effective Goal Setting with LLM-Based ChatbotsMichel Schimpf, Sebastian Maier, Anton Wyrowski, Lara Christoforakos et al.CHI 2026 · 2 citations
- Understanding User Needs Underlying the Expected Roles of LLM-Based Chatbots in Privacy Decision-MakingJian Jun, Yunjae Josephine Choi, Jeonghoon Han, Sangsu LeeCHI 2026 · 1 citation
