Lexara: A User-Centered Toolkit for Evaluating Large Language Models for Conversational Visual Analytics
Srishti Palani, Vidya Setlur
Abstract
Large Language Models (LLMs) are transforming Conversational Visual Analytics (CVA) by enabling data analysis through natural language. However, evaluating LLMs for CVA remains a challenge: requiring programming expertise, overlooking real-world complexity, and lacking interpretable metrics for multi-format (visualizations and text) outputs. Through interviews with 22 CVA developers and 16 end-users, we identified use cases, evaluation criteria and workflows. We present Lexara, a user-centered evaluation toolkit for CVA that operationalizes these insights into: (i) test cases spanning real-world scenarios; (ii) interpretable metrics covering visualization quality (data fidelity, semantic alignment, functional correctness, design clarity) and language quality (factual grounding, analytical reasoning, conversational coherence) using rule-based and LLM-as-a-Judge methods; and (iii) an interactive toolkit enabling experimental setup and multi-format and multilevel exploration of results without programming expertise. We conducted a two-week diary study with six CVA developers, drawn from our initial cohort of 22. Their feedback demonstrated Lexara's effectiveness for guiding appropriate model and prompt selection.
• Human-centered computing → Visualization systems and tools; Interactive systems and tools.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5ab9e095-1dcc-412d-b7d6-07a6fba82b60Builds on22
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- LIMA: Less Is More for AlignmentChunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer et al.NeurIPS 2023 · 1,486 citations
- Re-examining Whether, Why, and How Human-AI Interaction Is Uniquely Difficult to DesignQian Yang, Aaron Steinfeld, Carolyn P. Rosé, John ZimmermanCHI 2020 · 604 citations
- Least-to-Most Prompting Enables Complex Reasoning in Large Language ModelsDenny Zhou, Nathanael Schärli, Le Hou, Jason Wei et al.ICLR 2023 · 318 citations
- Interactive and Visual Prompt Engineering for Ad-hoc Task Adaptation with Large Language ModelsHendrik Strobelt, Albert Webson, Victor Sanh, Benjamin Hoover et al.IEEE VIS 2022 · 191 citations
Related papers
- LLM Comparator: Interactive Analysis of Side-by-Side Evaluation of Large Language ModelsMinsuk Kahng, Ian Tenney, Mahima Pushkarna, Michael Xieyang Liu et al.IEEE VIS 2024 · 23 citations
- Closing the Feedback Loop in Text2Vis: Refining Visualization with Vision-Language ModelsShengze Shi, Tao Ren, Guoliang Zhu, Guan Dong Feng et al.ACM MM 2025 · 2 citations
- VisEval: A Benchmark for Data Visualization in the Era of Large Language ModelsNan Chen, Yuge Zhang, Jiahang Xu, Kan Ren et al.IEEE VIS 2024 · 44 citations
- Explainable XR: Understanding User Behaviors of XR Environments Using LLM-Assisted Analytics FrameworkYoonsang Kim, Zainab Aamir, Mithilesh Kumar Singh, Saeed Boorboor et al.IEEE VR 2025 · 27 citations
- Towards Dataset-Scale and Feature-Oriented Evaluation of Text Summarization in Large Language Model PromptsSam Yu-Te Lee, Aryaman Bahukhandi, Dongyu Liu, Kwan-Liu MaIEEE VIS 2024 · 18 citations
