Towards Dataset-Scale and Feature-Oriented Evaluation of Text Summarization in Large Language Model Prompts
Sam Yu-Te Lee, Aryaman Bahukhandi, Dongyu Liu, Kwan-Liu Ma
Abstract
Recent advancements in Large Language Models (LLMs) and Prompt Engineering have made chatbot customization more accessible, significantly reducing barriers to tasks that previously required programming skills. However, prompt evaluation, especially at the dataset scale, remains complex due to the need to assess prompts across thousands of test instances within a dataset. Our study, based on a comprehensive literature review and pilot study, summarized five critical challenges in prompt evaluation. In response, we introduce a feature-oriented workflow for systematic prompt evaluation. In the context of text summarization, our workflow advocates evaluation with summary characteristics (feature metrics) such as complexity, formality, or naturalness, instead of using traditional quality metrics like ROUGE. This design choice enables a more user-friendly evaluation of prompts, as it guides users in sorting through the ambiguity inherent in natural language. To support this workflow, we introduce Awesum, a visual analytics system that facilitates identifying optimal prompt refinements for text summarization through interactive visualizations, featuring a novel Prompt Comparator design that employs a BubbleSet-inspired design enhanced by dimensionality reduction techniques. We evaluate the effectiveness and general applicability of the system with practitioners from various domains and found that (1) our design helps overcome the learning curve for non-technical people to conduct a systematic evaluation of summarization prompts, and (2) our feature-oriented workflow has the potential to generalize to other NLG and image-generation tasks. For future works, we advocate moving towards feature-oriented evaluation of LLM prompts and discuss unsolved challenges in terms of human-agent interaction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- RAGTrace: Understanding and Refining Retrieval-Generation Dynamics in Retrieval-Augmented GenerationSizhe Cheng, Jiaping Li, Huanchen Wang, Yuxin MaUIST 2025 · 6 citations
- Understanding Large Language Model Behaviors Through Interactive Counterfactual Generation and AnalysisFurui Cheng, Vilém Zouhar, Robin Shing Moon Chan, Daniel Fürst et al.IEEE VIS 2025 · 5 citations
- The Evolving Duet of Two Modalities: A Survey on Integrating Text and Visualization for Data CommunicationXingyu Lan, Xi Li, Yixing Zhang, Mengqin Cheng et al.CHI 2026 · 1 citation
- ZipLJP: Zipped Information Processor for Legal Judgment PredictionFanghao Lou, Qiqi Wang, Guanyu Chen, Kaiqi Zhao et al.AAAI 2026
Builds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order SensitivityYao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel et al.ACL 2022 · 1,494 citations
- Why Johnny Can't Prompt: How Non-AI Experts Try (and Fail) to Design LLM PromptsJ. D. Zamfirescu-Pereira, Richmond Y. Wong, Bjoern Hartmann, Qian YangCHI 2023 · 892 citations
Related papers
- One Prompt To Rule Them All: LLMs for Opinion Summary EvaluationTejpalsingh Siledar, Swaroop Nath, Sankara Sri Raghava Ravindra Muddu, Rupasai Rangaraju et al.ACL 2024
- FineSurE: Fine-grained Summarization Evaluation using LLMsHwanjun Song, Hang Su, Igor Shalyminov, Jason Cai et al.ACL 2024
- CriSPO: Multi-Aspect Critique-Suggestion-guided Automatic Prompt Optimization for Text GenerationHan He, Qianchu Liu, Lei Xu, Chaitanya Shivade et al.AAAI 2025 · 13 citations
- Just Adjust One Prompt: Enhancing In-Context Dialogue Scoring via Constructing the Optimal Subgraph of Demonstrations and PromptsJiashu Pu, Ling Cheng, Lu Fan, Tangjie Lv et al.EMNLP 2023 · 2 citations
- PrExMe! Large Scale Prompt Exploration of Open Source LLMs for Machine Translation and Summarization EvaluationChristoph Leiter, Steffen EgerEMNLP 2024 · 5 citations
