Answer Summarization for Technical Queries: Benchmark and New Approach
Chengran Yang, Bowen Xu, Ferdian Thung, Yucen Shi, Ting Zhang, Zhou Yang, Xin Zhou, Jieke Shi, Junda He, DongGyun Han, David Lo
Abstract
Prior studies have demonstrated that approaches to generate an answer summary for a given technical query in Software Question and Answer (SQA) sites are desired. We find that existing approaches are assessed solely through user studies. Hence, a new user study needs to be performed every time a new approach is introduced; this is time-consuming, slows down the development of the new approach, and results from different user studies may not be comparable to each other. There is a need for a benchmark with ground truth summaries as a complement assessment through user studies. Unfortunately, such a benchmark is non-existent for answer summarization for technical queries from SQA sites.
To fill the gap, we manually construct a high-quality benchmark to enable automatic evaluation of answer summarization for the technical queries for SQA sites. It contains 111 query-summary pairs extracted from 382 Stack Overflow answers with 2,014 sentence candidates. Using the benchmark, we comprehensively evaluate the performance of existing approaches and find that there is still a big room for improvements.
Motivated by the results, we propose a new approach Tech-SumBot with three key modules:1) Usefulness Ranking module; 2) Centrality Estimation module; and 3) Redundancy Removal module. We evaluate TechSumBot in both automatic (i.e., using our benchmark) and manual (i.e., via a user study) manners. The results from both evaluations consistently demonstrate that TechSum-Bot outperforms the best performing baseline approaches from both SE and NLP domains by a large margin, i.e., 10.83%-14.90%, 32.75%-36.59%, and 12.61%-17.54%, in terms of ROUGE-1, ROUGE-2, and ROUGE-L on automatic evaluation, and 5.79%-9.23% and 17.03%-17.68%, in terms of average usefulness and diversity score on human evaluation. This highlights that automatic evaluation on our benchmark can uncover findings similar to the ones found through user studies. More importantly, the automatic evaluation
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e7f346eb-1c2d-4d49-b2d1-67ca8bcc874fCited by top-tier papers4
- Unveiling Memorization in Code ModelsZhou Yang, Zhipeng Zhao, Chenyu Wang, Jieke Shi et al.ICSE 2024 · 33 citations
- AdaptAgent: A Multi-agent, Domain-Guided Reasoning Framework for Code AdaptationXiaokai Rong, Hridya Dhulipala, Aashish Yadavally, Tien N. NguyenISSTA 2026
- Cracking Query Bottlenecks: Towards Efficiency-Oriented Text-to-SQL GenerationLi Lin, Yunfeng Shen, Lingfeng Bao, Rongxin Wu et al.ISSTA 2026
- Think Like Human Developers: Harnessing Community Knowledge for Structured Code ReasoningChengran Yang, Zhensu Sun, Hong Jin Kang, Jieke Shi et al.ICSE 2026
Builds on7
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- Extractive Summarization as Text MatchingMing Zhong, Pengfei Liu, Yiran Chen, Danqing Wang et al.ACL 2020 · 410 citations
- TANDA: Transfer and Adapt Pre-Trained Transformer Models for Answer Sentence SelectionSiddhant Garg, Thuy Vu, Alessandro MoschittiAAAI 2020 · 229 citations
- Coarse-to-Fine Query Focused Multi-Document SummarizationYumo Xu, Mirella LapataEMNLP 2020 · 76 citations
- CLEAR: Contrastive Learning for API RecommendationMoshi Wei, Nima Shiri Harzevili, Yuchao Huang, Junjie Wang et al.ICSE 2022 · 43 citations
Related papers
- Summarizing Community-based Question-Answer PairsTing-Yao Hsu, Yoshi Suhara, Xiaolan WangEMNLP 2022 · 5 citations
- SEM-F1: an Automatic Way for Semantic Evaluation of Multi-Narrative Overlap Summaries at ScaleNaman Bansal, Mousumi Akter, Shubhra Kanti Karmaker SantuEMNLP 2022 · 2 citations
- CBench: Towards Better Evaluation of Question Answering Over Knowledge GraphsAbdelghny Orogat, Isabelle Liu, Ahmed El-RobyVLDB 2021 · 17 citations
- Unsupervised Reference-Free Summary Quality Evaluation via Contrastive LearningHanlu Wu, Tengfei Ma, Lingfei Wu, Tariro Manyumwa et al.EMNLP 2020 · 47 citations
- ASQA: Factoid Questions Meet Long-Form AnswersIvan Stelmakh, Yi Luan, Bhuwan Dhingra, Ming-Wei ChangEMNLP 2022 · 51 citations
