Automated Summarization of Stack Overflow Posts
Bonan Kou, Muhao Chen, Tianyi Zhang
摘要
Software developers often resort to Stack Overflow (SO) to fill their programming needs. Given the abundance of relevant posts, navigating them and comparing different solutions is tedious and time-consuming. Recent work has proposed to automatically summarize SO posts to concise text to facilitate the navigation of SO posts. However, these techniques rely only on information retrieval methods or heuristics for text summarization, which is insufficient to handle the ambiguity and sophistication of natural language. This paper presents a deep learning based framework called Assortfor SO post summarization. Assortincludes two complementary learning methods, <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex> and <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex> , to address the lack of labeled training data for SO post summarization. <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex> is designed to directly train a novel ensemble learning model with BERT embeddings and domain-specific features to account for the unique characteristics of SO posts. By contrast, <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex> is designed to reuse pre-trained models while addressing the domain shift challenge when no training data is present (i.e., zero-shot learning). Both <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex> and <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex> outperform six existing techniques by at least 13% and 7% respectively in terms of the F1 score. Furthermore, a human study shows that participants significantly preferred summaries generated by <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex> and <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex> over the best baseline, while the preference difference between <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex> and <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex> was small.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Is Stack Overflow Obsolete? An Empirical Study of the Characteristics of ChatGPT Answers to Stack Overflow QuestionsSamia Kabir, David N. Udo-Imeh, Bonan Kou, Tianyi ZhangCHI 2024 · 被引用 149 次
- Unsupervised Large Language Model Alignment for Information Retrieval via Contrastive FeedbackQian Dong, Yiding Liu, Qingyao Ai, Zhijing Wu 等SIGIR 2024 · 被引用 9 次
- Automating API Documentation from Crowdsourced KnowledgeBonan Kou, Zijie Zhou, Muhao Chen, Tianyi ZhangICSE 2026
- A Comparison of Conversational Models and Humans in Answering Technical Questions: the Firefox CaseJoão Correia, Daniel Coutinho, Marco Castelluccio, Caio Barbosa 等ICSE 2026
它引用的顶会 Paper4
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Evaluating the Factual Consistency of Abstractive Text SummarizationWojciech Kryscinski, Bryan McCann, Caiming Xiong, Richard SocherEMNLP 2020 · 被引用 67 次
- Automatic Solution Summarization for Crash BugsHaoye Wang, Xin Xia, David Lo, John C. Grundy 等ICSE 2021 · 被引用 14 次
- Code and Named Entity Recognition in StackOverflowJeniya Tabassum, Mounica Maddela, Wei Xu, Alan RitterACL 2020 · 被引用 9 次
相关 Paper
- Traceability Transformed: Generating more Accurate Links with Pre-Trained BERT ModelsJinfeng Lin, Yalin Liu, Qingkai Zeng, Meng Jiang 等ICSE 2021 · 被引用 124 次
- CLEAR: Contrastive Learning for API RecommendationMoshi Wei, Nima Shiri Harzevili, Yuchao Huang, Junjie Wang 等ICSE 2022 · 被引用 43 次
- PRCBERT: Prompt Learning for Requirement Classification using BERT-based Pretrained Language ModelsXianchang Luo, Yinxing Xue, Zhenchang Xing, Jiamou SunASE 2022 · 被引用 72 次
- Improving API Knowledge Discovery with ML: A Case Study of Comparable API MethodsDaye Nam, Brad A. Myers, Bogdan Vasilescu, Vincent J. HellendoornICSE 2023 · 被引用 6 次
- AUTOSUMM: Automatic Model Creation for Text SummarizationSharmila Reddy Nangi, Atharv Tyagi, Jay Mundra, Sagnik Mukherjee 等EMNLP 2021 · 被引用 1 次
