LiSum: Open Source Software License Summarization with Multi-Task Learning
Linyu Li, Sihan Xu, Yang Liu, Ya Gao, Xiangrui Cai, Jiarun Wu, Wenli Song, Zheli Liu
Abstract
Open source software (OSS) licenses regulate the conditions under which users can reuse, modify, and distribute the software legally. However, there exist various OSS licenses in the community, written in a formal language, which are typically long and complicated to understand. In this paper, we conducted a 661-participants online survey to investigate the perspectives and practices of developers towards OSS licenses. The user study revealed an indeed need for an automated tool to facilitate license understanding. Motivated by the user study and the fast growth of licenses in the community, we propose the first study towards automated license summarization. Specifically, we released the first high quality text summarization dataset and designed two tasks, i.e., license text summarization (LTS), aiming at generating a relatively short summary for an arbitrary license, and license term classification (LTC), focusing on the attitude inference towards a predefined set of key license terms (e.g., Distribute). Aiming at the two tasks, we present LiSum, a multi-task learning method to help developers overcome the obstacles of understanding OSS licenses. Comprehensive experiments demonstrated that the proposed jointly training objective boosted the performance on both tasks, surpassing state-of-the-art baselines with gains of at least 5 points w.r.t. F1 scores of four summarization metrics and achieving 95.13% micro average F1 score for classification simultaneously. We released all the datasets, the replication package, and the questionnaires for the community.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d9812dc8-1c6f-414d-825b-905ff4480a45Cited by top-tier papers1
Ask how each one uses itBuilds on8
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 2,453 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Extractive Summarization as Text MatchingMing Zhong, Pengfei Liu, Yiran Chen, Danqing Wang et al.ACL 2020 · 410 citations
- BRIO: Bringing Order to Abstractive SummarizationYixin Liu, Pengfei Liu, Dragomir R. Radev, Graham NeubigACL 2022 · 329 citations
Related papers
- SUMMIT: Scaffolding Open Source Software Issue Discussion Through SummarizationSaskia Gilmer, Avinash Bhat, Shuvam Shah, Kevin Cherry et al.CSCW 2023 · 10 citations
- LiResolver: License Incompatibility Resolution for Open Source SoftwareSihan Xu, Ya Gao, Lingling Fan, Linyu Li et al.ISSTA 2023 · 6 citations
- RNSum: A Large-Scale Dataset for Automatic Release Note Generation via Commit Logs SummarizationHisashi Kamezawa, Noriki Nishida, Nobuyuki Shimizu, Takashi Miyazaki et al.ACL 2022
- Hidden Licensing Risks in the PTMware EcosystemBo Wang, Yueyang Chen, Jieke Shi, Minghui Li et al.ISSTA 2026
- Your "Notice" Is Missing: Detecting and Fixing Violations of Modification Terms in Open Source Licenses during ForkingKaifeng Huang, Yingfeng Xia, Bihuan Chen, Siyang He et al.ISSTA 2024 · 3 citations
