AI-based Question Answering Assistance for Analyzing Natural-language Requirements
Saad Ezzini, Sallam Abualhaija, Chetan Arora, Mehrdad Sabetzadeh
Abstract
By virtue of being prevalently written in natural language (NL), requirements are prone to various defects, e.g., inconsistency and incompleteness. As such, requirements are frequently subject to quality assurance processes. These processes, when carried out entirely manually, are tedious and may further overlook important quality issues due to time and budget pressures. In this paper, we propose QAssist - a question-answering (QA) approach that provides automated assistance to stakeholders, including requirements engineers, during the analysis of NL requirements. Posing a question and getting an instant answer is beneficial in various quality-assurance scenarios, e.g., incompleteness detection. Answering requirements-related questions automatically is challenging since the scope of the search for answers can go beyond the given requirements specification. To that end, QAssist provides support for mining external domain-knowledge resources. Our work is one of the first initiatives to bring together QA and external domain knowledge for addressing requirements engineering challenges. We evaluate QAssist on a dataset covering three application domains and containing a total of 387 question-answer pairs. We experiment with state-of-the-art QA methods, based primarily on recent large-scale language models. In our empirical study, QAssist localizes the answer to a question to three passages within the requirements specification and within the external domain-knowledge resource with an average recall of 90.1% and 96.5%, respectively. QAssist extracts the actual answer to the posed question with an average accuracy of 84.2%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 79923fbe-082c-4ace-9ed4-224e460657c0Cited by top-tier papers1
Ask how each one uses itBuilds on8
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao et al.NeurIPS 2020 · 2,727 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- Improving Question Answering Model Robustness with Synthetic Adversarial Data GenerationMax Bartolo, Tristan Thrush, Robin Jia, Sebastian Riedel et al.EMNLP 2021 · 68 citations
Related papers
- Using Domain-specific Corpora for Improved Handling of Ambiguity in RequirementsSaad Ezzini, Sallam Abualhaija, Chetan Arora, Mehrdad Sabetzadeh et al.ICSE 2021 · 51 citations
- Identifying and Answering Questions with False Assumptions: An Interpretable ApproachZijie Wang, Eduardo BlancoEMNLP 2025
- Automated Handling of Anaphoric Ambiguity in Requirements: A Multi-solution StudySaad Ezzini, Sallam Abualhaija, Chetan Arora, Mehrdad SabetzadehICSE 2022 · 52 citations
- Where am I? Large Language Models Wandering between Semantics and Structures in Long ContextsSeonmin Koo, Jinsung Kim, YoungJoon Jang, Chanjun Park et al.EMNLP 2024
- On the Generation of Medical Question-Answer PairsSheng Shen, Yaliang Li, Nan Du, Xian Wu et al.AAAI 2020 · 25 citations
