A Non-Factoid Question-Answering Taxonomy
Valeria Bolotova, Vladislav Blinov, Falk Scholer, W. Bruce Croft, Mark Sanderson
Abstract
Non-factoid question answering (NFQA) is a challenging and underresearched task that requires constructing long-form answers, such as explanations or opinions, to open-ended non-factoid questions -NFQs. There is still little understanding of the categories of NFQs that people tend to ask, what form of answers they expect to see in return, and what the key research challenges of each category are.
This work presents the first comprehensive taxonomy of NFQ categories and the expected structure of answers. The taxonomy was constructed with a transparent methodology and extensively evaluated via crowdsourcing. The most challenging categories were identified through an editorial user study. We also release a dataset of categorised NFQs and a question category classifier 1 .
Finally, we conduct a quantitative analysis of the distribution of question categories using major NFQA datasets, showing that the NFQ categories that are the most challenging for current NFQA systems are poorly represented in these datasets. This imbalance may lead to insufficient system performance for challenging categories. The new taxonomy, along with the category classifier, will aid research in the area, helping to create more balanced benchmarks and to focus models on addressing specific categories.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 24b98cd7-ca4b-4964-aa52-68e81f9cfae9Cited by top-tier papers7
- MemoryBench: A Benchmark for Memory and Continual Learning in LLM SystemsQingyao Ai, Yichen Tang, Changyue Wang, Jianming Long et al.ICML 2026 · 47 citations
- WikiHowQA: A Comprehensive Benchmark for Multi-Document Non-Factoid Question AnsweringValeria Bolotova-Baranova, Vladislav Blinov, Sofya Filippova, Falk Scholer et al.ACL 2023 · 12 citations
- A User-Centric Multi-Intent Benchmark for Evaluating Large Language ModelsJiayin Wang, Fengran Mo, Weizhi Ma, Peijie Sun et al.EMNLP 2024 · 10 citations
- Gesture and Audio-Haptic Guidance Techniques to Direct Conversations with Intelligent Voice InterfacesShwetha Rajaram, Hemant Bhaskar Surale, Codie McConkey, Carine Rognon et al.CHI 2025 · 5 citations
- Effective Contrastive Weighting for Dense Query ExpansionXiao Wang, Sean MacAvaney, Craig Macdonald, Iadh OunisACL 2023 · 2 citations
Builds on2
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- SubjQA: A Dataset for Subjectivity and Review ComprehensionJohannes Bjerva, Nikita Bhutani, Behzad Golshan, Wang-Chiew Tan et al.EMNLP 2020
Related papers
- ASQA: Factoid Questions Meet Long-Form AnswersIvan Stelmakh, Yi Luan, Bhuwan Dhingra, Ming-Wei ChangEMNLP 2022 · 51 citations
- An Empirical Study of Evaluating Long-form Question AnsweringNing Xian, Yixing Fan, Ruqing Zhang, Maarten de Rijke et al.SIGIR 2025 · 2 citations
- A Critical Evaluation of Evaluations for Long-form Question AnsweringFangyuan Xu, Yixiao Song, Mohit Iyyer, Eunsol ChoiACL 2023 · 25 citations
- A Taxonomy of Empathetic Questions in Social DialogsEkaterina Svikhnushina, Iuliana Voinea, Anuradha Welivita, Pearl PuACL 2022 · 17 citations
- How Do We Answer Complex Questions: Discourse Structure of Long-form AnswersFangyuan Xu, Junyi Jessy Li, Eunsol ChoiACL 2022 · 25 citations
