COEUR: COhesion and Exhaustiveness of User Stories Representations
Marius Ortega, Hassan Imhah, Nédra Mellouli, Christophe Rodrigues, Nicolas Travers
Abstract
User Stories are key artifacts in Requirement and Software Engineering. Despite their wide adoption, their writing in industrial contexts tends to diverge from the principles initially stated in Agile methodologies. In this context, sets of metrics such as INVEST or QUS emerged to qualify these items. In this paper, we argue that the said sets of metrics are only partially efficient at capturing the quality of user stories contextualized in a project, and that their actual adoption in business contexts is limited due to multiple aspects: their unfitness to specific contexts, the difficulty of implementation requiring human intervention, the absence of reproducibility or their misalignment with actual quality of user stories. Such limitations, prevent practitioners from efficiently applying them for downstream tasks such as LLM-based user story generation. To address these challenges, we introduce COEUR, a framework comprising two metrics: Cohesion and Exhaustiveness. These metrics are designed to mirror core Product Owner responsibilities. Specifically, Cohesion evaluates the structural organization and logical grouping of the backlog, while Exhaustiveness monitors the semantic alignment between the elicited needs and the proposed technical solutions. Additionally, they are quantitative, reproducible, and automatically computable. To evaluate COEUR, we conduct four empirical experiments organized into two validation tracks. The first track utilizes two noising-based experiments to assess our metrics' sensitivity to requirement degradation. The second track evaluates their performance in generative contexts through two standard learning paradigms for LLMs: In-Context Learning (ICL) and Supervised Fine-Tuning (SFT), both applied to our user story generation task. Subsequently, COEUR provides a turnkey measurement of Product Backlogs' quality for both project monitoring by human experts and LLM benchmarking applied to automatic user stories generation.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 20bb082b-d3a1-4bd1-90a5-61e4b9d2d8c6Related papers
- CoQuIR: A Comprehensive Benchmark for Code Quality-Aware Information RetrievalJiahui Geng, Fengyu Cai, Shaobo Cui, Qing Li et al.ACL 2026 · 3 citations
- HoarePrompt: Structural Reasoning About Program Correctness in Natural LanguageDimitrios Stamatios Bouras, Yihan Dai, Tairan Wang, Yingfei Xiong et al.ICSE 2026
- OpenMEVA: A Benchmark for Evaluating Open-ended Story Generation MetricsJian Guan, Zhexin Zhang, Zhuoer Feng, Zitao Liu et al.ACL 2021
- From Bugs to Benefits: Improving User Stories by Leveraging Crowd Knowledge with CrUISE-ACStefan Schwedt, Thomas StröderICSE 2025 · 3 citations
- SWR-Bench: Assessing LLM Performance in Real-World Code Review Comment GenerationZhengran Zeng, Ruikai Shi, Keke Han, Yixin Li et al.FSE 2026
