A New HOPE: Domain-agnostic Automatic Evaluation of Text Chunking
Henrik Brådland, Morten Goodwin, Per-Arne Andersen, Alexander Salveson Nossum, Aditya Gupta
摘要
Document chunking fundamentally impacts Retrieval-Augmented Generation (RAG) by determining how source materials are segmented before indexing. Despite evidence that Large Language Models (LLMs) are sensitive to the layout and structure of retrieved data, there is currently no framework to analyze the impact of different chunking methods. In this paper, we introduce a novel methodology that defines essential characteristics of the chunking process at three levels: intrinsic passage properties, extrinsic passage properties, and passages-document coherence. We propose HOPE (Holistic Passage Evaluation), a domain-agnostic, automatic evaluation metric that quantifies and aggregates these characteristics. Our empirical evaluations across seven domains demonstrate that the HOPE metric correlates significantly (𝜌 > 0.13) with various RAG performance indicators, revealing contrasts between the importance of extrinsic and intrinsic properties of passages. Semantic independence between passages proves essential for system performance with a performance gain of up to 56.2% in factual correctness and 21.1% in answer correctness. On the contrary, traditional assumptions about maintaining concept unity within passages show minimal impact. These findings provide actionable insights for optimizing chunking strategies, thus improving RAG system design to produce more factually correct responses.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- The Power of Noise: Redefining Retrieval for RAG SystemsFlorin Cuconasu, Giovanni Trappolini, Federico Siciliano, Simone Filice 等SIGIR 2024 · 被引用 212 次
- A Question Answering Framework for Decontextualizing User-facing Snippets from Scientific DocumentsBenjamin Newman, Luca Soldaini, Raymond Fok, Arman Cohan 等EMNLP 2023 · 被引用 6 次
- HelpSteer2-Preference: Complementing Ratings with PreferencesZhilin Wang, Alexander Bukharin, Olivier Delalleau, Daniel Egert 等ICLR 2025
相关 Paper
- SePer: Measure Retrieval Utility Through The Lens Of Semantic Perplexity ReductionLu Dai, Yijie Xu, Jinhui Ye, Hao Liu 等ICLR 2025
- RAGEval: Scenario Specific RAG Evaluation Dataset Generation FrameworkKunlun Zhu, Yifan Luo, Dingling Xu, Yukun Yan 等ACL 2025 · 被引用 53 次
- HiChunk: Evaluating and Enhancing Retrieval Augmented Generation with Hierarchical ChunkingWensheng Lu, Keyu Chen, Zhifeng Shen, Ruizhi Qiao 等ACL 2026 · 被引用 10 次
- Testing Retrieval-Augmented Generation Systems with Chunk CoverageJinhan Kim, Samuele Pasini, Paolo TonellaISSTA 2026
- MoC: Mixtures of Text Chunking Learners for Retrieval-Augmented Generation SystemJihao Zhao, Zhiyuan Ji, Zhaoxin Fan, Hanyu Wang 等ACL 2025 · 被引用 21 次
