QChunker: Learning Question-Aware Text Chunking for Domain RAG via Multi-Agent Debate
Jihao Zhao, Daixuan Li, Pengfei Li, Shuaishuai Zu, Biao Qin, Hongyan Liu
Abstract
The effectiveness upper bound of retrieval-augmented generation (RAG) is fundamentally constrained by the semantic integrity and information granularity of text chunks in its knowledge base. Moreover, domain documents are characterized by dense terminology and strong contextual dependencies, which exacerbate the semantic fragmentation of text chunks, thereby making it difficult to efficiently utilize their key information. To address these challenges, this paper proposes QChunker, which restructures the RAG paradigm from retrieval-augmentation to understanding-retrieval-augmentation. Firstly, QChunker models the text chunking as a composite task of text segmentation and knowledge completion to ensure the logical coherence and integrity of text chunks. Drawing inspiration from Hal Gregersen's ''Questions Are the Answer'' theory, we design a multi-agent debate framework comprising four specialized components: a question outline generator, text segmenter, integrity reviewer, and knowledge completer. This framework operates on the principle that questions serve as catalysts for profound insights. Through this pipeline, we successfully construct a high-quality dataset of 45K entries and transfer this capability to small language models. Additionally, to handle long evaluation chains and low efficiency in existing chunking evaluation methods, which overly rely on downstream QA tasks, we introduce a novel direct evaluation metric, ChunkScore. Both theoretical and experimental validations demonstrate that ChunkScore can directly and efficiently discriminate the quality of text chunks. Furthermore, during the text segmentation phase, we utilize document outlines for multi-path sampling to generate multiple candidate chunks and select the optimal solution employing ChunkScore. Extensive experimental results across four heterogeneous domains exhibit that QChunker effectively resolves aforementioned issues by providing RAG with more logically coherent and information-rich text chunks. Notably, this study also establishes a small-domain QA dataset concerning hazardous chemical safety, which fully reveals the significant value of RAG in specialized domains and the generalization capability of the QChunker framework.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2876c98b-e077-4cf9-8ef4-4cf48cb15818Builds on13
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil et al.ICLR 2024 · 1,798 citations
- Large Language Models Can Be Easily Distracted by Irrelevant ContextFreda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales et al.ICML 2023 · 970 citations
- End-to-End Training of Multi-Document Reader and Retriever for Open-Domain Question AnsweringDevendra Singh Sachan, Siva Reddy, William L. Hamilton, Chris Dyer et al.NeurIPS 2021 · 197 citations
- LongBench: A Bilingual, Multitask Benchmark for Long Context UnderstandingYushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu et al.ACL 2024 · 94 citations
Related papers
- A New HOPE: Domain-agnostic Automatic Evaluation of Text ChunkingHenrik Brådland, Morten Goodwin, Per-Arne Andersen, Alexander Salveson Nossum et al.SIGIR 2025 · 9 citations
- SmartChunk Retrieval: Query-Aware Chunk Compression with Planning for Efficient Document RAGXuechen Zhang, Koustava Goswami, Samet Oymak, Jiasi Chen et al.ICLR 2026 · 1 citation
- HiChunk: Evaluating and Enhancing Retrieval Augmented Generation with Hierarchical ChunkingWensheng Lu, Keyu Chen, Zhifeng Shen, Ruizhi Qiao et al.ACL 2026 · 10 citations
- MCoRe: Multi-Entry Complementary Retrieval with Reflection-Guided Iteration for Multi-Hop QAJuxiang Zeng, Zhuohui Gao, Zhe Hou, Pinghui Wang et al.KDD 2026
- Clue-RAG: Towards Accurate and Cost-Efficient Graph-Based RAG Via Multi-Partite Graph-Based IndexYaodong Su, Yixiang Fang, Yingli Zhou, Chuanhui YangICDE 2026
