MoC: Mixtures of Text Chunking Learners for Retrieval-Augmented Generation System
Jihao Zhao, Zhiyuan Ji, Zhaoxin Fan, Hanyu Wang, Simin Niu, Bo Tang, Feiyu Xiong, Zhiyu Li
Abstract
Retrieval-Augmented Generation (RAG), while serving as a viable complement to large language models (LLMs), often overlooks the crucial aspect of text chunking within its pipeline. This paper initially introduces a dual-metric evaluation method, comprising Boundary Clarity and Chunk Stickiness, to enable the direct quantification of chunking quality. Leveraging this assessment method, we highlight the inherent limitations of traditional and semantic chunking in handling complex contextual nuances, thereby substantiating the necessity of integrating LLMs into chunking process. To address the inherent trade-off between computational efficiency and chunking precision in LLM-based approaches, we devise the granularity-aware Mixture-of-Chunkers (MoC) framework, which consists of a three-stage processing mechanism. Notably, our objective is to guide the chunker towards generating a structured list of chunking regular expressions, which are subsequently employed to extract chunks from the original text. Extensive experiments demonstrate that both our proposed metrics and the MoC framework effectively settle challenges of the chunking task, revealing the chunking kernel while enhancing the performance of the RAG system 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- HiChunk: Evaluating and Enhancing Retrieval Augmented Generation with Hierarchical ChunkingWensheng Lu, Keyu Chen, Zhifeng Shen, Ruizhi Qiao et al.ACL 2026 · 10 citations
- QChunker: Learning Question-Aware Text Chunking for Domain RAG via Multi-Agent DebateJihao Zhao, Daixuan Li, Pengfei Li, Shuaishuai Zu et al.WWW 2026
- HiKEY: Hierarchical Multimodal Retrieval for Open-Domain Document Question AnsweringJoongmin Shin, Gyuho Shim, Jeongbae Park, Jaehyung Seo et al.ACL 2026
- Disco-RAG: Discourse-Aware Retrieval-Augmented GenerationDongqi Liu, Hang Ding, Qiming Feng, Xurong Xie et al.ACL 2026
Builds on11
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat et al.ICML 2020 · 2,937 citations
- Large Language Models Can Be Easily Distracted by Irrelevant ContextFreda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales et al.ICML 2023 · 970 citations
- End-to-End Training of Multi-Document Reader and Retriever for Open-Domain Question AnsweringDevendra Singh Sachan, Siva Reddy, William L. Hamilton, Chris Dyer et al.NeurIPS 2021 · 197 citations
- LongBench: A Bilingual, Multitask Benchmark for Long Context UnderstandingYushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu et al.ACL 2024 · 94 citations
Related papers
- A New HOPE: Domain-agnostic Automatic Evaluation of Text ChunkingHenrik Brådland, Morten Goodwin, Per-Arne Andersen, Alexander Salveson Nossum et al.SIGIR 2025 · 9 citations
- SAGE: A Framework of Precise Retrieval for RAGJintao Zhang, Guoliang Li, Jinyang SuICDE 2025 · 9 citations
- SePer: Measure Retrieval Utility Through The Lens Of Semantic Perplexity ReductionLu Dai, Yijie Xu, Jinhui Ye, Hao Liu et al.ICLR 2025
- CARROT: A Learned Cost-Constrained Retrieval Optimization System for RAGZiting Wang, Haitao Yuan, Wei Dong, Gao Cong et al.ICDE 2026 · 1 citation
- PRGB Benchmark: A Robust Placeholder-Assisted Algorithm for Benchmarking Retrieval-Augmented GenerationZhehao Tan, Yihan Jiao, Dan Yang, Junwei Liu et al.AAAI 2026
