Dynamic Chunking for End-to-End Hierarchical Sequence Modeling
Sukjun Hwang, Brandon Wang, Albert Gu
摘要
Major progress on language models (LMs) in recent years has largely resulted from moving away from specialized models designed for specific tasks, to general models based on powerful architectures (e.g. the Transformer) that learn everything from raw data. Despite this trend, pre-processing steps such as tokenization remain a barrier to true end-to-end foundation models. We introduce a collection of new techniques that enable a dynamic chunking mechanism which automatically learns content-and context-dependent segmentation strategies learned jointly with the rest of the model. Incorporating this into an explicit hierarchical network (H-Net) allows replacing the (implicitly hierarchical) tokenization-LMdetokenization pipeline with a single model learned fully end-to-end. When compute-and data-matched, an H-Net with one stage of hierarchy operating at the byte level outperforms a strong Transformer language model operating over BPE tokens. Iterating the hierarchy to multiple stages further increases its performance by modeling multiple levels of abstraction, demonstrating significantly better scaling with data and matching the token-based Transformer of twice its size. H-Nets pretrained on English show significantly increased character-level robustness, and qualitatively learn meaningful data-dependent chunking strategies without any heuristics or explicit supervision. Finally, the H-Net's improvement over tokenized pipelines is further increased in languages and modalities with weaker tokenization heuristics, such as Chinese and code, or DNA sequences (nearly 4× improvement in data efficiency over baselines), showing the potential of true end-to-end models that learn and scale better from unprocessed data. 1 Many other edge cases have been discussed in informal online discourse rather than papers; we defer to Andrej Karpathy's lectures and tweets. 2 An extended related work can be found in Appendix A, which is summarized in Table 6 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- From Bytes to Ideas: Language Modeling with Autoregressive U-NetsMathurin Videau, Badr Youbi Idrissi, Alessandro Ferreira Leite, Marc Schoenauer 等NeurIPS 2025 · 被引用 14 次
- Context-level Language Modeling by Learning Predictive Context Embeddingsbeiya dai, Yuliang Liu, Yunchong Song, Daozheng Xue 等ICML 2026 · 被引用 5 次
- zip2zip: Inference-Time Adaptive Tokenization via Online CompressionSaibo Geng, Nathan Ranchin, Yunzhen Yao, Maxime Peyrard 等NeurIPS 2025 · 被引用 5 次
- Proxy Compression for Language ModelingLin Zheng, Li Xinyu, Qian Liu, Xiachong Feng 等ICML 2026 · 被引用 3 次
- MergeDNA: Context-Aware Genome Modeling with Dynamic Tokenization Through Token MergingSiyuan Li, Kai Yu, Anna Wang, Zicheng Liu 等AAAI 2026 · 被引用 2 次
它引用的顶会 Paper54
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
相关 Paper
- ByteFlow: Language Modeling through Adaptive Byte Compression without a TokenizerChunyuan Deng, Sanket Lokegaonkar, Colin Lockard, Besnik Fetahu 等ICLR 2026 · 被引用 1 次
- LDARNet: DNA Adaptive Representation Network with Learnable Tokenization for Genomic ModelingDaria Ledneva, Denis KuznetsovICML 2026 · 被引用 1 次
- Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language ModelsPit Neitemeier, Björn Deiseroth, Constantin Eichenberg, Lukas BallesICLR 2025
- From Characters to Words: Hierarchical Pre-trained Language Model for Open-vocabulary Language UnderstandingLi Sun, Florian Luisier, Kayhan Batmanghelich, Dinei A. F. Florêncio 等ACL 2023
- dnaHNet: A Scalable and Hierarchical Foundation Model for Genomic Sequence LearningArnav Shah, Junzhe Li, Parsa Idehpour, Adibvafa Fallahpour 等ICML 2026
