From Token to Token Pair: Efficient Prompt Compression for Large Language Models in Clinical Prediction
Mingcheng Zhu, Zhiyao Luo, Yu Liu, Tingting Zhu
Abstract
By processing electronic health records (EHRs) as natural language sequences, large language models (LLMs) have shown potential in clinical prediction tasks such as mortality prediction and phenotyping. However, longitudinal or highly frequent EHRs often yield excessively long token sequences that result in high computational costs and even reduced performance. Existing solutions either add modules for compression or remove less important tokens, which introduce additional inference latency or risk losing clinical information. To achieve lossless compression of token sequences without additional cost or loss of performance, we propose Medical Token-Pair Encoding (MedTPE), a layered method that extends standard tokenisation for EHR sequences. MedTPE merges frequently co-occurring medical token pairs into composite tokens, providing lossless compression while preserving the computational complexity through a dependency-aware replacement strategy. Only the embeddings of the newly introduced tokens of merely 0.5-1.0% of the LLM’s parameters are fine-tuned via self-supervised learning. Experiments on real-world datasets for two clinical scenarios demonstrate that MedTPE reduces input token length by up to 31% and inference latency by 34-63%, while maintaining or even improving both predictive performance and output format compliance across multiple LLMs and four clinical prediction tasks. Furthermore, MedTPE demonstrates robustness across different input context lengths and generalisability to scientific and financial domains and different languages. The code is available in the GitHub repository.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on12
- Are More LLM Calls All You Need? Towards the Scaling Properties of Compound AI SystemsLingjiao Chen, Jared Quincy Davis, Boris Hanin, Peter Bailis et al.NeurIPS 2024 · 110 citations
- LLMLingua: Compressing Prompts for Accelerated Inference of Large Language ModelsHuiqiang Jiang, Qianhui Wu, Chin-Yew Lin, Yuqing Yang et al.EMNLP 2023 · 94 citations
- Prompt Compression with Context-Aware Sentence Encoding for Fast and Improved LLM InferenceBarys Liskavets, Maxim Ushakov, Shuvendu Roy, Mark Klibanov et al.AAAI 2025 · 41 citations
- Zero-Shot Tokenizer TransferBenjamin Minixhofer, Edoardo Maria Ponti, Ivan VulicNeurIPS 2024 · 37 citations
- Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement FinetuningMinheng Ni, Zhengyuan Yang, Linjie Li, Chung-Ching Lin et al.NeurIPS 2025 · 35 citations
Related papers
- zip2zip: Inference-Time Adaptive Tokenization via Online CompressionSaibo Geng, Nathan Ranchin, Yunzhen Yao, Maxime Peyrard et al.NeurIPS 2025 · 5 citations
- Multimodal Medical Code TokenizerXiaorui Su, Shvat Messica, Yepeng Huang, Ruth Johnson et al.ICML 2025 · 2 citations
- Byte Pair Encoding for Symbolic MusicNathan Fradet, Nicolas Gutowski, Fabien Chhel, Jean-Pierre BriotEMNLP 2023 · 8 citations
- Retrofitting Large Language Models with Dynamic TokenizationDarius Feher, Ivan Vulic, Benjamin MinixhoferACL 2025
- HeartLLM: Discretized ECG Tokenization for LLM-Based Diagnostic ReasoningJinning Yang, Wenjie Sun, Wen ShiAAAI 2026
