Glyph: Scaling Context Windows via Visual-Text Compression
Jiale Cheng, Yusen Liu, Xinyu Zhang, Yulin Fei, Wenyi Hong, Ruiliang Lyu, Weihan Wang, Zhe Su, Xiaotao Gu, Xiao Liu, Yushi Bai, Jie Tang
Abstract
Large language models (LLMs) increasingly rely on long-context modeling for tasks such as document understanding, code analysis, and multi-step reasoning. However, scaling context windows to the million-token level brings prohibitive computational and memory costs, limiting the practicality of long-context LLMs. In this work, we take a different perspective-visual context scaling-to tackle this challenge. Instead of extending token-based sequences, we propose Glyph, a framework that renders long texts into images and processes them with vision-language models (VLMs). This approach substantially compresses textual input while preserving semantic information, and we further design an LLM-driven genetic search to identify optimal visual rendering configurations for balancing accuracy and compression. Through extensive experiments, we demonstrate that our method achieves 3-4× token compression while maintaining accuracy comparable to leading LLMs such as Qwen3-8B on various long-context benchmarks. This compression also leads to around 4× faster prefilling and decoding, and approximately 2× faster SFT training. Furthermore, under extreme compression, a 128K-context VLM could scale to handle 1M-token-level text tasks. In addition, the rendered text data benefits realworld multimodal tasks, such as document understanding. Our code and model are released at https://github.com/thu-coai/Glyph .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 41e2af92-ee25-46ac-878d-051364e91786Cited by top-tier papers1
Ask how each one uses itBuilds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 1,168 citations
- YaRN: Efficient Context Window Extension of Large Language ModelsBowen Peng, Jeffrey Quesnelle, Honglu Fan, Enrico ShippoleICLR 2024 · 508 citations
- Gated Linear Attention Transformers with Hardware-Efficient TrainingSonglin Yang, Bailin Wang, Yikang Shen, Rameswar Panda et al.ICML 2024 · 390 citations
Related papers
- Inference Optimal VLMs Need Fewer Visual Tokens and More ParametersKevin Y. Li, Sachin Goyal, João D. Semedo, J. Zico KolterICLR 2025
- Vision-centric Token Compression in Large Language ModelLing Xing, Alex Jinpeng Wang, Rui Yan, Xiangbo Shu et al.NeurIPS 2025 · 32 citations
- Leveraging Visual Tokens for Extended Text Contexts in Multi-Modal LearningAlex Jinpeng Wang, Linjie Li, Yiqi Lin, Min Li et al.NeurIPS 2024 · 21 citations
- EvoComp: Learning Visual Token Compression for Multimodal Large Language Models via Semantic-Guided Evolutionary LabelingJiafei Song, Fengwei Zhou, Jin Qu, Wenjin Jason Li et al.CVPR 2026 · 4 citations
- Vision-aligned Latent Reasoning for Multi-modal Large Language ModelByungwoo Jeon, Yoonwoo Jeong, Hyunseok Lee, Minsu Cho et al.ICML 2026 · 7 citations
