Autoregressive Pre-Training on Pixels and Texts
Yekun Chai, Qingyi Liu, Jingwu Xiao, Shuohuan Wang, Yu Sun, Hua Wu
Abstract
The integration of visual and textual information represents a promising direction in the advancement of language models. In this paper, we explore the dual modality of language-both visual and textual-within an autoregressive framework, pre-trained on both document images and texts. Our method employs a multimodal training strategy, utilizing visual data through next patch prediction with a regression head and/or textual data through next token prediction with a classification head. We focus on understanding the interaction between these two modalities and their combined impact on model performance. Our extensive evaluation across a wide range of benchmarks shows that incorporating both visual and textual data significantly improves the performance of pixel-based language models. Remarkably, we find that a unidirectional pixelbased model trained solely on visual data can achieve comparable results to state-of-the-art bidirectional models on several language understanding tasks. This work uncovers the untapped potential of integrating visual and textual modalities for more effective language modeling. We release our code, data, and model checkpoints at https://github.com/ ernie-research/pixelgpt .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Beyond Text Compression: Evaluating Tokenizers Across ScalesJonas F. Lotz, António Vilarinho Lopes, Stephan Peitz, Hendra Setiawan et al.ACL 2025 · 3 citations
- Multilingual Pretraining for Pixel Language ModelsIlker Kesen, Jonas F. Lotz, Ingo Ziegler, Phillip Rust et al.EMNLP 2025 · 1 citation
- Understanding Subword Compositionality of Large Language ModelsQiwei Peng, Yekun Chai, Anders SøgaardEMNLP 2025
Builds on13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu et al.ICML 2020 · 1,773 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Scalable Pre-training of Large Autoregressive Image ModelsAlaaeldin El-Nouby, Michal Klein, Shuangfei Zhai, Miguel Ángel Bautista et al.ICML 2024 · 130 citations
Related papers
- Language Modelling with PixelsPhillip Rust, Jonas F. Lotz, Emanuele Bugliarello, Elizabeth Salesky et al.ICLR 2023 · 17 citations
- Grounding Language Models to Images for Multimodal Inputs and OutputsJing Yu Koh, Ruslan Salakhutdinov, Daniel FriedICML 2023 · 160 citations
- UNIMO: Towards Unified-Modal Understanding and Generation via Cross-Modal Contrastive LearningWei Li, Can Gao, Guocheng Niu, Xinyan Xiao et al.ACL 2021
- CLIPPO: Image-and-Language Understanding from Pixels OnlyMichael Tschannen, Basil Mustafa, Neil HoulsbyCVPR 2023
- HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language EmbeddingChenxin Tao, Shiqian Su, Xizhou Zhu, Chenyu Zhang et al.CVPR 2025
