Lune

CVPR2026Top-tier venue

Prompt Yourself: Awakening Textual Semantics in 1D Visual Tokenizers

Hualiang Wang, Siming Fu, Weinan Jia, Yuning Lu, Mu Liu, Jidong Jiang, Xiaomeng Li

2026Year

Abstract

One-dimensional (1D) visual tokenizers offer notable semantic compactness by discarding local spatial priors, and have become increasingly popular for image reconstruction and generation tasks. However, such global and sequential representations struggle to preserve fine-grained visual content; simply increasing network size or token count offers only superficial mitigation. To address this, we introduce VLTok, a novel 1D hybrid tokenizer that unifies Visual and Language representations in a shared Token space through a self-prompted training paradigm. During training, VLTok simultaneously generates 1D visual and textual tokens from images, aligning the textual tokens with embeddings from a pre-trained language model. This crossmodal alignment infuses implicit linguistic cues into the tokenizer, enhancing fine-grained image encoding. During inference, the self-prompted paradigm eliminates the need for external text, maintaining the simplicity of the image-only framework while benefiting from multi-modal guidance. Extensive experiments on the ImageNet benchmark demonstrate that VLTok achieves state-of-the-art performance in both image reconstruction and image generation. For example, under the same model parameter budget, our method achieves a relative reduction of 11.1% in rFID and 18.7% in gFID compared to GigaTok.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 522cb502-ae40-4648-9c8e-a0d90be3e87c

Builds on32

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines