Lune

CVPR2026顶会

Prompt Yourself: Awakening Textual Semantics in 1D Visual Tokenizers

Hualiang Wang, Siming Fu, Weinan Jia, Yuning Lu, Mu Liu, Jidong Jiang, Xiaomeng Li

出版方
2026年份

摘要

One-dimensional (1D) visual tokenizers offer notable semantic compactness by discarding local spatial priors, and have become increasingly popular for image reconstruction and generation tasks. However, such global and sequential representations struggle to preserve fine-grained visual content; simply increasing network size or token count offers only superficial mitigation. To address this, we introduce VLTok, a novel 1D hybrid tokenizer that unifies Visual and Language representations in a shared Token space through a self-prompted training paradigm. During training, VLTok simultaneously generates 1D visual and textual tokens from images, aligning the textual tokens with embeddings from a pre-trained language model. This crossmodal alignment infuses implicit linguistic cues into the tokenizer, enhancing fine-grained image encoding. During inference, the self-prompted paradigm eliminates the need for external text, maintaining the simplicity of the image-only framework while benefiting from multi-modal guidance. Extensive experiments on the ImageNet benchmark demonstrate that VLTok achieves state-of-the-art performance in both image reconstruction and image generation. For example, under the same model parameter budget, our method achieves a relative reduction of 11.1% in rFID and 18.7% in gFID compared to GigaTok.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper32

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖