Towards Semantic Equivalence of Tokenization in Multimodal LLM
Shengqiong Wu, Hao Fei, Xiangtai Li, Jiayi Ji, Hanwang Zhang, Tat-Seng Chua, Shuicheng Yan
Abstract
Multimodal Large Language Models (MLLMs) have demonstrated exceptional capabilities in processing vision-language tasks. One of the crux of MLLMs lies in vision tokenization, which involves efficiently transforming input visual signals into feature representations that are most beneficial for LLMs. However, existing vision tokenizers, essential for semantic alignment between vision and language, remain problematic. Existing methods aggressively fragment visual input, corrupting the visual semantic integrity. To address this, this work presents a novel dynamic Semantic-Equivalent Vision Tokenizer (SeTok), which groups visual features into semantic units via a dynamic clustering algorithm, flexibly determining the number of tokens based on image complexity. The resulting vision tokens effectively preserve semantic integrity and capture both low-frequency and high-frequency visual features. The proposed MLLM (SETOKIM) equipped with SeTok significantly demonstrates superior performance across various tasks, as evidenced by our experimental results. The project page is https://sqwu. top/SeTok-web/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers31
- OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and UnderstandingTao Zhang, Xiangtai Li, Hao Fei, Haobo Yuan et al.NeurIPS 2024 · 186 citations
- Vitron: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, EditingHao Fei, Shengqiong Wu, Hanwang Zhang, Tat-Seng Chua et al.NeurIPS 2024 · 100 citations
- What Factors Affect Multi-Modal In-Context Learning? An In-Depth ExplorationLibo Qin, Qiguang Chen, Hao Fei, Zhi Chen et al.NeurIPS 2024 · 37 citations
- AToken: A Unified Tokenizer for VisionJiasen Lu, Liangchen Song, Mingze Xu, Byeongjoo Ahn et al.CVPR 2026 · 33 citations
- Towards Unified Multimodal Editing with Enhanced Knowledge CollaborationKaihang Pan, Zhaoyu Fan, Juncheng Li, Qifan Yu et al.NeurIPS 2024 · 27 citations
Builds on56
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
Related papers
- A More Word-like Image Tokenization for MLLMsHyun Lee, Hyemin Jeong, Yejin Kim, Hyungwook Choi et al.CVPR 2026 · 2 citations
- Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned RepresentationsJiaming Han, Hao Chen, Yang Zhao, Hanyu Wang et al.NeurIPS 2025 · 50 citations
- Blink: Dynamic Visual Token Resolution for Enhanced Multimodal UnderstandingYuchen Feng, Zhenyu Zhang, Naibin Gu, Yilong Chen et al.CVPR 2026 · 2 citations
- Unified Language-Vision Pretraining in LLM with Dynamic Discrete Visual TokenizationYang Jin, Kun Xu, Liwei Chen, Chao Liao et al.ICLR 2024 · 87 citations
- VisionTrim: Unified Vision Token Compression for Training-Free MLLM AccelerationHanxun Yu, Wentong Li, Xuan Qu, Song Wang et al.ICLR 2026 · 17 citations
