Subobject-level Image Tokenization
Delong Chen, Samuel Cahyawijaya, Jianfeng Liu, Baoyuan Wang, Pascale Fung
摘要
Patch-based image tokenization ignores the morphology of the visual world, limiting effective and efficient learning of image understanding. Inspired by subword tokenization, we introduce subobject-level adaptive token segmentation and explore several approaches, including superpixel, SAM, and a proposed Efficient and PanOptiC (EPOC) image tokenizer. Our EPOC combines boundary detection-a simple task that can be handled well by a compact model-with watershed segmentation, which inherently guarantees no pixels are left unsegmented. Intrinsic evaluations across 5 datasets demonstrate that EPOC's segmentation aligns well with human annotations of both object-and part-level visual morphology, producing more monosemantic tokens and offering substantial efficiency advantages. For extrinsic evaluation, we designed a token embedding that handles arbitrary-shaped tokens, and trained VLMs with different tokenizers on 4 datasets of object recognition and detailed captioning. The results reveal that subobject tokenization enables faster convergence and better generalization while using fewer visual tokens. Project website: https://github. com/ChenDelong1999/subobjects .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Understanding the Limits of Vision Language Models Through the Lens of the Binding ProblemDeclan Campbell, Sunayana Rane, Tyler Giallanza, Nicolò De Sabbata 等NeurIPS 2024 · 被引用 101 次
- VL-JEPA: Joint Embedding Predictive Architecture for Vision-languageDelong Chen, Mustafa Shukor, Théo Moutakanni, Willy Chung 等ICLR 2026 · 被引用 60 次
- RemoteSAM: Towards Segment Anything for Earth ObservationLiang Yao, Fan Liu, Delong Chen, Chuanyi Zhang 等ACM MM 2025 · 被引用 28 次
- Model Decides How to Tokenize: Adaptive DNA Sequence Tokenization with MxDNALifeng Qiao, Peng Ye, Yuchen Ren, Weiqiang Bai 等NeurIPS 2024 · 被引用 23 次
- RemoteReasoner: Towards Unifying Geospatial Reasoning WorkflowLiang Yao, Fan Liu, Hongbo Lu, Chuanyi Zhang 等AAAI 2026 · 被引用 16 次
它引用的顶会 Paper28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
相关 Paper
- SAMTok: Representing Any Mask with Two WordsYikang Zhou, Tao Zhang, Dengxian Gong, Yuanzheng Wu 等CVPR 2026 · 被引用 10 次
- HiMTok: Learning Hierarchical Mask Tokens for Image Segmentation with Large Multimodal ModelTao Wang, Changxu Cheng, Lingfeng Wang, Senda Chen 等ICCV 2025 · 被引用 5 次
- Learning Hierarchical Image Segmentation For Recognition and By RecognitionTsung-Wei Ke, Sangwoo Mo, Stella X. YuICLR 2024 · 被引用 20 次
- SemHiTok: A Unified Image Tokenizer via Semantic-Guided Hierarchical Codebook for Multimodal Understanding and GenerationZisheng Chen, Chunwei Wang, Runhui Huang, Hongbin Xu 等ICLR 2026 · 被引用 24 次
- UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and GenerationZhengrong Yue, Haiyu Zhang, Xiangyu Zeng, Boyu Chen 等ICLR 2026 · 被引用 25 次
