Highly Compressed Tokenizer Can Generate Without Training
Lukas Lao Beyer, Tianhong Li, Xinlei Chen, Sertac Karaman, Kaiming He
摘要
Commonly used image tokenizers produce a 2D grid of spatially arranged tokens. In contrast, socalled 1D image tokenizers represent images as highly compressed one-dimensional sequences of as few as 32 discrete tokens. We find that the high degree of compression achieved by a 1D tokenizer with vector quantization enables image editing and generative capabilities through heuristic manipulation of tokens, demonstrating that even very crude manipulations -such as copying and replacing tokens between latent representations of images -enable fine-grained image editing by transferring appearance and semantic attributes. Motivated by the expressivity of the 1D tokenizer's latent space, we construct an image generation pipeline leveraging gradient-based testtime optimization of tokens with plug-and-play loss functions such as reconstruction or CLIP similarity. Our approach is demonstrated for inpainting and text-guided image editing use cases, and can generate diverse and realistic samples without requiring training of any generative model. Code is available at https://github.com/ lukaslaobeyer/token-opt .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Group Critical-token Policy Optimization for Autoregressive Image GenerationGuohui Zhang, Hu Yu, Xiaoxiao Ma, JingHao Zhang 等ICLR 2026 · 被引用 16 次
- EditCtrl: Disentangled Local and Global Control for Real-Time Generative Video EditingYehonathan Litman, Shikun Liu, Dario Seyb, Nicholas Milef 等CVPR 2026 · 被引用 5 次
- From 2D Grids to 1D Tokens: Reforming Shared Representations for Multimodal Image FusionYuchen Xian, Yunqiu Xu, Yang He, Yi YangICML 2026 · 被引用 2 次
- Evaluating Generative Models via One-Dimensional Code DistributionsZexi Jia, Pengcheng Luo, Yijia Zhong, Jinchao Zhang 等CVPR 2026 · 被引用 2 次
- ImgCoT: Compressing Long Chain of Thought into Compact Visual Tokens for Efficient Reasoning of Large Language ModelXiaoshu Chen, sihang zhou, KE LIANG, Taichun Zhou 等ICML 2026 · 被引用 1 次
它引用的顶会 Paper22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
相关 Paper
- An Image is Worth 32 Tokens for Reconstruction and GenerationQihang Yu, Mark Weber, Xueqing Deng, Xiaohui Shen 等NeurIPS 2024 · 被引用 331 次
- End-to-End Autoregressive Image Generation with 1D Semantic TokenizerWenda Chu, Bingliang Zhang, Jiaqi Han, Yizhuo Li 等ICML 2026 · 被引用 2 次
- Spectral Image TokenizerCarlos Esteves, Mohammed Suhail, Ameesh MakadiaICCV 2025 · 被引用 1 次
- Prompt Yourself: Awakening Textual Semantics in 1D Visual TokenizersHualiang Wang, Siming Fu, Weinan Jia, Yuning Lu 等CVPR 2026
- Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional TokensDongwon Kim, Ju He, Qihang Yu, Chenglin Yang 等ICCV 2025 · 被引用 8 次
