ICML2025
FlexTok: Resampling Images into 1D Token Sequences of Flexible Length
Roman Bachmann, Jesse Allardice, David Mizrahi, Enrico Fini, Oguzhan Fatih Kar, Elmira Amirloo, Alaaeldin El-Nouby, Amir Zamir, Afshin Dehghan
摘要
FlexTok: Resampling Images into 1D Token Sequences of Flexible Length 256 tokens 128 tokens 64 tokens 32 tokens 16 tokens 8 tokens 4 tokens 2 tokens 1 token Original RGB Figure 2: Reconstruction examples using FlexTok d18-d28 trained on DFN. Notice how most of the images' semantic and geometric content is captured by fewer than 16 tokens. The first tokens already capture the high-level semantic concepts (e.g., gray bird, people in colorful garments, mountain scene, yellow flower), while more tokens are required to reconstruct more intricate scene details (e.g., position and clothing of every person, brushstroke placement, etc.). To showcase out-of-distribution reconstruction, we generated the original images using Midjourney v6.1 (Midjourney, 2024). * Note. For evaluation of the class-conditioned image generation results, we follow the common practice of measuring the generation FID (gFID) of 50K generated samples relative to the reference statistics calculated over the entire training split of the ImageNet-1k dataset (Dhariwal & Nichol, 2021).
