From 2D Grids to 1D Tokens: Reforming Shared Representations for Multimodal Image Fusion
Yuchen Xian, Yunqiu Xu, Yang He, Yi Yang
Abstract
Multimodal image fusion aims to integrate complementary information from different modalities into a fused image that preserves rich local details while maintaining globally consistent appearance. Existing approaches build shared representations on 2D feature grids, which excel at modeling local structures but offer limited leverage over image-level global appearance factors. To balance these objectives, we introduce a compact 1D token interface based on a frozen pretrained image tokenizer for modeling non-local appearance/base factors. Rather than using the tokenizer as a reconstruction backbone, our design uses the 1D token space as a global carrier while retaining the 2D spatial pathway for local structure restoration. Specifically, we introduce Selective Token Editing (STE), which sparsely updates/replaces a small set of critical tokens, providing a lightweight mechanism to steer global appearance coherence while keeping the fusion backbone unchanged and avoiding extra losses. Experiments on four commonly used benchmarks show that our method achieves the best overall performance, with consistent, multi-metric improvements in both global coherence and local fidelity. Project page: https://zju-xyc.github.io/1D-Fusion-Project-Page/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3c4db464-5808-4ed8-91cd-1998000a9306Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Target-aware Dual Adversarial Learning and a Multi-scenario Multi-Modality Benchmark to Fuse Infrared and Visible for Object DetectionJinyuan Liu, Xin Fan, Zhanbo Huang, Guanyao Wu et al.CVPR 2022 · 929 citations
- FusionDN: A Unified Densely Connected Network for Image FusionHan Xu, Jiayi Ma, Zhuliang Le, Junjun Jiang et al.AAAI 2020 · 559 citations
Related papers
- Prompt Yourself: Awakening Textual Semantics in 1D Visual TokenizersHualiang Wang, Siming Fu, Weinan Jia, Yuning Lu et al.CVPR 2026
- Highly Compressed Tokenizer Can Generate Without TrainingLukas Lao Beyer, Tianhong Li, Xinlei Chen, Sertac Karaman et al.ICML 2025
- Vision Foundation Models as Effective Visual Tokenizers for Autoregressive GenerationAnlin Zheng, Xin Wen, Xuanyang Zhang, Chuofan Ma et al.NeurIPS 2025 · 21 citations
- Multimodal Token Fusion for Vision TransformersYikai Wang, Xinghao Chen, Lele Cao, Wenbing Huang et al.CVPR 2022 · 214 citations
- End-to-End Autoregressive Image Generation with 1D Semantic TokenizerWenda Chu, Bingliang Zhang, Jiaqi Han, Yizhuo Li et al.ICML 2026 · 2 citations
