Directed-Tokens: A Robust Multi-Modality Alignment Approach to Large Language-Vision Models
Thanh-Dat Truong, Huu-Thien Tran, Tran Thai Son, Bhiksha Raj, Khoa Luu
Abstract
Large multimodal models (LMMs) have gained impressive performance due to their outstanding capability in various understanding tasks. However, these models still suffer from some fundamental limitations related to robustness and generalization due to the alignment and correlation between visual and textual features. In this paper, we introduce a simple but efficient learning mechanism for improving the robust alignment between visual and textual modalities by solving shuffling problems. In particular, the proposed approach can improve reasoning capability, visual understanding, and cross-modality alignment by introducing two new tasks: reconstructing the image order and the text order into the LMM's pre-training and fine-tuning phases. In addition, we propose a new directed-token approach to capture visual and textual knowledge, enabling the capability to reconstruct the correct order of visual inputs. Then, we introduce a new Image-to-Response Guided loss to further improve the visual understanding of the LMM in its responses. The proposed approach consistently achieves state-of-the-art (SoTA) performance compared with prior LMMs on academic task-oriented and instruction-following LMM benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on23
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
Related papers
- Multi-modal Auto-regressive Modeling via Visual TokensTianshuo Peng, Zuchao Li, Lefei Zhang, Hai Zhao et al.ACM MM 2024 · 1 citation
- SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-rewardsJixiang Hong, Yiran Zhang, Guanzhong Wang, Yi Liu et al.KDD 2026 · 4 citations
- SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMsYuanyang Yin, Yaqi Zhao, Yajie Zhang, Yuanxing Zhang et al.EMNLP 2025
- Unified Multimodal Understanding via Byte-Pair Visual EncodingWanpeng Zhang, Yicheng Feng, Hao Luo, Yijiang Li et al.ICCV 2025 · 2 citations
- HiMTok: Learning Hierarchical Mask Tokens for Image Segmentation with Large Multimodal ModelTao Wang, Changxu Cheng, Lingfeng Wang, Senda Chen et al.ICCV 2025 · 5 citations
