UNIALIGN: Scaling Multimodal Alignment within One Unified Model
Bo Zhou, Liulei Li, Yujia Wang, Huafeng Liu, Yazhou Yao, Wenguan Wang
摘要
We present UNIALIGN, a unified model to align an arbitrary number of modalities (e.g., image, text, audio, 3D point cloud, etc.) through one encoder and a single training phase. Existing solutions typically employ distinct encoders for each modality, resulting in increased parameters as the number of modalities grows. In contrast, UNIALIGN proposes a modality-aware adaptation of the powerful mixtureof-experts (MoE) schema and further integrates it with Low-Rank Adaptation (LoRA), efficiently scaling the encoder to accommodate inputs in diverse modalities while maintaining a fixed computational overhead. Moreover, prior work often requires separate training for each extended modality. This leads to task-specific models and further hinders the communication between modalities. To address this, we propose a soft modality binding strategy that aligns all modalities using unpaired data samples across datasets. Two additional training objectives are introduced to distill knowledge from well-aligned anchor modalities and prior multimodal models, elevating UNIALIGN into a high performance multimodal foundation model. Experiments on 11 benchmarks across 6 different modalities demonstrate that UNIALIGN could achieve comparable performance to SOTA approaches, while using merely 7.8M trainable parameters and maintaining an identical model with the same weight across all tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Efficiency Follows Global-Local DecouplingZhenyu Yang, Gensheng Pei, Tao Chen, Yichao Zhou 等CVPR 2026 · 被引用 3 次
- Iris: Bringing Real-World Priors into Diffusion Model for Monocular Depth EstimationXinhao Cai, Gensheng Pei, Zeren Sun, Yazhou Yao 等CVPR 2026 · 被引用 2 次
- Unifying Heterogeneous Multi-Modal Remote Sensing Detection Via Language-Pivoted PretrainingYuxuan Li, Yuming Chen, Yunheng Li, Ming-Ming Cheng 等ICML 2026
- Beyond Quadratic: Linear-Time Change Detection with RWKVZhenyu Yang, Gensheng Pei, Tao Chen, Xia Yuan 等AAAI 2026
它引用的顶会 Paper50
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
相关 Paper
- Each Rank Could be an Expert: Single-Ranked Mixture of Experts LoRA for Multi-task LearningZiyu Zhao, Yixiao Zhou, Xin Yu, Zhi Zhang 等KDD 2026 · 被引用 13 次
- Multi-Task Dense Prediction via Mixture of Low-Rank ExpertsYuqi Yang, Peng-Tao Jiang, Qibin Hou, Hao Zhang 等CVPR 2024 · 被引用 32 次
- Make LoRA Great Again: Boosting LoRA with Adaptive Singular Values and Mixture-of-Experts Optimization AlignmentChenghao Fan, Zhenyi Lu, Sichen Liu, Chengfeng Gu 等ICML 2025
- DeLo: Dual Decomposed Low-Rank Experts Collaboration for Continual Missing Modality LearningXiwei Liu, Yulong Li, Feilong Tang, Imran RazzakAAAI 2026
- UniBind: LLM-Augmented Unified and Balanced Representation Space to Bind Them AllYuanhuiyi Lyu, Xu Zheng, Jiazhou Zhou, Lin WangCVPR 2024 · 被引用 8 次
