Toward Structural Multimodal Representations: Specialization, Selection, and Sparsification via Mixture-of-Experts
Hahyeon Choi, NOJUN KWAK
摘要
We propose S3 (Specialization, Selection, Sparsification), a framework that rethinks multimodal learning through a structural perspective. Instead of encoding all signals into a fixed embedding, S3 decomposes multimodal inputs into semantic experts and selectively routes them for each task. Specialization forms concept-level experts in a shared latent space, Selection adapts routing for task-specific needs, and Sparsification prunes low-utility paths to yield compact, information-minimal representations. Across four MultiBench benchmarks, S3 improves accuracy and exhibits consistent sparsity-performance dynamics, exhibiting a reverse U-shaped trend, with performance peaking at intermediate sparsity. These results suggest that structuring multimodal representations as selectable semantic components provides a practical and principled alternative to contrastive learning or InfoMax-driven approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna 等NeurIPS 2020 · 被引用 7,049 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann 等NeurIPS 2021 · 被引用 1,213 次
相关 Paper
- Sparse CLIP: Co-Optimizing Interpretability and Performance in Contrastive LearningChuan Qin, Constantin Venhoff, Sonia Joseph, Fanyi Xiao 等ICLR 2026 · 被引用 4 次
- Retrv-MoE: Scaling Unified Multimodal Retrieval with Sparse Mixture-of-ExpertsTongxu Lin, Jiayin XiaoKDD 2026
- MM-Spectrum: Multimodal Multi-spectral Molecular Structural Elucidation with a Stable MoE FrameworkHaitao YU, Nan Min, Zheng Fang, Hongyu Zhan 等ICML 2026
- PLUM-Net: Prototype-Induced Label Structuring for Disentangled Multimodal Representation NetworkKehan Wang, Huan Zhao, Yong Wei, Xupeng Zha 等AAAI 2026
- Distillation-Guided Structural Transfer for Continual Learning Beyond Sparse Distributed MemoryHuiyan Xue, Xuming Ran, Yaxin Li, Qi Xu 等AAAI 2026 · 被引用 2 次
