OmniVec2 - A Novel Transformer Based Network for Large Scale Multimodal and Multitask Learning
Siddharth Srivastava, Gaurav Sharma
摘要
We present a novel multimodal multitask network and associated training algorithm. The method is capable of ingesting data from approximately 12 different modalities namely image, video, audio, text, depth, point cloud, time series, tabular, graph, X-ray, infrared, IMU, and hyper-spectral. The proposed approach utilizes modality specialized tokenizers, a shared transformer architecture, and cross-attention mechanisms to project the data from different modalities into a unified embedding space. It addresses multimodal and multitask scenarios by incorpo-rating modality-specific task heads for different tasks in respective modalities. We propose a novel pretraining strategy with iterative modality switching to initialize the network, and a training algorithm which trades off fully joint training over all modalities, with training on pairs of modalities at a time. We provide comprehensive evaluation across 25 datasets from 12 modalities and show state of the art performances, demonstrating the effectiveness of the proposed architecture, pretraining strategy and adapted multitask training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- TRIBE: TRImodal Brain Encoder for whole-brain fMRI response predictionStéphane d'Ascoli, Jérémy Rapin, Yohann Benchetrit, Hubert Banville 等ICLR 2026 · 被引用 28 次
- GenRec: Unifying Video Generation and Recognition with Diffusion ModelsZejia Weng, Xitong Yang, Zhen Xing, Zuxuan Wu 等NeurIPS 2024 · 被引用 19 次
- VAEmo: Efficient Representation Learning for Visual-Audio Emotion With Knowledge InjectionHao Cheng, Zhiwei Zhao, Yichao He, Zhenzhen Hu 等ACM MM 2025 · 被引用 9 次
- ReasonAct: Progressive Training for Fine-Grained Video Reasoning in Small ModelsJiaxin Liu, Zhaolu KangAAAI 2026 · 被引用 4 次
- ProM3E: Probabilistic Masked MultiModal Embedding Model for EcologySrikumar Sastry, Subash Khanal, Aayush Dhakal, Jiayu Lin 等CVPR 2026 · 被引用 1 次
它引用的顶会 Paper32
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang 等AAAI 2021 · 被引用 7,289 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu 等NeurIPS 2020 · 被引用 1,957 次
- Do Transformers Really Perform Badly for Graph Representation?Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng 等NeurIPS 2021 · 被引用 1,632 次
相关 Paper
- mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and VideoHaiyang Xu, Qinghao Ye, Ming Yan, Yaya Shi 等ICML 2023 · 被引用 237 次
- UniT: Multimodal Multitask Learning with a Unified TransformerRonghang Hu, Amanpreet SinghICCV 2021 · 被引用 354 次
- M3P: Learning Universal Representations via Multitask Multilingual Multimodal Pre-TrainingMinheng Ni, Haoyang Huang, Lin Su, Edward Cui 等CVPR 2021
- Alternating Gradient Descent and Mixture-of-Experts for Integrated Multimodal PerceptionHassan Akbari, Dan Kondratyuk, Yin Cui, Rachel Hornung 等NeurIPS 2023 · 被引用 33 次
- Uni-Perceiver: Pre-training Unified Architecture for Generic Perception for Zero-shot and Few-shot TasksXizhou Zhu, Jinguo Zhu, Hao Li, Xiaoshi Wu 等CVPR 2022
