Vokenization: Improving Language Understanding with Contextualized, Visual-Grounded Supervision
Hao Tan, Mohit Bansal
摘要
Humans learn language by listening, speaking, writing, reading, and also, via interaction with the multimodal real world. Existing language pre-training frameworks show the effectiveness of text-only self-supervision while we explore the idea of a visually-supervised language model in this paper. We find that the main reason hindering this exploration is the large divergence in magnitude and distributions between the visually-grounded language datasets and pure-language corpora. Therefore, we develop a technique named "vokenization" that extrapolates multimodal alignments to language-only data by contextually mapping language tokens to their related images (which we call "vokens"). The "vokenizer" is trained on relatively small image captioning datasets and we then apply it to generate vokens for large language corpora. Trained with these contextually generated vokens, our visually-supervised language models show consistent improvements over self-supervised alternatives on multiple purelanguage tasks such as GLUE, SQuAD, and SWAG. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper28
- Keeping Your Eye on the Ball: Trajectory Attention in Video TransformersMandela Patrick, Dylan Campbell, Yuki M. Asano, Ishan Misra 等NeurIPS 2021 · 被引用 382 次
- UniT: Multimodal Multitask Learning with a Unified TransformerRonghang Hu, Amanpreet SinghICCV 2021 · 被引用 354 次
- BridgeTower: Building Bridges between Encoders in Vision-Language Representation LearningXiao Xu, Chenfei Wu, Shachar Rosenman, Vasudev Lal 等AAAI 2023 · 被引用 99 次
- On Uni-Modal Feature Learning in Supervised Multi-Modal LearningChenzhuang Du, Jiaye Teng, Tingle Li, Yichen Liu 等ICML 2023 · 被引用 79 次
- VidLanKD: Improving Language Understanding via Video-Distilled Knowledge TransferZineng Tang, Jaemin Cho, Hao Tan, Mohit BansalNeurIPS 2021 · 被引用 36 次
它引用的顶会 Paper8
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li 等ICLR 2020 · 被引用 1,825 次
- Unified Vision-Language Pre-Training for Image Captioning and VQALuowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu 等AAAI 2020 · 被引用 1,047 次
- Climbing towards NLU: On Meaning, Form, and Understanding in the Age of DataEmily M. Bender, Alexander KollerACL 2020 · 被引用 914 次
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
相关 Paper
- Cross-Lingual Transfer of Large Language Model by Visually-Derived Supervision Toward Low-Resource LanguagesMasayasu Muraoka, Bishwaranjan Bhattacharjee, Michele Merler, Graeme Blackwood 等ACM MM 2023 · 被引用 6 次
- Visually-Augmented Language ModelingWeizhi Wang, Li Dong, Hao Cheng, Haoyu Song 等ICLR 2023 · 被引用 5 次
- Enhancing Sentence Representation with Visually-supervised Multimodal Pre-trainingZhe Li, Laurence T. Yang, Xin Nie, Bocheng Ren 等ACM MM 2023 · 被引用 5 次
- Expand BERT Representation with Visual Information via Grounded Language Learning with Multimodal Partial AlignmentCong-Duy Nguyen, The-Anh Vu-Le, Thong Nguyen, Tho Quan 等ACM MM 2023 · 被引用 3 次
- Language-Driven Cross-Modal Classifier for Zero-Shot Multi-Label Image RecognitionYicheng Liu, Jie Wen, Chengliang Liu, Xiaozhao Fang 等ICML 2024 · 被引用 7 次
