Lune

ACL2021顶会

E2E-VLP: End-to-End Vision-Language Pre-training Enhanced by Visual Learning

Haiyang Xu, Ming Yan, Chenliang Li, Bin Bi, Songfang Huang, Wenming Xiao, Fei Huang

2021年份
33顶会引用

摘要

Vision-language pre-training (VLP) on largescale image-text pairs has achieved huge success for the cross-modal downstream tasks. The most existing pre-training methods mainly adopt a two-step training procedure, which firstly employs a pre-trained object detector to extract region-based visual features, then concatenates the image representation and text embedding as the input of Transformer to train. However, these methods face problems of using task-specific visual representation of the specific object detector for generic crossmodal understanding, and the computation inefficiency of two-stage pipeline.

In this paper, we propose the first end-to-end vision-language pre-trained model for both V+L understanding and generation, namely E2E-VLP, where we build a unified Transformer framework to jointly learn visual representation, and semantic alignments between image and text. We incorporate the tasks of object detection and image captioning into pretraining with a unified Transformer encoderdecoder architecture for enhancing visual learning. An extensive set of experiments have been conducted on well-established visionlanguage downstream tasks to demonstrate the effectiveness of this novel VLP paradigm.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 97050a06-71c2-4b75-8aa7-0b926905e364

引用它的顶会 Paper33

问问它们各自怎么用它

它引用的顶会 Paper11

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖