Lune

CVPR2024Top-tier venue

Enhancing Vision-Language Pre-Training with Rich Supervisions

Yuan Gao, Kunyu Shi, Pengkai Zhu, Edouard Belval, Oren Nuriel, Srikar Appalaraju, Shabnam Ghadar, Zhuowen Tu, Vijay Mahadevan, Stefano Soatto

2024Year
4Top-tier citations

Abstract

We propose Strongly Supervised pre-training with ScreenShots (S4) -a novel pre-training paradigm for Vision-Language Models using data from large-scale web screenshot rendering. Using web screenshots unlocks a treasure trove of visual and textual cues that are not present in using image-text pairs. In S4, we leverage the inherent tree-structured hierarchy of HTML elements and the spatial localization to carefully design 10 pre-training tasks with large scale annotated data. These tasks resemble downstream tasks across different domains and the annotations are cheap to obtain. We demonstrate that, compared to current screenshot pre-training objectives, our innovative pre-training method significantly enhances performance of image-to-text model in nine varied and popular downstream tasks -up to 76.1% improvements on Table Detection , and at least 1% on Widget Captioning.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext df1d0567-e6d1-4327-97ee-10d64098f9ed

Cited by top-tier papers4

Ask how each one uses it

Builds on43

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines