Lune

CVPR2026顶会

InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding

Ashutosh Kumar, Rajat Saini, Jingjing Pan, Mustafa Erdogan, Mingfang Zhang, Betty Le, Norimasa Kobori, Quan Kong

2026年份

摘要

Current vision–language pre-training (VLP) paradigms excel at global scene understanding but struggle with instance-level reasoning due to global-only supervision. We introduce InstAP\textbf{InstAP}, an Inst\textbf{Inst}ance-A\textbf{A}ware P\textbf{P}re-training framework that jointly optimizes global image–text alignment and fine-grained, instance-level contrastive alignment by grounding textual mentions to specific spatial–temporal regions. To support this, we present InstVL\textbf{InstVL}, a large-scale dataset (22 million images, 50,00050,000 videos) with dual-granularity annotations: holistic scene captions and dense, grounded instance descriptions. On the InstVL benchmark, InstAP substantially outperforms existing VLP models on instance-level retrieval, and also surpasses a strong VLP baseline trained on the exact same data corpus, isolating the benefit of our instance-aware objective. Moreover, instance-centric pre-training improves global understanding: InstAP achieves competitive zero-shot performance on multiple video benchmarks, including MSR-VTT and DiDeMo. Qualitative visualizations further show that InstAP localizes textual mentions to the correct instances, while global-only models exhibit more diffuse, scene-level attention.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext a7805fb0-bebd-4cdf-a58f-2c4dd7ed585e

它引用的顶会 Paper36

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖