Lune

EMNLP2024顶会

Training-free Deep Concept Injection Enables Language Models for Video Question Answering

Xudong Lin, Manling Li, Richard S. Zemel, Heng Ji, Shih-Fu Chang

2024年份
1被引次数
1顶会引用

摘要

Recently, enabling pretrained language models (PLMs) to perform zero-shot crossmodal tasks such as video question answering has been extensively studied. A popular approach is to learn a projection network that projects visual features into the input text embedding space of a PLM, as well as feed-forward adaptation layers, with the weights of the PLM frozen. However, is it really necessary to learn such additional layers? In this paper, we make the first attempt to demonstrate that the PLM is able to perform zero-shot crossmodal tasks without any crossmodal pretraining, when the observed visual concepts are injected as both additional input text tokens and augmentation in the intermediate features within each feed-forward network for the PLM. Specifically, inputting observed visual concepts as text tokens helps to inject them through the self-attention layers in the PLM; to augment the intermediate features in a way that is compatible with the PLM, we propose to construct adaptation layers based on the intermediate representation of concepts (obtained by solely inputting them to the PLM). These two complementary injection mechanisms form the proposed Deep Concept Injection, which comprehensively enables the PLM to perceive instantly without crossmodal pretraining. Extensive empirical analysis on zero-shot video question answering, as well as visual question answering, shows Deep Concept Injection achieves competitive or even better results in both zero-shot and fine-tuning settings, compared to state-of-the-art methods that require crossmodal pretraining. Input Video 𝑣𝑣 What is the woman wearing? [mask] ℱ 𝑉𝑉 Similarity Calculation 𝑃𝑃(𝑎𝑎|𝑣𝑣, 𝑡𝑡) Concept Vocabulary ℱ 𝑇𝑇 Concept Retrieval Eq. 1 Question 𝒢𝒢 Cook kitchen … What is the woman wearing ? [mask] Self-attention Feed-Forward × 𝑛𝑛 𝐵𝐵 Linear 2 of Block j Linear 1 of Block j 𝑒𝑒 0,𝑗𝑗

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper1

问问它们各自怎么用它

它引用的顶会 Paper24

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖