Transitional Adaptation of Pretrained Models for Visual Storytelling
Youngjae Yu, Jiwan Chung, Heeseung Yun, Jongseok Kim, Gunhee Kim
摘要
Previous models for vision-to-language generation tasks usually pretrain a visual encoder and a language generator in the respective domains and jointly finetune them with the target task. However, this direct transfer practice may suffer from the discord between visual specificity and language fluency since they are often separately trained from large corpora of visual and text data with no common ground. In this work, we claim that a transitional adaptation task is required between pretraining and finetuning to harmonize the visual encoder and the language model for challenging downstream target tasks like visual storytelling. We propose a novel approach named Transitional Adaptation of Pretrained Model (TAPM) that adapts the multi-modal modules to each other with a simpler alignment task between visual inputs only with no need for text labels. Through extensive experiments, we show that the adaptation step significantly improves the performance of multiple language models for sequential video and image captioning tasks. We achieve new state-of-the-art performance on both language metrics and human evaluation in the multi-sentence description task of LSMDC 2019 [50] and the image storytelling task of VIST [18]. Our experiments reveal that this improvement in caption quality does not depend on the specific choice of language models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- AutoAD II: The Sequel - Who, When, and What in Movie Audio DescriptionTengda Han, Max Bain, Arsha Nagrani, Gül Varol 等ICCV 2023 · 被引用 55 次
- Memory Reviver: Supporting Photo-Collection Reminiscence for People with Visual Impairment via a Proactive ChatbotShuchang Xu, Chang Chen, Zichen Liu, Xiaofu Jin 等UIST 2024 · 被引用 21 次
- Ordered Attention for Coherent Visual StorytellingTom Braude, Idan Schwartz, Alexander G. Schwing, Ariel ShamirACM MM 2022 · 被引用 12 次
- Text-Only Training for Visual StorytellingYuechen Wang, Wengang Zhou, Zhenbo Lu, Houqiang LiACM MM 2023 · 被引用 4 次
- Engage for All: Making Ordinary Image Descriptions Appealing Again!Yuyan Chen, Yifan Jiang, Li Zhou, Jinghan Cao 等ICCV 2025 · 被引用 1 次
它引用的顶会 Paper7
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li 等ICLR 2020 · 被引用 1,825 次
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy 等ICCV 2019 · 被引用 1,396 次
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong 等AAAI 2020 · 被引用 966 次
- Taking a HINT: Leveraging Explanations to Make Vision and Language Models More GroundedRamprasaath Ramasamy Selvaraju, Stefan Lee, Yilin Shen, Hongxia Jin 等ICCV 2019 · 被引用 288 次
- MART: Memory-Augmented Recurrent Transformer for Coherent Video Paragraph CaptioningJie Lei, Liwei Wang, Yelong Shen, Dong Yu 等ACL 2020 · 被引用 168 次
相关 Paper
- Adaptively Building a Video-language Model for Video Captioning and Retrieval without Massive Video PretrainingZihao Liu, Xiaoyu Wu, Shengjin Wang, Jiayao QianACM MM 2024 · 被引用 1 次
- Bridging Vision and Language Spaces with Assignment PredictionJungin Park, Jiyoung Lee, Kwanghoon SohnICLR 2024 · 被引用 15 次
- TAP: Text-Aware Pre-Training for Text-VQA and Text-CaptionZhengyuan Yang, Yijuan Lu, Jianfeng Wang, Xi Yin 等CVPR 2021
- E2E-VLP: End-to-End Vision-Language Pre-training Enhanced by Visual LearningHaiyang Xu, Ming Yan, Chenliang Li, Bin Bi 等ACL 2021
- eP-ALM: Efficient Perceptual Augmentation of Language ModelsMustafa Shukor, Corentin Dancette, Matthieu CordICCV 2023 · 被引用 36 次
