MoS2: Mixture of Scale and Shift Experts for Text-Only Video Captioning
Heng Jia, Yunqiu Xu, Linchao Zhu, Guang Chen, Yufei Wang, Yi Yang
Abstract
Video captioning is a challenging task and typically requires paired video-text data for training. However, manually annotating coherent textual descriptions for videos is laborious and time-consuming. To address this challenge, we propose a novel approach that enhances video captioning using only synthetic text data. Leveraging the exceptional text generation capabilities of large language models (LLMs), we produce high-quality and diverse video captions tailored to the target domain. Our approach employs a two-stage prompting strategy: first prompt GPT-4 with few-shot target-domain captions to create a set of high-quality captions, and then continue prompting with the generated captions to acquire large-scale synthetic data. To effectively utilize these captions, we introduce Mixture of Scale and Shift experts (MoS2), an efficient adaptation method for pre-trained captioning models. MoS2 employs lightweight routing networks to estimate probability distributions over a collection of scale and shift experts, dynamically allocating tokens to the appropriate experts. This dynamic adjustment mechanism enhances the model's ability to handle data variations and mitigates the distribution shift between synthetic and real captions. Moreover, our method reduces the number of learnable parameters, facilitating more efficient adaptation. Our method achieves superior performance with only synthetic text data, narrowing the gap between zero-shot and fine-tuned models and reducing the dependency on paired data from the target domain.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 1ebcdd2d-658b-4d31-ae02-6d65ceffc5b2Cited by top-tier papers5
- MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical ProblemsShuhang Chen, Hangjie Yuan, Yunqiu Xu, Pengwei Liu et al.ACL 2026 · 9 citations
- MC-Bench: A Benchmark for Multi-Context Visual Grounding in the Era of MLLMsYunqiu Xu, Linchao Zhu, Yi YangICCV 2025 · 7 citations
- Clear Nights Ahead: Towards Multi-Weather Nighttime Image RestorationYuetong Liu, Yunqiu Xu, Yang Wei, Xiuli Bi et al.AAAI 2026 · 5 citations
- CogFlow: Bridging Perception and Reasoning through Knowledge Internalization for Visual Mathematical Problem SolvingShuhang Chen, Yunqiu Xu, Junjie Xie, Aojun Lu et al.ICLR 2026 · 4 citations
- VideoGrain: Modulating Space-Time Attention for Multi-Grained Video EditingXiangpeng Yang, Linchao Zhu, Hehe Fan, Yi YangICLR 2025
Related papers
- Cap4Video: What Can Auxiliary Captions Do for Text-Video Retrieval?Wenhao Wu, Haipeng Luo, Bo Fang, Jingdong Wang et al.CVPR 2023
- MultiCapCLIP: Auto-Encoding Prompts for Zero-Shot Multilingual Visual CaptioningBang Yang, Fenglin Liu, Xian Wu, Yaowei Wang et al.ACL 2023 · 10 citations
- Image Captioning with Multi-Context Synthetic DataFeipeng Ma, Yizhou Zhou, Fengyun Rao, Yueyi Zhang et al.AAAI 2024 · 22 citations
- TOPA: Extending Large Language Models for Video Understanding via Text-Only Pre-AlignmentWei Li, Hehe Fan, Yongkang Wong, Mohan S. Kankanhalli et al.NeurIPS 2024 · 18 citations
- Distilling Vision-Language Models on Millions of VideosYue Zhao, Long Zhao, Xingyi Zhou, Jialin Wu et al.CVPR 2024
