SF2T: Self-supervised Fragment Finetuning of Video-LLMs for Fine-Grained Understanding
Yangliu Hu, Zikai Song, Na Feng, Yawei Luo, Junqing Yu, Yi-Ping Phoebe Chen, Wei Yang
Abstract
Video-based Large Language Models (Video-LLMs) have witnessed substantial advancements in recent years, propelled by the advancement in multi-modal LLMs. Although these models have demonstrated proficiency in providing the overall description of videos, they struggle with fine-grained understanding, particularly in aspects such as visual dynamics and video details inquiries. To tackle these shortcomings, we find that fine-tuning Video-LLMs on self-supervised fragment tasks, greatly improve their finegrained video understanding abilities. Hence we propose two key contributions: (1) Self-Supervised Fragment Fine-Tuning (SF 2 T), a novel effortless fine-tuning method, employs the rich inherent characteristics of videos for training, while unlocking more fine-grained understanding ability of Video-LLMs. Moreover, it relieves researchers from labor-intensive annotations and smartly circumvents the limitations of natural language, which often fails to capture the complex spatiotemporal variations in videos; (2) A novel benchmark dataset, namely FineVidBench, for rigorously assessing Video-LLMs' performance at both the scene and fragment levels, offering a comprehensive evaluation of their capabilities. We assessed multiple models and validated the effectiveness of SF 2 T on them. Experimental results reveal that our approach improves their ability to capture and interpret spatiotemporal details.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ec912792-5bfc-49cb-bd69-3e952fbdf021Cited by top-tier papers13
- Logical Phase Transitions: Understanding Collapse in LLM Logical ReasoningXinglang Zhang, Yunyao Zhang, ZeLiang Chen, Junqing Yu et al.ACL 2026 · 28 citations
- Semantic-Aware Logical Reasoning via a Semiotic FrameworkYunyao Zhang, Xinglang Zhang, Junxi Sheng, Wenbing Li et al.ACL 2026 · 27 citations
- ReTrack: Evidence-Driven Dual-Stream Directional Anchor Calibration Network for Composed Video RetrievalZixu Li, Yupeng Hu, Zhiwei Chen, Qinlei Huang et al.AAAI 2026 · 24 citations
- LoRA-Mixer: Coordinate Modular LoRA Experts Through Serial Attention RoutingWenbing Li, Zikai Song, Hang Zhou, Junqing Yu et al.ICLR 2026 · 20 citations
- ReasonMap: Towards Fine-Grained Visual Reasoning from Transit MapsSicheng Feng, Song Wang, Shuyi Ouyang, Lingdong Kong et al.CVPR 2026 · 19 citations
Builds on17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- TIME: Temporal-Sensitive Multi-Dimensional Instruction Tuning and Robust Benchmarking for Video-LLMsYunxiao Wang, Meng Liu, Wenqi Liu, Xuemeng Song et al.AAAI 2026 · 1 citation
- MVBench: A Comprehensive Multi-modal Video Understanding BenchmarkKunchang Li, Yali Wang, Yinan He, Yizhuo Li et al.CVPR 2024
- Momentor: Advancing Video Large Language Model with Fine-Grained Temporal ReasoningLong Qian, Juncheng Li, Yu Wu, Yaobo Ye et al.ICML 2024 · 121 citations
- CaReBench: A Fine-grained Benchmark for Video Captioning and RetrievalYifan Xu, Xinhao Li, Yichun Yang, Desen Meng et al.ICLR 2026 · 10 citations
- VideoSSR: Video Self-Supervised Reinforcement LearningZefeng He, Xiaoye Qu, Yafu Li, Siyuan Huang et al.CVPR 2026 · 4 citations
