SF2T: Self-supervised Fragment Finetuning of Video-LLMs for Fine-Grained Understanding
Yangliu Hu, Zikai Song, Na Feng, Yawei Luo, Junqing Yu, Yi-Ping Phoebe Chen, Wei Yang
摘要
Video-based Large Language Models (Video-LLMs) have witnessed substantial advancements in recent years, propelled by the advancement in multi-modal LLMs. Although these models have demonstrated proficiency in providing the overall description of videos, they struggle with fine-grained understanding, particularly in aspects such as visual dynamics and video details inquiries. To tackle these shortcomings, we find that fine-tuning Video-LLMs on self-supervised fragment tasks, greatly improve their finegrained video understanding abilities. Hence we propose two key contributions: (1) Self-Supervised Fragment Fine-Tuning (SF 2 T), a novel effortless fine-tuning method, employs the rich inherent characteristics of videos for training, while unlocking more fine-grained understanding ability of Video-LLMs. Moreover, it relieves researchers from labor-intensive annotations and smartly circumvents the limitations of natural language, which often fails to capture the complex spatiotemporal variations in videos; (2) A novel benchmark dataset, namely FineVidBench, for rigorously assessing Video-LLMs' performance at both the scene and fragment levels, offering a comprehensive evaluation of their capabilities. We assessed multiple models and validated the effectiveness of SF 2 T on them. Experimental results reveal that our approach improves their ability to capture and interpret spatiotemporal details.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Logical Phase Transitions: Understanding Collapse in LLM Logical ReasoningXinglang Zhang, Yunyao Zhang, ZeLiang Chen, Junqing Yu 等ACL 2026 · 被引用 28 次
- Semantic-Aware Logical Reasoning via a Semiotic FrameworkYunyao Zhang, Xinglang Zhang, Junxi Sheng, Wenbing Li 等ACL 2026 · 被引用 27 次
- ReTrack: Evidence-Driven Dual-Stream Directional Anchor Calibration Network for Composed Video RetrievalZixu Li, Yupeng Hu, Zhiwei Chen, Qinlei Huang 等AAAI 2026 · 被引用 24 次
- LoRA-Mixer: Coordinate Modular LoRA Experts Through Serial Attention RoutingWenbing Li, Zikai Song, Hang Zhou, Junqing Yu 等ICLR 2026 · 被引用 20 次
- ReasonMap: Towards Fine-Grained Visual Reasoning from Transit MapsSicheng Feng, Song Wang, Shuyi Ouyang, Lingdong Kong 等CVPR 2026 · 被引用 19 次
它引用的顶会 Paper17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
相关 Paper
- TIME: Temporal-Sensitive Multi-Dimensional Instruction Tuning and Robust Benchmarking for Video-LLMsYunxiao Wang, Meng Liu, Wenqi Liu, Xuemeng Song 等AAAI 2026 · 被引用 1 次
- MVBench: A Comprehensive Multi-modal Video Understanding BenchmarkKunchang Li, Yali Wang, Yinan He, Yizhuo Li 等CVPR 2024
- Momentor: Advancing Video Large Language Model with Fine-Grained Temporal ReasoningLong Qian, Juncheng Li, Yu Wu, Yaobo Ye 等ICML 2024 · 被引用 121 次
- CaReBench: A Fine-grained Benchmark for Video Captioning and RetrievalYifan Xu, Xinhao Li, Yichun Yang, Desen Meng 等ICLR 2026 · 被引用 10 次
- VideoSSR: Video Self-Supervised Reinforcement LearningZefeng He, Xiaoye Qu, Yafu Li, Siyuan Huang 等CVPR 2026 · 被引用 4 次
