TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions
Linli Yao, Yuancheng Wei, Yaojie Zhang, Lei Li, Xinlong Chen, Feifan Song, Ziyue Wang, Kun Ouyang, Yuanxin Liu, Lingpeng Kong, Qi Liu, Pengfei Wan
Abstract
Timestamp:00:34 -00:41 Detailed Events:From an overhead view, the white car continues to drive in circles. Inside, the curly-haired male driver is Xialuo's brotherin-law. He is a middle-aged East Asian man wearing a dark suit jacket and a dark blue shirt. He urges his brother-in-law, Xialuo, to stop showing off, explaining that he must leave immediately because it's his girlfriend's 60th birthday and he took the car without permission. Visual Background:The scene, from a bird's-eye view, shows the paved driveway of the manor and lush green trees. Acoustics Content:(1)Speaking tone:The brother-in-law's tone is impatient. (2)Background sound or music: Light, cheerful music continues to play. Dialogue Content:Driver: 'Listen, brother-in-law, that's enough. I have to hurry back. It's my girlfriend's 60th birthday today, and she has no idea I took the car out.' Camera State:The shot begins with a high-angle medium-long shot of a car. The camera then moves down and pans to the upper right, capturing a full shot of the car. Subsequently, the scene cuts to an exterior extreme close-up of the car, showing a shaky close-up of the driver through the windshield. Shot Editing Style:The editing cuts between the high-angle shot and the driver's close-up, creating a contrast between the external scene of the car driving and the internal monologue that explains the driver's motivation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ee4e8412-0371-4bd5-bbaf-d3b4acaed417Builds on11
- End-to-End Dense Video Captioning with Parallel DecodingTeng Wang, Ruimao Zhang, Zhichao Lu, Feng Zheng et al.ICCV 2021 · 238 citations
- Time-R1: Post-Training Large Vision Language Model for Temporal Video GroundingYe Wang, Ziheng Wang, Boshen Xu, Yang Du et al.NeurIPS 2025 · 143 citations
- video-SALMONN: Speech-Enhanced Audio-Visual Large Language ModelsGuangzhi Sun, Wenyi Yu, Changli Tang, Xianzhao Chen et al.ICML 2024 · 92 citations
- TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video UnderstandingShuhuai Ren, Linli Yao, Shicheng Li, Xu Sun et al.CVPR 2024 · 83 citations
- OneStory: Coherent Multi-Shot Video Generation with Adaptive MemoryZhaochong An, Menglin Jia, Haonan Qiu, Zijian Zhou et al.CVPR 2026 · 33 citations
Related papers
- Cafe-Talk: Generating 3D Talking Face Animation with Multimodal Coarse- and Fine-grained ControlHejia Chen, Haoxian Zhang, Shoulong Zhang, Xiaoqiang Liu et al.ICLR 2025
- Crab: A Unified Audio-Visual Scene Understanding Model with Explicit CooperationHenghui Du, Guangyao Li, Chang Zhou, Chunjie Zhang et al.CVPR 2025
- Score Jacobian Chaining: Lifting Pretrained 2D Diffusion Models for 3D GenerationHaochen Wang, Xiaodan Du, Jiahao Li, Raymond A. Yeh et al.CVPR 2023
- EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question AnsweringSheng Zhou, Junbin Xiao, Qingyun Li, Yicong Li et al.CVPR 2025
- InfiniBench: A Benchmark for Large Multi-Modal Models in Long-Form Movies and TV ShowsKirolos Ataallah, Eslam Mohamed Bakr, Mahmoud Ahmed, Chenhui Gou et al.EMNLP 2025
