TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions
Linli Yao, Yuancheng Wei, Yaojie Zhang, Lei Li, Xinlong Chen, Feifan Song, Ziyue Wang, Kun Ouyang, Yuanxin Liu, Lingpeng Kong, Qi Liu, Pengfei Wan
摘要
Timestamp:00:34 -00:41 Detailed Events:From an overhead view, the white car continues to drive in circles. Inside, the curly-haired male driver is Xialuo's brotherin-law. He is a middle-aged East Asian man wearing a dark suit jacket and a dark blue shirt. He urges his brother-in-law, Xialuo, to stop showing off, explaining that he must leave immediately because it's his girlfriend's 60th birthday and he took the car without permission. Visual Background:The scene, from a bird's-eye view, shows the paved driveway of the manor and lush green trees. Acoustics Content:(1)Speaking tone:The brother-in-law's tone is impatient. (2)Background sound or music: Light, cheerful music continues to play. Dialogue Content:Driver: 'Listen, brother-in-law, that's enough. I have to hurry back. It's my girlfriend's 60th birthday today, and she has no idea I took the car out.' Camera State:The shot begins with a high-angle medium-long shot of a car. The camera then moves down and pans to the upper right, capturing a full shot of the car. Subsequently, the scene cuts to an exterior extreme close-up of the car, showing a shaky close-up of the driver through the windshield. Shot Editing Style:The editing cuts between the high-angle shot and the driver's close-up, creating a contrast between the external scene of the car driving and the internal monologue that explains the driver's motivation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- End-to-End Dense Video Captioning with Parallel DecodingTeng Wang, Ruimao Zhang, Zhichao Lu, Feng Zheng 等ICCV 2021 · 被引用 238 次
- Time-R1: Post-Training Large Vision Language Model for Temporal Video GroundingYe Wang, Ziheng Wang, Boshen Xu, Yang Du 等NeurIPS 2025 · 被引用 143 次
- video-SALMONN: Speech-Enhanced Audio-Visual Large Language ModelsGuangzhi Sun, Wenyi Yu, Changli Tang, Xianzhao Chen 等ICML 2024 · 被引用 92 次
- TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video UnderstandingShuhuai Ren, Linli Yao, Shicheng Li, Xu Sun 等CVPR 2024 · 被引用 83 次
- OneStory: Coherent Multi-Shot Video Generation with Adaptive MemoryZhaochong An, Menglin Jia, Haonan Qiu, Zijian Zhou 等CVPR 2026 · 被引用 33 次
相关 Paper
- Cafe-Talk: Generating 3D Talking Face Animation with Multimodal Coarse- and Fine-grained ControlHejia Chen, Haoxian Zhang, Shoulong Zhang, Xiaoqiang Liu 等ICLR 2025
- Crab: A Unified Audio-Visual Scene Understanding Model with Explicit CooperationHenghui Du, Guangyao Li, Chang Zhou, Chunjie Zhang 等CVPR 2025
- Score Jacobian Chaining: Lifting Pretrained 2D Diffusion Models for 3D GenerationHaochen Wang, Xiaodan Du, Jiahao Li, Raymond A. Yeh 等CVPR 2023
- EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question AnsweringSheng Zhou, Junbin Xiao, Qingyun Li, Yicong Li 等CVPR 2025
- InfiniBench: A Benchmark for Large Multi-Modal Models in Long-Form Movies and TV ShowsKirolos Ataallah, Eslam Mohamed Bakr, Mahmoud Ahmed, Chenhui Gou 等EMNLP 2025
