InfiniBench: A Benchmark for Large Multi-Modal Models in Long-Form Movies and TV Shows
Kirolos Ataallah, Eslam Mohamed Bakr, Mahmoud Ahmed, Chenhui Gou, Khushbu Pahwa, Jian Ding, Mohamed Elhoseiny
摘要
Link ing Ev en ts Deep Contex t Un de rs ta n d in g S p o il er U nd er sta nding S u m m a ri za tio n T V S h o w s & Mo vies M o v ie s R e a s o n in g (O p e n - E n d ed ) Glob al Ap pe ar a n c e S c e n e T ra n si tio ns Cha racter Ac tio ns C h r o n o lo g ic a l U n d er sta nding TV S h o w s & M o v ie s T V S h o w s G r o u n ding (M C Q ) 00:00 30:26 Q: How does Celia react to Ross's monkey during the date? A: Celia yells as the monkey pulls her hair until Ross takes him away Q: What is the connection between Monica's failed dinner and Phoebe's reaction during Steve's next massage appointment? A: Monica's dinner is ruined by Steve showing up high. Later, an annoyed Phoebe gets revenge by giving him a painful massage.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Eyes Wide Open: Ego Proactive Video-LLM for Streaming VideoXueyang Yu, Cheng Shi, Yang Wang, Sibei YangNeurIPS 2025 · 被引用 34 次
- MovieRecapsQA: A Multimodal Open-Ended Video Question-Answering BenchmarkShaden Shaar, Bradon Thymes, Sirawut Chaixanien, Claire Cardie 等CVPR 2026 · 被引用 4 次
- WeaveTime: Streaming from Earlier Frames into Emergent Memory in VideoLLMsYulin Zhang, Cheng Shi, Sibei YangCVPR 2026 · 被引用 1 次
- PRIM:Cooperative Dynamic Token Compression for Efficient Large Multimodal ModelsSong Li, yongping xiongICML 2026
它引用的顶会 Paper8
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
- HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-trainingLinjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan 等EMNLP 2020 · 被引用 387 次
- Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language ModelsMuhammad Maaz, Hanoona Abdul Rasheed, Salman Khan, Fahad KhanACL 2024 · 被引用 279 次
- Video-LLaVA: Learning United Visual Representation by Alignment Before ProjectionBin Lin, Yang Ye, Bin Zhu, Jiaxi Cui 等EMNLP 2024 · 被引用 231 次
- VideoChat-Flash: Hierarchical Compression for Long-Context Video ModelingXinhao Li, Yi Wang, Jiashuo Yu, Xiangyu Zeng 等ICLR 2026 · 被引用 172 次
相关 Paper
- Digital Life Project: Autonomous 3D Characters with Social IntelligenceZhongang Cai, Jianping Jiang, Zhongfei Qing, Xinying Guo 等CVPR 2024
- MECD: Unlocking Multi-Event Causal Discovery in Video ReasoningTieyuan Chen, Huabin Liu, Tianyao He, Yihang Chen 等NeurIPS 2024 · 被引用 36 次
- Learning Interactions and Relationships Between Movie CharactersAnna Kukleva, Makarand Tapaswi, Ivan LaptevCVPR 2020
- TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual CaptionsLinli Yao, Yuancheng Wei, Yaojie Zhang, Lei Li 等ICML 2026
- Mind the Time: Temporally-Controlled Multi-Event Video GenerationZiyi Wu, Aliaksandr Siarohin, Willi Menapace, Ivan Skorokhodov 等CVPR 2025
