InfiniBench: A Benchmark for Large Multi-Modal Models in Long-Form Movies and TV Shows
Kirolos Ataallah, Eslam Mohamed Bakr, Mahmoud Ahmed, Chenhui Gou, Khushbu Pahwa, Jian Ding, Mohamed Elhoseiny
Abstract
Link ing Ev en ts Deep Contex t Un de rs ta n d in g S p o il er U nd er sta nding S u m m a ri za tio n T V S h o w s & Mo vies M o v ie s R e a s o n in g (O p e n - E n d ed ) Glob al Ap pe ar a n c e S c e n e T ra n si tio ns Cha racter Ac tio ns C h r o n o lo g ic a l U n d er sta nding TV S h o w s & M o v ie s T V S h o w s G r o u n ding (M C Q ) 00:00 30:26 Q: How does Celia react to Ross's monkey during the date? A: Celia yells as the monkey pulls her hair until Ross takes him away Q: What is the connection between Monica's failed dinner and Phoebe's reaction during Steve's next massage appointment? A: Monica's dinner is ruined by Steve showing up high. Later, an annoyed Phoebe gets revenge by giving him a painful massage.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fcfb585a-8624-41d2-84db-04faf50cdf40Cited by top-tier papers4
- Eyes Wide Open: Ego Proactive Video-LLM for Streaming VideoXueyang Yu, Cheng Shi, Yang Wang, Sibei YangNeurIPS 2025 · 34 citations
- MovieRecapsQA: A Multimodal Open-Ended Video Question-Answering BenchmarkShaden Shaar, Bradon Thymes, Sirawut Chaixanien, Claire Cardie et al.CVPR 2026 · 4 citations
- WeaveTime: Streaming from Earlier Frames into Emergent Memory in VideoLLMsYulin Zhang, Cheng Shi, Sibei YangCVPR 2026 · 1 citation
- PRIM:Cooperative Dynamic Token Compression for Efficient Large Multimodal ModelsSong Li, yongping xiongICML 2026
Builds on8
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-trainingLinjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan et al.EMNLP 2020 · 387 citations
- Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language ModelsMuhammad Maaz, Hanoona Abdul Rasheed, Salman Khan, Fahad KhanACL 2024 · 279 citations
- Video-LLaVA: Learning United Visual Representation by Alignment Before ProjectionBin Lin, Yang Ye, Bin Zhu, Jiaxi Cui et al.EMNLP 2024 · 231 citations
- VideoChat-Flash: Hierarchical Compression for Long-Context Video ModelingXinhao Li, Yi Wang, Jiashuo Yu, Xiangyu Zeng et al.ICLR 2026 · 172 citations
Related papers
- Digital Life Project: Autonomous 3D Characters with Social IntelligenceZhongang Cai, Jianping Jiang, Zhongfei Qing, Xinying Guo et al.CVPR 2024
- MECD: Unlocking Multi-Event Causal Discovery in Video ReasoningTieyuan Chen, Huabin Liu, Tianyao He, Yihang Chen et al.NeurIPS 2024 · 36 citations
- Learning Interactions and Relationships Between Movie CharactersAnna Kukleva, Makarand Tapaswi, Ivan LaptevCVPR 2020
- TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual CaptionsLinli Yao, Yuancheng Wei, Yaojie Zhang, Lei Li et al.ICML 2026
- Mind the Time: Temporally-Controlled Multi-Event Video GenerationZiyi Wu, Aliaksandr Siarohin, Willi Menapace, Ivan Skorokhodov et al.CVPR 2025
