SEED-Bench: Benchmarking Multimodal Large Language Models
Bohao Li, Yuying Ge, Yixiao Ge, Guangzhi Wang, Rui Wang, Ruimao Zhang, Ying Shan
摘要
L 4 H i e r a r c h i c a l T a s k L e v e l C o r r e s p o n d i n g M o d e l C o r r e s p o n d i n g B e n c h m a r k * Equal Contribution. † Correspondence Author. Can you recognize the actions that occur in this video and list them in order? A. Cook breakfast, switch stove on, close fridge, carry milk, peel banana B. Scoop ice cream, squeeze chocolate syrup, pour sprinkles, close fridge C. Close fridge, carry milk, screw open milk cap, pour milk, screw close milk cap D. Reach for cereal box, grab bowl, pour milk, stir cereal, close fridge Procedure Understanding time What action do you anticipate following the end of this video? A. Stir potatoes B. Wash potatoes C. Add potatoes D. Slice potatoes Action Prediction time What is the action being carried out in the video? A. Throwing something in the air and letting it fall B. Throwing something in the air and catching it C. Lifting up one end of something, then letting it drop down D. Poking something so that it falls over Action Recognition time What are the differences between the two image? A. In the second image, there are two people standing on the sidewalk instead of three and a car is just entering the parking lot. B. In the second image, there are four people standing on the sidewalk instead of three and a car is just leaving the parking lot. C. In the second image, there are three people standing on the sidewalk instead of two and a car is just entering the parking lot. D. In the second image, there are two people standing on the sidewalk instead of three and a car is just leaving the parking lot. Difference Spotting What is funny about this comic strip? A. The polar bear entered the bus pavilion with a Dalmatian, but the bus pavilion was a dog without Dalmatian. B. The Dalmatian and bear are in the rain. C. This is a fake Dalmatian. D. The rain cleaned the Dalmatian's spots.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper144
- We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?Runqi Qiao, Qiuna Tan, Guanting Dong, Minhui Wu 等ACL 2025 · 被引用 236 次
- WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMsJack Hong, Shilin Yan, Jiayin Cai, Xiaolong Jiang 等ICLR 2026 · 被引用 162 次
- Cambrian-S: Towards Spatial Supersensing in VideoShusheng Yang, Jihan Yang, Pinzhi Huang, Ellis Brown 等ICLR 2026 · 被引用 139 次
- OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMsCaorui Li, Yu Chen, Yiyan Ji, Jin Xu 等ICLR 2026 · 被引用 53 次
- UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and EditingHao Tang, Chen-Wei Xie, Xiaoyi Bao, Tingyu Weng 等ICLR 2026 · 被引用 45 次
它引用的顶会 Paper10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
相关 Paper
- Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric VideosChiara Plizzari, Alessio Tonioni, Yongqin Xian, Achin Kulshrestha 等CVPR 2025
- Action-Sketcher: From Reasoning to Action via Visual Sketches for Robotic ManipulationHuajie Tan, Peterson Co, Yijie Xu, Shanyu Rong 等CVPR 2026
- UFVideo: Towards Unified Fine-Grained Video Cooperative Understanding with Large Language ModelsHewen Pan, Cong Wei, Dashuang Liang, Zepeng Huang 等CVPR 2026 · 被引用 3 次
- Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware ReasoningSARA GHAZANFARI, Francesco Croce, Nicolas Flammarion, Prashanth Krishnamurthy 等CVPR 2026 · 被引用 35 次
- Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video ContentQiuheng Wang, Yukai Shi, Jiarong Ou, Rui Chen 等CVPR 2025
