SEED-Bench: Benchmarking Multimodal Large Language Models
Bohao Li, Yuying Ge, Yixiao Ge, Guangzhi Wang, Rui Wang, Ruimao Zhang, Ying Shan
Abstract
L 4 H i e r a r c h i c a l T a s k L e v e l C o r r e s p o n d i n g M o d e l C o r r e s p o n d i n g B e n c h m a r k * Equal Contribution. † Correspondence Author. Can you recognize the actions that occur in this video and list them in order? A. Cook breakfast, switch stove on, close fridge, carry milk, peel banana B. Scoop ice cream, squeeze chocolate syrup, pour sprinkles, close fridge C. Close fridge, carry milk, screw open milk cap, pour milk, screw close milk cap D. Reach for cereal box, grab bowl, pour milk, stir cereal, close fridge Procedure Understanding time What action do you anticipate following the end of this video? A. Stir potatoes B. Wash potatoes C. Add potatoes D. Slice potatoes Action Prediction time What is the action being carried out in the video? A. Throwing something in the air and letting it fall B. Throwing something in the air and catching it C. Lifting up one end of something, then letting it drop down D. Poking something so that it falls over Action Recognition time What are the differences between the two image? A. In the second image, there are two people standing on the sidewalk instead of three and a car is just entering the parking lot. B. In the second image, there are four people standing on the sidewalk instead of three and a car is just leaving the parking lot. C. In the second image, there are three people standing on the sidewalk instead of two and a car is just entering the parking lot. D. In the second image, there are two people standing on the sidewalk instead of three and a car is just leaving the parking lot. Difference Spotting What is funny about this comic strip? A. The polar bear entered the bus pavilion with a Dalmatian, but the bus pavilion was a dog without Dalmatian. B. The Dalmatian and bear are in the rain. C. This is a fake Dalmatian. D. The rain cleaned the Dalmatian's spots.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers144
- We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?Runqi Qiao, Qiuna Tan, Guanting Dong, Minhui Wu et al.ACL 2025 · 236 citations
- WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMsJack Hong, Shilin Yan, Jiayin Cai, Xiaolong Jiang et al.ICLR 2026 · 162 citations
- Cambrian-S: Towards Spatial Supersensing in VideoShusheng Yang, Jihan Yang, Pinzhi Huang, Ellis Brown et al.ICLR 2026 · 139 citations
- OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMsCaorui Li, Yu Chen, Yiyan Ji, Jin Xu et al.ICLR 2026 · 53 citations
- UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and EditingHao Tang, Chen-Wei Xie, Xiaoyi Bao, Tingyu Weng et al.ICLR 2026 · 45 citations
Builds on10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
Related papers
- Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric VideosChiara Plizzari, Alessio Tonioni, Yongqin Xian, Achin Kulshrestha et al.CVPR 2025
- Action-Sketcher: From Reasoning to Action via Visual Sketches for Robotic ManipulationHuajie Tan, Peterson Co, Yijie Xu, Shanyu Rong et al.CVPR 2026
- UFVideo: Towards Unified Fine-Grained Video Cooperative Understanding with Large Language ModelsHewen Pan, Cong Wei, Dashuang Liang, Zepeng Huang et al.CVPR 2026 · 3 citations
- Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware ReasoningSARA GHAZANFARI, Francesco Croce, Nicolas Flammarion, Prashanth Krishnamurthy et al.CVPR 2026 · 35 citations
- Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video ContentQiuheng Wang, Yukai Shi, Jiarong Ou, Rui Chen et al.CVPR 2025
