VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection
Songhao Han, Wei Huang, Hairong Shi, Le Zhuo, Xiu Su, Shifeng Zhang, Xu Zhou, Xiaojuan Qi, Yue Liao, Si Liu
2025Year
35Top-tier citations
Abstract
Evidence: The <obj_start> man <obj_end> <box_start> [[575, 513, 544, 972]] ... with <obj_start> pink pills <obj_end> <box_start> [[355, 443, 33, 61]] <box_end> later in the images.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers35
- VideoRFT: Incentivizing Video Reasoning Capability in MLLMs via Reinforced Fine-TuningQi (Cheems) Wang, Yanrui Yu, Ye Yuan, Rui Mao et al.NeurIPS 2025 · 103 citations
- Scaling RL to Long VideosYukang Chen, Wei Huang, Baifeng Shi, Qinghao Hu et al.NeurIPS 2025 · 91 citations
- Open-o3-Video: Grounded Video Reasoning with Explicit Spatio-Temporal EvidenceJiahao Meng, Xiangtai Li, Haochen Wang, Tan Yue et al.ICML 2026 · 43 citations
- Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware ReasoningSARA GHAZANFARI, Francesco Croce, Nicolas Flammarion, Prashanth Krishnamurthy et al.CVPR 2026 · 35 citations
- VideoITG: Multimodal Video Understanding with Instructed Temporal GroundingShihao Wang, Guo Chen, De-An Huang, Zhiqi Li et al.CVPR 2026 · 35 citations
Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
Related papers
- Learning Asynchronous and Sparse Human-Object Interaction in VideosRomero F. A. B. de Morais, Vuong Le, Svetha Venkatesh, Truyen TranCVPR 2021
- ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion ModelsOzgur Kara, Krishna Kumar Singh, Feng Liu, Duygu Ceylan et al.CVPR 2025
- CelebV-Text: A Large-Scale Facial Text-Video DatasetJianhui Yu, Hao Zhu, Liming Jiang, Chen Change Loy et al.CVPR 2023
- FaceInsight: A Multimodal Large Language Model for Face PerceptionJingzhi Li, Changjiang Luo, Ruoyu Chen, Hua Zhang et al.ACM MM 2025 · 3 citations
- SimpleNet: A Simple Network for Image Anomaly Detection and LocalizationZhikang Liu, Yiming Zhou, Yuansheng Xu, Zilei WangCVPR 2023
