Artemis: Towards Referential Understanding in Complex Videos
Jihao Qiu, Yuan Zhang, Xi Tang, Lingxi Xie, Tianren Ma, Pengyu Yan, David S. Doermann, Qixiang Ye, Yunjie Tian
Abstract
Videos carry rich visual information including object description, action, interaction, etc., but the existing multimodal large language models (MLLMs) fell short in referential understanding scenarios such as video-based referring. In this paper, we present Artemis, an MLLM that pushes video-based referential understanding to a finer level. Given a video, Artemis receives a natural-language question with a bounding box in any video frame and describes the referred target in the entire video. The key to achieving this goal lies in extracting compact, target-specific video features, where we set a solid baseline by tracking and selecting spatiotemporal features from the video. We train Artemis on the newly established VideoRef45K dataset with 45K video-QA pairs and design a computationally efficient, three-stage training procedure. Results are promising both quantitatively and qualitatively. Additionally, we show that can be integrated with video grounding and text summarization tools to understand more complex scenarios. Code and data are available at https://github.com/qiujihao19/Artemis.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c5043d74-9e52-420a-a5a7-6ba3d3777edfCited by top-tier papers15
- Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and VideosWeifeng Lin, Xinyu Wei, Ruichuan An, Tianhe Ren et al.NeurIPS 2025 · 47 citations
- UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual ReasoningYe Liu, Zongyang Ma, Junfu Pu, Zhongang Qi et al.NeurIPS 2025 · 39 citations
- 3D Aware Region Prompted Vision Language ModelAn-Chieh Cheng, Yang Fu, Yukang Chen, Zhijian Liu et al.ICLR 2026 · 30 citations
- Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMsHaochen Wang, Yuhao Wang, Tao Zhang, Yikang Zhou et al.ICLR 2026 · 18 citations
- Describe Anything: Detailed Localized Image and Video CaptioningLong Lian, Yifan Ding, Yunhao Ge, Sifei Liu et al.ICCV 2025 · 14 citations
Builds on23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- Object-Centric Video Question Answering with Visual Grounding and ReferringHaochen Wang, Qirui Chen, Cilin Yan, Jiayin Cai et al.ICCV 2025 · 2 citations
- VideoOrion: Tokenizing Object Dynamics in VideosYicheng Feng, Yijiang Li, Wanpeng Zhang, Sipeng Zheng et al.ICCV 2025
- VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLMYuqian Yuan, Hang Zhang, Wentong Li, Zesen Cheng et al.CVPR 2025
- SPARROW: Learning Spatial Precision and Temporal Referential Consistency in Pixel-Grounded Video MLLMsMohamad Alansari, Naufal Suryanto, Divya Velayudhan, Sajid Javed et al.CVPR 2026 · 2 citations
- VP-Bench: A Comprehensive Benchmark for Visual Prompting in Multimodal Large Language ModelsMingjie Xu, Jinpeng Chen, Yuzhi Zhao, Jason Chun Lok Li et al.AAAI 2026
