Learning to Segment Referred Objects from Narrated Egocentric Videos
Yuhan Shen, Huiyu Wang, Xitong Yang, Matt Feiszli, Ehsan Elhamifar, Lorenzo Torresani, Effrosyni Mavroudi
Abstract
Egocentric videos provide a first-person perspective of the wearer's activities, involving simultaneous interactions with multiple objects. In this work, we propose the task of weakly-supervised Narration-based Video Object Segmentation (NVOS). Given an egocentric video clip and a narration of the wearer's activities, our aim is to segment object instances mentioned in the narration, without using any spatial annotations during training. Existing weakly-supervised video object grounding methods typically yield bounding boxes for referred objects. In contrast, we propose ROSA, a weakly-supervised pixel-level grounding framework learning alignments between referred objects and segmentation mask proposals. Our model harnesses vision-language models pre-trained on image-text pairs to embed region masks and object phrases. During training, we combine (a) a video-narration contrastive loss that implicitly supervises the alignment between regions and phrases, and (b) a region-phrase contrastive loss based on inferred latent alignments. To address the lack of annotated NVOS datasets in egocentric videos, we create a new evaluation benchmark, VISOR-NVOS, leveraging existing annotations of segmentation masks from VISOR alongside 14.6k newly-collected, object-based video clip narrations. Our approach achieves state-of-the-art zero-shot pixel-level grounding performance compared to strong baselines under similar supervision. Additionally, we demonstrate generalization capabilities for zero-shot video object grounding on YouCook2, a third-person instructional video dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 70cc1a97-a4c7-4055-8c9b-6629d41fae26Cited by top-tier papers7
- Eyes Wide Open: Ego Proactive Video-LLM for Streaming VideoXueyang Yu, Cheng Shi, Yang Wang, Sibei YangNeurIPS 2025 · 34 citations
- MOSCATO: Predicting Multiple Object State Change through ActionsParnian Zameni, Yuhan Shen, Ehsan ElhamifarICCV 2025 · 4 citations
- Robust Egocentric Referring Video Object Segmentation via Dual-Modal Causal InterventionHaijing Liu, Zhiyuan Song, Hefeng Wu, Tao Pu et al.NeurIPS 2025 · 2 citations
- EgoLife: Towards Egocentric Life AssistantJingkang Yang, Shuai Liu, Hongming Guo, Yuhao Dong et al.CVPR 2025
- ANNEXE: Unified Analyzing, Answering, and Pixel Grounding for Egocentric InteractionYuejiao Su, Yi Wang, Qiongyang Hu, Chuang Yang et al.CVPR 2025
Builds on38
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or et al.ICCV 2021 · 1,437 citations
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
Related papers
- Look at What I'm Doing: Self-Supervised Spatial Grounding of Narrations in Instructional VideosReuben Tan, Bryan A. Plummer, Kate Saenko, Hailin Jin et al.NeurIPS 2021 · 30 citations
- Point-VOS: Pointing Up Video Object SegmentationSabarinath Mahadevan, Idil Esen Zulfikar, Paul Voigtlaender, Bastian LeibeCVPR 2024 · 3 citations
- SOC: Semantic-Assisted Object Cluster for Referring Video Object SegmentationZhuoyan Luo, Yicheng Xiao, Yong Liu, Shuyan Li et al.NeurIPS 2023 · 89 citations
- Contrastive Learning with Expectation-Maximization for Weakly Supervised Phrase GroundingKeqin Chen, Richong Zhang, Samuel Mensah, Yongyi MaoEMNLP 2022 · 3 citations
- Semantic and Sequential Alignment for Referring Video Object SegmentationFeiyu Pan, Hao Fang, Fangkai Li, Yanyu Xu et al.CVPR 2025
