EgoTV: Egocentric Task Verification from Natural Language Task Descriptions
Rishi Hazra, Brian Chen, Akshara Rai, Nitin Kamra, Ruta Desai
Abstract
Natural Language-based Egocentric Task Verification (NLETV) aims to verify the alignment between action sequences in egocentric videos and their corresponding textual descriptions. However, existing NLETV approaches are still facing two critical challenges: (1) These methods are designed for simulating environments, ignoring the domain gap between synthetic and realistic data. (2) The matching processes are regarded as a simple binary classification problem, which undermines model reliability due to evaluation bias and uncalibrated decision settings. To address these challenges, we propose a novel method termed Prototypical Evidential Learning (PEL), which can be adapted to existing NLETV approaches and boost the model generalization and mitigate prediction bias. Our method leverages prototypes to guide cross-domain alignment and evidence collection. Specifically, PEL consists of two key components: (1) Prototypical Domain Adaptation module enabling cross-domain feature alignment and intra-domain prototype preservation between synthetic and realistic domains; (2) Matching Evidence Collector module, which quantifies prediction uncertainty on the prototypical representations through evidential deep learning. It enforces the model to collect the vision-text consistency and discrepancy evidence, thus addressing the issues of biased decisions in binary classification. Extensive experiments on two public datasets demonstrate that our PEL method outperforms existing state-of-the-art NLETV methods and shows remarkable generalizability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e1227dde-3ddc-4a33-b061-3dcc2139b1edCited by top-tier papers5
- EAGLE: Egocentric AGgregated Language-video EngineJing Bi, Yunlong Tang, Luchuan Song, Ali Vosoughi et al.ACM MM 2024 · 3 citations
- De-biased Natural Language Egocentric Task Verification via Prototypical Evidence LearningChong Liu, Xun Jiang, Fumin Shen, Lei Zhu et al.AAAI 2026
- SVLTA: Benchmarking Vision-Language Temporal Alignment via Synthetic Video SituationHao Du, Bo Wu, Yan Lu, Zhendong MaoCVPR 2025
- PHGC: Procedural Heterogeneous Graph Completion for Natural Language Task Verification in Egocentric VideosXun Jiang, Zhiyi Huang, Xing Xu, Jingkuan Song et al.CVPR 2025
- REvolve: Reward Evolution with Large Language Models using Human FeedbackRishi Hazra, Alkis Sygkounas, Andreas Persson, Amy Loutfi et al.ICLR 2025
Builds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- CLEVRER: Collision Events for Video Representation and ReasoningKexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli et al.ICLR 2020 · 584 citations
- RegionCLIP: Region-based Language-Image PretrainingYiwu Zhong, Jianwei Yang, Pengchuan Zhang, Chunyuan Li et al.CVPR 2022 · 481 citations
Related papers
- Prototypical Action Reasoning Facilitated by Vision-Language Alignment for Egocentric Action AnticipationJiang Shao, Xinbo Zhao, Wenyin Tuo, Xiaochun ZouCVPR 2026
- Adaptive Evidential Learning for Temporal-Semantic Robustness in Moment RetrievalHaojian Huang, Kaijing Ma, Jin Chen, Haodong Chen et al.AAAI 2026
- Interactive Prototype Learning for Egocentric Action RecognitionXiaohan Wang, Linchao Zhu, Heng Wang, Yi YangICCV 2021 · 78 citations
- DCEL: Deep Cross-modal Evidential Learning for Text-Based Person RetrievalShenshen Li, Xing Xu, Yang Yang, Fumin Shen et al.ACM MM 2023 · 56 citations
- OPEL: Optimal Transport Guided ProcedurE LearningSayeed Shafayet Chowdhury, Soumyadeep Chandra, Kaushik RoyNeurIPS 2024 · 11 citations
