De-biased Natural Language Egocentric Task Verification via Prototypical Evidence Learning
Chong Liu, Xun Jiang, Fumin Shen, Lei Zhu, Jingkuan Song, Heng Tao Shen, Xing Xu
Abstract
Natural Language-based Egocentric Task Verification (NLETV) aims to verify the alignment between action sequences in egocentric videos and their corresponding textual descriptions. However, existing NLETV approaches are still facing two critical challenges: (1) These methods are designed for simulating environments, ignoring the domain gap between synthetic and realistic data. (2) The matching processes are regarded as a simple binary classification problem, which undermines model reliability due to evaluation bias and uncalibrated decision settings. To address these challenges, we propose a novel method termed Prototypical Evidential Learning (PEL), which can be adapted to existing NLETV approaches and boost the model generalization and mitigate prediction bias. Our method leverages prototypes to guide cross-domain alignment and evidence collection. Specifically, PEL consists of two key components: (1) Prototypical Domain Adaptation module enabling cross-domain feature alignment and intra-domain prototype preservation between synthetic and realistic domains; (2) Matching Evidence Collector module, which quantifies prediction uncertainty on the prototypical representations through evidential deep learning. It enforces the model to collect the vision-text consistency and discrepancy evidence, thus addressing the issues of biased decisions in binary classification. Extensive experiments on two public datasets demonstrate that our PEL method outperforms existing state-of-the-art NLETV methods and shows remarkable generalizability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8f9cb97e-9d44-49ef-bc9e-f39844d0e3f1Builds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Evidential Deep Learning for Open Set Action RecognitionWentao Bao, Qi Yu, Yu KongICCV 2021 · 204 citations
- Assembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural ActivitiesFadime Sener, Dibyadip Chatterjee, Daniel Shelepov, Kun He et al.CVPR 2022 · 168 citations
- GroundingGPT: Language Enhanced Multi-modal Grounding ModelZhaowei Li, Qi Xu, Dong Zhang, Hang Song et al.ACL 2024 · 29 citations
- SVIP: Sequence VerIfication for Procedures in VideosYicheng Qian, Weixin Luo, Dongze Lian, Xu Tang et al.CVPR 2022 · 19 citations
Related papers
- EgoTV: Egocentric Task Verification from Natural Language Task DescriptionsRishi Hazra, Brian Chen, Akshara Rai, Nitin Kamra et al.ICCV 2023
- PHGC: Procedural Heterogeneous Graph Completion for Natural Language Task Verification in Egocentric VideosXun Jiang, Zhiyi Huang, Xing Xu, Jingkuan Song et al.CVPR 2025
- Prototypical Action Reasoning Facilitated by Vision-Language Alignment for Egocentric Action AnticipationJiang Shao, Xinbo Zhao, Wenyin Tuo, Xiaochun ZouCVPR 2026
- Adaptive Evidential Learning for Temporal-Semantic Robustness in Moment RetrievalHaojian Huang, Kaijing Ma, Jin Chen, Haodong Chen et al.AAAI 2026
- DCEL: Deep Cross-modal Evidential Learning for Text-Based Person RetrievalShenshen Li, Xing Xu, Yang Yang, Fumin Shen et al.ACM MM 2023 · 56 citations
