De-biased Natural Language Egocentric Task Verification via Prototypical Evidence Learning
Chong Liu, Xun Jiang, Fumin Shen, Lei Zhu, Jingkuan Song, Heng Tao Shen, Xing Xu
摘要
Natural Language-based Egocentric Task Verification (NLETV) aims to verify the alignment between action sequences in egocentric videos and their corresponding textual descriptions. However, existing NLETV approaches are still facing two critical challenges: (1) These methods are designed for simulating environments, ignoring the domain gap between synthetic and realistic data. (2) The matching processes are regarded as a simple binary classification problem, which undermines model reliability due to evaluation bias and uncalibrated decision settings. To address these challenges, we propose a novel method termed Prototypical Evidential Learning (PEL), which can be adapted to existing NLETV approaches and boost the model generalization and mitigate prediction bias. Our method leverages prototypes to guide cross-domain alignment and evidence collection. Specifically, PEL consists of two key components: (1) Prototypical Domain Adaptation module enabling cross-domain feature alignment and intra-domain prototype preservation between synthetic and realistic domains; (2) Matching Evidence Collector module, which quantifies prediction uncertainty on the prototypical representations through evidential deep learning. It enforces the model to collect the vision-text consistency and discrepancy evidence, thus addressing the issues of biased decisions in binary classification. Extensive experiments on two public datasets demonstrate that our PEL method outperforms existing state-of-the-art NLETV methods and shows remarkable generalizability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Evidential Deep Learning for Open Set Action RecognitionWentao Bao, Qi Yu, Yu KongICCV 2021 · 被引用 204 次
- Assembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural ActivitiesFadime Sener, Dibyadip Chatterjee, Daniel Shelepov, Kun He 等CVPR 2022 · 被引用 168 次
- GroundingGPT: Language Enhanced Multi-modal Grounding ModelZhaowei Li, Qi Xu, Dong Zhang, Hang Song 等ACL 2024 · 被引用 29 次
- SVIP: Sequence VerIfication for Procedures in VideosYicheng Qian, Weixin Luo, Dongze Lian, Xu Tang 等CVPR 2022 · 被引用 19 次
相关 Paper
- EgoTV: Egocentric Task Verification from Natural Language Task DescriptionsRishi Hazra, Brian Chen, Akshara Rai, Nitin Kamra 等ICCV 2023
- PHGC: Procedural Heterogeneous Graph Completion for Natural Language Task Verification in Egocentric VideosXun Jiang, Zhiyi Huang, Xing Xu, Jingkuan Song 等CVPR 2025
- Prototypical Action Reasoning Facilitated by Vision-Language Alignment for Egocentric Action AnticipationJiang Shao, Xinbo Zhao, Wenyin Tuo, Xiaochun ZouCVPR 2026
- Adaptive Evidential Learning for Temporal-Semantic Robustness in Moment RetrievalHaojian Huang, Kaijing Ma, Jin Chen, Haodong Chen 等AAAI 2026
- DCEL: Deep Cross-modal Evidential Learning for Text-Based Person RetrievalShenshen Li, Xing Xu, Yang Yang, Fumin Shen 等ACM MM 2023 · 被引用 56 次
