CREST: Cross-modal Resonance through Evidential Deep Learning for Enhanced Zero-Shot Learning
Haojian Huang, Xiaozhen Qiao, Zhuo Chen, Haodong Chen, Bingyu Li, Zhe Sun, Mulin Chen, Xuelong Li
摘要
Zero-shot learning (ZSL) enables the recognition of novel classes by leveraging semantic knowledge transfer from known to unknown categories. This knowledge, typically encapsulated in attribute descriptions, aids in identifying class-specific visual features, thus facilitating visual-semantic alignment and improving ZSL performance. However, real-world challenges such as distribution imbalances and attribute co-occurrence among instances often hinder the discernment of local variances in images, a problem exacerbated by the scarcity of fine-grained, region-specific attribute annotations. Moreover, the variability in visual presentation within categories can also skew attribute-category associations. In response, we propose a bidirectional cross-modal ZSL approach CREST. It begins by extracting representations for attribute and visual localization and employs Evidential Deep Learning (EDL) to measure underlying epistemic uncertainty, thereby enhancing the model's resilience against hard negatives. CREST incorporates dual learning pathways, focusing on both visual-category and attribute-category alignments, to ensure robust correlation between latent and observable spaces. Moreover, we introduce an uncertainty-informed cross-modal fusion technique to refine visual-attribute inference. Extensive experiments demonstrate our model's effectiveness and unique explainability across multiple datasets. Our code and data are available at: https://github.com/JethroJames/CREST
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- FineCLIPER: Multi-modal Fine-grained CLIP for Dynamic Facial Expression Recognition with AdaptERsHaodong Chen, Haojian Huang, Junhao Dong, Mingzhe Zheng 等ACM MM 2024 · 被引用 26 次
- Trusted Unified Feature-Neighborhood Dynamics for Multi-View ClassificationHaojian Huang, Chuanyu Qin, Zhe Liu, Kaijing Ma 等AAAI 2025 · 被引用 25 次
- SeFAR: Semi-supervised Fine-grained Action Recognition with Temporal Perturbation and Learning StabilizationYongle Huang, Haodong Chen, Zhenbang Xu, Zihan Jia 等AAAI 2025 · 被引用 13 次
- EvoEmpirBench: Dynamic Spatial Reasoning with Agent-ExpVerPukun Zhao, Longxiang Wang, Miaowei Wang, Chen Chen 等AAAI 2026 · 被引用 2 次
- Find, Fix, Reason: Context Repair for Video ReasoningHaojian Huang, Chuanyu Qin, Yinchuan Li, YINGCONG CHENICML 2026 · 被引用 1 次
它引用的顶会 Paper32
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied AgentsWenlong Huang, Pieter Abbeel, Deepak Pathak, Igor MordatchICML 2022 · 被引用 1,539 次
- VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text UnderstandingHu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko 等EMNLP 2021 · 被引用 399 次
- Attribute Prototype Network for Zero-Shot LearningWenjia Xu, Yongqin Xian, Jiuniu Wang, Bernt Schiele 等NeurIPS 2020 · 被引用 392 次
相关 Paper
- TransZero: Attribute-Guided Transformer for Zero-Shot LearningShiming Chen, Ziming Hong, Yang Liu, Guo-Sen Xie 等AAAI 2022 · 被引用 185 次
- Goal-Oriented Gaze Estimation for Zero-Shot LearningYang Liu, Lei Zhou, Xiao Bai, Yifei Huang 等CVPR 2021
- TOMCAT: Test-time Comprehensive Knowledge Accumulation for Compositional Zero-Shot LearningXudong Yan, Songhe FengNeurIPS 2025
- DUET: Cross-Modal Semantic Grounding for Contrastive Zero-Shot LearningZhuo Chen, Yufeng Huang, Jiaoyan Chen, Yuxia Geng 等AAAI 2023 · 被引用 97 次
- VGSE: Visually-Grounded Semantic Embeddings for Zero-Shot LearningWenjia Xu, Yongqin Xian, Jiuniu Wang, Bernt Schiele 等CVPR 2022 · 被引用 61 次
