Learning Robust Representations with Information Bottleneck and Memory Network for RGB-D-based Gesture Recognition
Yunan Li, Huizhou Chen, Guanwen Feng, Qiguang Miao
摘要
Although previous RGB-D-based gesture recognition methods have shown promising performance, researchers often overlook the interference of task-irrelevant cues like illumination and background. These unnecessary factors are learned together with the predictive ones by the network and hinder accurate recognition. In this paper, we propose a convenient and analytical framework to learn a robust feature representation that is impervious to gesture-irrelevant factors. Based on the Information Bottleneck theory, two rules of Sufficiency and Compactness are derived to develop a new information-theoretic loss function, which cultivates a more sufficient and compact representation from the feature encoding and mitigates the impact of gesture-irrelevant information. To highlight the predictive information, we further integrate a memory network. Using our proposed content-based and contextual memory addressing scheme, we weaken the nuisances while preserving the task-relevant information, providing guidance for refining the feature representation. Experiments conducted on three public datasets demonstrate that our approach leads to a better feature representation and achieves better performance than state-of-the-art methods. The code of our method is available at: https://github.com/Carpumpkin/InBoMem.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper10
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- Learning Robust Representations via Multi-View Information BottleneckMarco Federici, Anjan Dutta, Patrick Forré, Nate Kushman 等ICLR 2020 · 被引用 330 次
- Gloss-free Sign Language Translation: Improving from Visual-Language PretrainingBenjia Zhou, Zhigang Chen, Albert Clapés, Jun Wan 等ICCV 2023 · 被引用 123 次
- Video-based Person Re-identification with Spatial and Temporal Memory NetworksChanho Eom, Geon Lee, Junghyup Lee, Bumsub HamICCV 2021 · 被引用 107 次
- Temporal Feature Alignment and Mutual Information Maximization for Video-Based Human Pose EstimationZhenguang Liu, Runyang Feng, Haoming Chen, Shuang Wu 等CVPR 2022 · 被引用 76 次
相关 Paper
- Drop-Bottleneck: Learning Discrete Compressed Representation for Noise-Robust ExplorationJaekyeom Kim, Minjung Kim, Dongyeon Woo, Gunhee KimICLR 2021 · 被引用 20 次
- Information Retention via Learning Supplemental FeaturesZhipeng Xie, Yahe LiICLR 2024 · 被引用 1 次
- Bridging and Compressing Feature and Semantic Spaces for Robust Graph Neural Networks: An Information Theory PerspectiveLuying Zhong, Renjie Lin, Jiayin Li, Shiping Wang 等KDD 2024 · 被引用 3 次
- Representation Learning with Conditional Information Flow MaximizationDou Hu, Lingwei Wei, Wei Zhou, Songlin HuACL 2024
- DTR: An Information Bottleneck Based Regularization Framework for Video Action RecognitionJiawei Fan, Yu Zhao, Xie Yu, Lihua Ma 等ACM MM 2022 · 被引用 3 次
