Knowledge As Priors: Cross-Modal Knowledge Generalization for Datasets Without Superior Knowledge
Long Zhao, Xi Peng, Yuxiao Chen, Mubbasir Kapadia, Dimitris N. Metaxas
Abstract
Cross-modal knowledge distillation deals with transferring knowledge from a model trained with superior modalities (Teacher) to another model trained with weak modalities (Student). Existing approaches require paired training examples exist in both modalities. However, accessing the data from superior modalities may not always be feasible. For example, in the case of 3D hand pose estimation, depth maps, point clouds, or stereo images usually capture better hand structures than RGB images, but most of them are expensive to be collected. In this paper, we propose a novel scheme to train the Student in a Target dataset where the Teacher is unavailable. Our key idea is to generalize the distilled cross-modal knowledge learned from a Source dataset, which contains paired examples from both modalities, to the Target dataset by modeling knowledge as priors on parameters of the Student. We name our method "Cross-Modal Knowledge Generalization" and demonstrate that our scheme results in competitive performance for 3D hand pose estimation on standard benchmark datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers21
- SMIL: Multimodal Learning with Severely Missing ModalityMengmeng Ma, Jian Ren, Long Zhao, Sergey Tulyakov et al.AAAI 2021 · 393 citations
- Maximum-Entropy Adversarial Data Augmentation for Improved Generalization and RobustnessLong Zhao, Ting Liu, Xi Peng, Dimitris N. MetaxasNeurIPS 2020 · 207 citations
- LIGA-Stereo: Learning LiDAR Geometry Aware Representations for Stereo-based 3D DetectorXiaoyang Guo, Shaoshuai Shi, Xiaogang Wang, Hongsheng LiICCV 2021 · 132 citations
- Towards Accurate Alignment in Real-time 3D Hand-Mesh ReconstructionXiao Tang, Tianyu Wang, Chi-Wing FuICCV 2021 · 83 citations
- ToAlign: Task-Oriented Alignment for Unsupervised Domain AdaptationGuoqiang Wei, Cuiling Lan, Wenjun Zeng, Zhizheng Zhang et al.NeurIPS 2021 · 80 citations
Builds on5
- FreiHAND: A Dataset for Markerless Capture of Hand Pose and Shape From Single RGB ImagesChristian Zimmermann, Duygu Ceylan, Jimei Yang, Bryan C. Russell et al.ICCV 2019 · 493 citations
- End-to-End Hand Mesh Recovery From a Monocular RGB ImageXiong Zhang, Qiang Li, Hong Mo, Wenbo Zhang et al.ICCV 2019 · 248 citations
- Aligning Latent Spaces for 3D Hand Pose EstimationLinlin Yang, Shile Li, Dongheui Lee, Angela YaoICCV 2019 · 93 citations
- Distill Knowledge From NRSfM for Weakly Supervised 3D Pose LearningChaoyang Wang, Chen Kong, Simon LuceyICCV 2019 · 52 citations
- Learning to Learn Single Domain GeneralizationFengchun Qiao, Long Zhao, Xi PengCVPR 2020
Related papers
- Cross-Modal Knowledge Distillation without Paired Data: Theoretical Foundation and AlgorithmT. K Tran, Duc Chu Anh, Quang Hung Pham, Phi Le Nguyen et al.ICML 2026
- Cross-Domain 3D Hand Pose Estimation with Dual ModalitiesQiuxia Lin, Linlin Yang, Angela YaoCVPR 2023
- Learning Scene Structure Guidance via Cross-Task Knowledge Transfer for Single Depth Super-ResolutionBaoli Sun, Xinchen Ye, Baopu Li, Haojie Li et al.CVPR 2021
- Cross-Domain and Cross-Modal Knowledge Distillation in Domain Adaptation for 3D Semantic SegmentationMiaoyu Li, Yachao Zhang, Yuan Xie, Zuodong Gao et al.ACM MM 2022 · 30 citations
- C2KD: Bridging the Modality Gap for Cross-Modal Knowledge DistillationFushuo Huo, Wenchao Xu, Jingcai Guo, Haozhao Wang et al.CVPR 2024
