RobustVisH: Robust Visual-Haptic Cross-Modal Recognition under Transmission Interference
Rouqi Zhang, Chengdi Lu, Hancheng Lu, Yang Cao, Tiesong Zhao
Abstract
Embodied AI calls for a reliable, cross-modal object recognition that deeply mines High-Quality (HQ) object appearance (i.e., visual information) and touch details (i.e., haptic information). While in real-world scenarios, cross-modal data is usually degraded due to data acquisition and delivery in complex environments. In this paper, we propose a Robust Visual-Haptic recognition (RobustVisH) model that identifies Low-Quality (LQ) visual-haptic data with transmission distortion for the first time. First, we introduce the WIreless Transmission Interference-based Multi-modal benchmark (WITIM) as a visual-haptic dataset under transmission interference. In particular, the dataset consists of WITIM/AU and WITIM/PHAC-2, in which the original signals are obtained from AU and PHAC-2, respectively. Second, we design a trainable weighted fusion and a Transformer encoder based on the bi-directional self-attention mechanism, enabling RobustVisH to form and learn fused visual-haptic features after modality-specific one-dimensional feature encoding. Third, we employ a covariate shift paradigm, transferring knowledge of RobustVisH from HQ data to LQ data, thereby increasing its robustness against transmission-interference inputs. Experimental results demonstrate that the proposed RobustVisH improves the accuracy of the state-of-the-art method by 2.06% and 9.28% on WITIM/AU and WITIM/PHAC-2, respectively. Source code is available at: https://github.com/lylibylily/RobustVisH.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get ec9d78a8-b5e4-4ff6-ad0b-e2c0063b7919Related papers
- DECO: Decoupled Multimodal Diffusion Transformer for Bimanual Dexterous Manipulation with a Plugin Tactile AdapterXukun Li, Yu Sun, Lei Zhang, Bo-Sheng Huang et al.ICML 2026
- 3D Shape Reconstruction from Vision and TouchEdward J. Smith, Roberto Calandra, Adriana Romero, Georgia Gkioxari et al.NeurIPS 2020 · 90 citations
- TouchFormer: A Robust Transformer-based Framework for Multimodal Material PerceptionKailin Lyu, Long Xiao, Jianing Zeng, Junhao Dong et al.AAAI 2026
- Efficient Semantic Codec for Real-time Vibrotactile TransmissionRunjie Wang, Kemi Chen, Shuijie Li, Mingkai Chen et al.ACM MM 2025
- Seeing Through Touch: Tactile-Driven Visual Localization of Material RegionsSeongyu Kim, Seungwoo Lee, Hyeonggon Ryu, Joon Chung et al.CVPR 2026 · 2 citations
