Seeing Through Touch: Tactile-Driven Visual Localization of Material Regions
Seongyu Kim, Seungwoo Lee, Hyeonggon Ryu, Joon Chung, Arda Senocak
Abstract
We address the problem of tactile localization, where the goal is to identify image regions that share the same material properties as a tactile input. Existing visuo-tactile methods rely on global alignment and thus fail to capture the fine-grained local correspondences required for this task. The challenge is amplified by existing datasets, which predominantly contain close-up, low-diversity images. We propose a model that learns local visuo-tactile alignment via dense cross-modal feature interactions, producing tactile saliency maps for touch-conditioned material segmentation. To overcome dataset constraints, we introduce: (i) in-the-wild multi-material scene images that expand visual diversity, and (ii) a material-diversity pairing strategy that aligns each tactile sample with visually varied yet tactilely consistent images, improving contextual localization and robustness to weak signals. We also construct two new tactile-grounded material segmentation datasets for quantitative evaluation. Experiments on both new and existing benchmarks show that our approach substantially outperforms prior visuo-tactile methods in tactile localization. Project page: https://mm.kaist.ac. kr/projects/SeeingThroughTouch/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 40b0fb1c-407b-4954-9208-bec9e30cc77eBuilds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Image Segmentation Using Text and Image PromptsTimo Lüddecke, Alexander S. EckerCVPR 2022 · 457 citations
- A Touch, Vision, and Language Dataset for Multimodal AlignmentLetian Fu, Gaurav Datta, Huang Huang, William Chung-Ho Panitch et al.ICML 2024 · 89 citations
- ObjectFolder 2.0: A Multisensory Object Dataset for Sim2Real TransferRuohan Gao, Zilin Si, Yen-Yu Chang, Samuel Clarke et al.CVPR 2022 · 58 citations
Related papers
- RA-Touch: Retrieval-Augmented Touch Understanding with Enriched Visual DataYoorhim Cho, Hongyeob Kim, Semin Kim, Youjia Zhang et al.ACM MM 2025
- Cross-Tactile Sensor Representation LearningYan Zhang, Zheng WANG, Pengpeng Zeng, Xing Xu et al.ICML 2026
- Augmenting Imagery with Multimodal Vibrotactile Representations: Touch, Feel, and HearMazen Salous, Matthias Kramer, Wilko Heuten, Charles Hudin et al.CHI 2026 · 3 citations
- RobustVisH: Robust Visual-Haptic Cross-Modal Recognition under Transmission InterferenceRouqi Zhang, Chengdi Lu, Hancheng Lu, Yang Cao et al.ACM MM 2025 · 2 citations
- Universal Visuo-Tactile Video Understanding for Embodied InteractionYifan Xie, Mingyang Li, Shoujie Li, Xingting Li et al.NeurIPS 2025 · 16 citations
