Seeing Through Touch: Tactile-Driven Visual Localization of Material Regions
Seongyu Kim, Seungwoo Lee, Hyeonggon Ryu, Joon Chung, Arda Senocak
摘要
We address the problem of tactile localization, where the goal is to identify image regions that share the same material properties as a tactile input. Existing visuo-tactile methods rely on global alignment and thus fail to capture the fine-grained local correspondences required for this task. The challenge is amplified by existing datasets, which predominantly contain close-up, low-diversity images. We propose a model that learns local visuo-tactile alignment via dense cross-modal feature interactions, producing tactile saliency maps for touch-conditioned material segmentation. To overcome dataset constraints, we introduce: (i) in-the-wild multi-material scene images that expand visual diversity, and (ii) a material-diversity pairing strategy that aligns each tactile sample with visually varied yet tactilely consistent images, improving contextual localization and robustness to weak signals. We also construct two new tactile-grounded material segmentation datasets for quantitative evaluation. Experiments on both new and existing benchmarks show that our approach substantially outperforms prior visuo-tactile methods in tactile localization. Project page: https://mm.kaist.ac. kr/projects/SeeingThroughTouch/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Image Segmentation Using Text and Image PromptsTimo Lüddecke, Alexander S. EckerCVPR 2022 · 被引用 457 次
- A Touch, Vision, and Language Dataset for Multimodal AlignmentLetian Fu, Gaurav Datta, Huang Huang, William Chung-Ho Panitch 等ICML 2024 · 被引用 89 次
- ObjectFolder 2.0: A Multisensory Object Dataset for Sim2Real TransferRuohan Gao, Zilin Si, Yen-Yu Chang, Samuel Clarke 等CVPR 2022 · 被引用 58 次
相关 Paper
- RA-Touch: Retrieval-Augmented Touch Understanding with Enriched Visual DataYoorhim Cho, Hongyeob Kim, Semin Kim, Youjia Zhang 等ACM MM 2025
- Cross-Tactile Sensor Representation LearningYan Zhang, Zheng WANG, Pengpeng Zeng, Xing Xu 等ICML 2026
- Augmenting Imagery with Multimodal Vibrotactile Representations: Touch, Feel, and HearMazen Salous, Matthias Kramer, Wilko Heuten, Charles Hudin 等CHI 2026 · 被引用 3 次
- RobustVisH: Robust Visual-Haptic Cross-Modal Recognition under Transmission InterferenceRouqi Zhang, Chengdi Lu, Hancheng Lu, Yang Cao 等ACM MM 2025 · 被引用 2 次
- Universal Visuo-Tactile Video Understanding for Embodied InteractionYifan Xie, Mingyang Li, Shoujie Li, Xingting Li 等NeurIPS 2025 · 被引用 16 次
