AnyTouch: Learning Unified Static-Dynamic Representation across Multiple Visuo-tactile Sensors
Ruoxuan Feng, Jiangyu Hu, Wenke Xia, Tianci Gao, Ao Shen, Yuhao Sun, Bin Fang, Di Hu
Abstract
Visuo-tactile sensors aim to emulate human tactile perception, enabling robots to precisely understand and manipulate objects. Over time, numerous meticulously designed visuo-tactile sensors have been integrated into robotic systems, aiding in completing various tasks. However, the distinct data characteristics of these low-standardized visuo-tactile sensors hinder the establishment of a powerful tactile perception system. We consider that the key to addressing this issue lies in learning unified multi-sensor representations, thereby integrating the sensors and promoting tactile knowledge transfer between them. To achieve unified representation of this nature, we introduce TacQuad, an aligned multi-modal multi-sensor tactile dataset from four different visuo-tactile sensors, which enables the explicit integration of various sensors. Recognizing that humans perceive the physical environment by acquiring diverse tactile information such as texture and pressure changes, we further propose to learn unified multi-sensor representations from both static and dynamic perspectives. By integrating tactile images and videos, we present AnyTouch, a unified static-dynamic multi-sensor representation learning framework with a multi-level structure, aimed at both enhancing comprehensive perceptual abilities and enabling effective cross-sensor transfer. This multi-level architecture captures pixel-level details from tactile data via masked modeling and enhances perception and transferability by learning semantic-level sensor-agnostic features through multi-modal alignment and cross-sensor matching. We provide a comprehensive analysis of multi-sensor transferability, and validate our method on various offline datasets and in the real-world pouring task. Experimental results show that our method outperforms existing methods, exhibits outstanding static and dynamic perception capabilities across various sensors. The code, TacQuad dataset and AnyTouch model are fully available at gewu-lab.github.io/AnyTouch/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c7a011e6-8136-4e84-9b5b-7122093fe3ccCited by top-tier papers8
- Touch in the Wild: Learning Fine-Grained Manipulation with a Portable Visuo-Tactile GripperXinyue Zhu, Binghao Huang, Yunzhu LiNeurIPS 2025 · 62 citations
- AnyTouch 2: General Optical Tactile Representation Learning For Dynamic Tactile PerceptionRuoxuan Feng, Yuxuan Zhou, Siyu Mei, Dongzhan Zhou et al.ICLR 2026 · 25 citations
- Universal Visuo-Tactile Video Understanding for Embodied InteractionYifan Xie, Mingyang Li, Shoujie Li, Xingting Li et al.NeurIPS 2025 · 16 citations
- TaCo: A Benchmark for Lossless and Lossy Codecs of Heterogeneous Tactile DataZhengxue Cheng, Yan Zhao, Keyu Wang, Hengdi Zhang et al.ICLR 2026 · 3 citations
- Toward Artificial Palpation: Representation Learning of Touch on Soft BodiesZohar Rimon, Elisei Shafer, Tal Tepper, Efrat Shimron et al.NeurIPS 2025 · 1 citation
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 2,336 citations
- ViTGAN: Training GANs with Vision TransformersKwonjoon Lee, Huiwen Chang, Lu Jiang, Han Zhang et al.ICLR 2022 · 225 citations
- Omnivore: A Single Model for Many Visual ModalitiesRohit Girdhar, Mannat Singh, Nikhila Ravi, Laurens van der Maaten et al.CVPR 2022 · 185 citations
Related papers
- Cross-Tactile Sensor Representation LearningYan Zhang, Zheng WANG, Pengpeng Zeng, Xing Xu et al.ICML 2026
- Collaborative Representation Learning for Alignment of Tactile, Language, and Vision ModalitiesYiyun Zhou, Mingjing Xu, Jingwei Shi, Quanjiang Li et al.AAAI 2026 · 1 citation
- Binding Touch to Everything: Learning Unified Multimodal Tactile RepresentationsFengyu Yang, Chao Feng, Ziyang Chen, Hyoungseob Park et al.CVPR 2024 · 47 citations
- VTDexManip: A Dataset and Benchmark for Visual-tactile Pretraining and Dexterous Manipulation with Reinforcement LearningQingtao Liu, Yu Cui, Zhengnan Sun, Gaofeng Li et al.ICLR 2025
- X-Capture: An Open-Source Portable Device for Multi-Sensory LearningSamuel Clarke, Suzannah Wistreich, Yanjie Ze, Jiajun WuICCV 2025 · 1 citation
