TextToucher: Fine-Grained Text-to-Touch Generation
Jiahang Tu, Hao Fu, Fengyu Yang, Hanbin Zhao, Chao Zhang, Hui Qian
摘要
Tactile sensation plays a crucial role in the development of multi-modal large models and embodied intelligence. To collect tactile data with minimal cost as possible, a series of studies have attempted to generate tactile images by vision-to-touch image translation. However, compared to text modality, visual modality-driven tactile generation cannot accurately depict human tactile sensation. In this work, we analyze the characteristics of tactile images in detail from two granularities: object-level (tactile texture, tactile shape), and sensor-level (gel status). We model these granularities of information through text descriptions and propose a fine-grained Text-to-Touch generation method (TextToucher) to generate high-quality tactile samples. Specifically, we introduce a multimodal large language model to build the text sentences about object-level tactile information and employ a set of learnable text prompts to represent the sensor-level tactile information. To better guide the tactile generation process with the built text information, we fuse the dual grains of text information and explore various dual-grain text conditioning methods within the diffusion transformer architecture. Furthermore, we propose a Contrastive Text-Touch Pre-training (CTTP) metric to precisely evaluate the quality of text-driven generated tactile data. Extensive experiments demonstrate the superiority of our TextToucher method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- How to Continually Adapt Text-to-Image Diffusion Models for Flexible Customization?Jiahua Dong, Wenqi Liang, Hongliu Li, Duzhen Zhang 等NeurIPS 2024 · 被引用 42 次
- BELM: Bidirectional Explicit Linear Multi-step Sampler for Exact Inversion in Diffusion ModelsFangyikang Wang, Hubery Yin, Yuejiang Dong, Huminhao Zhu 等NeurIPS 2024 · 被引用 37 次
- RSA: Resolving Scale Ambiguities in Monocular Depth Estimators through Language DescriptionsZiyao Zeng, Yangchao Wu, Hyoungseob Park, Daniel Wang 等NeurIPS 2024 · 被引用 26 次
- FG-OrIU: Towards Better Forgetting via Feature-Gradient Orthogonality for Incremental UnlearningQian Feng, Jiahang Tu, Mintong Kang, Hanbin Zhao 等ICCV 2025 · 被引用 9 次
- Unleashing High-Quality Image Generation in Diffusion Sampling Using Second-Order Levenberg-Marquardt-LangevinFangyikang Wang, Hubery Yin, Lei Qian, Yinan Li 等ICCV 2025
它引用的顶会 Paper19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
相关 Paper
- A Touch, Vision, and Language Dataset for Multimodal AlignmentLetian Fu, Gaurav Datta, Huang Huang, William Chung-Ho Panitch 等ICML 2024 · 被引用 89 次
- Universal Visuo-Tactile Video Understanding for Embodied InteractionYifan Xie, Mingyang Li, Shoujie Li, Xingting Li 等NeurIPS 2025 · 被引用 16 次
- Tactile DreamFusion: Exploiting Tactile Sensing for 3D GenerationRuihan Gao, Kangle Deng, Gengshan Yang, Wenzhen Yuan 等NeurIPS 2024 · 被引用 13 次
- DECO: Decoupled Multimodal Diffusion Transformer for Bimanual Dexterous Manipulation with a Plugin Tactile AdapterXukun Li, Yu Sun, Lei Zhang, Bo-Sheng Huang 等ICML 2026
- Touch2Shape: Touch-Conditioned 3D Diffusion for Shape Exploration and ReconstructionYuanbo Wang, Zhaoxuan Zhang, Jiajin Qiu, Dilong Sun 等CVPR 2025
