Generating Visual Scenes from Touch
Fengyu Yang, Jiacheng Zhang, Andrew Owens
Abstract
An emerging line of work has sought to generate plausible imagery from touch. Existing approaches, however, tackle only narrow aspects of the visuo-tactile synthesis problem, and lag significantly behind the quality of cross-modal synthesis methods in other domains. We draw on recent advances in latent diffusion to create a model for synthesizing images from tactile signals (and vice versa) and apply it to a number of visuo-tactile synthesis tasks. Using this model, we significantly outperform prior work on the tactile-driven stylization problem, i.e., manipulating an image to match a touch signal, and we are the first to successfully generate images from touch without additional sources of information about the scene. We also successfully use our model to address two novel synthesis problems: generating images that do not contain the touch sensor or the hand holding it, and estimating an image’s shading from its reflectance and touch. Project Page: https://fredfyyang.github.io/vision-from-touch/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b9e88a59-4fe1-4f62-b2fb-fee9e8a12396Cited by top-tier papers10
- Touch in the Wild: Learning Fine-Grained Manipulation with a Portable Visuo-Tactile GripperXinyue Zhu, Binghao Huang, Yunzhu LiNeurIPS 2025 · 62 citations
- Binding Touch to Everything: Learning Unified Multimodal Tactile RepresentationsFengyu Yang, Chao Feng, Ziyang Chen, Hyoungseob Park et al.CVPR 2024 · 47 citations
- WorDepth: Variational Language Prior for Monocular Depth EstimationZiyao Zeng, Daniel Wang, Fengyu Yang, Hyoungseob Park et al.CVPR 2024 · 20 citations
- TextToucher: Fine-Grained Text-to-Touch GenerationJiahang Tu, Hao Fu, Fengyu Yang, Hanbin Zhao et al.AAAI 2025 · 16 citations
- Inversion-Free Image Editing with Language-Guided Diffusion ModelsSihan Xu, Yidong Huang, Jiayi Pan, Ziqiao Ma et al.CVPR 2024 · 12 citations
Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- Touch2Shape: Touch-Conditioned 3D Diffusion for Shape Exploration and ReconstructionYuanbo Wang, Zhaoxuan Zhang, Jiajin Qiu, Dilong Sun et al.CVPR 2025
- Tactile-Augmented Radiance FieldsYiming Dou, Fengyu Yang, Yi Liu, Antonio Loquercio et al.CVPR 2024
- A Temporal and Content Co-Awareness Latent Diffusion for Controllable Hand Image GenerationShuang Hao, Pengfei Ren, Haifeng Sun, Pan Ting et al.CVPR 2026
- TouchDream: 3D Object Completion through Imagined TouchYuanbo Wang, Xinning Wang, Zhaoxuan Zhang, Changlong Wang et al.CVPR 2026
- Tactile DreamFusion: Exploiting Tactile Sensing for 3D GenerationRuihan Gao, Kangle Deng, Gengshan Yang, Wenzhen Yuan et al.NeurIPS 2024 · 13 citations
