Learning to Jointly Understand Visual and Tactile Signals
Yichen Li, Yilun Du, Chao Liu, Chao Liu, Francis Williams, Michael Foshey, Benjamin Eckart, Jan Kautz, Joshua B. Tenenbaum, Antonio Torralba, Wojciech Matusik
摘要
Modeling and analyzing objects and shapes has been well-studied in the past. However, manipulation of these complex tools and articulated objects remains difficult for autonomous agents. Our human hands, however, are dexterous and adaptive. We can easily adapt a manipulation skill on one object to all objects in the class and to other similar classes. Our intuition comes from that there is a close connection between manipulations and topology and articulation of objects. The possible articulation of objects indicates the types of manipulation necessary to operate the object. In this work, we aim to take a manipulation perspective to understand everyday objects and tools. We collect a multi-modal visual-tactile dataset that contains paired full-hand force pressure maps and manipulation videos. We also propose a novel method to learn a cross-modal latent manifold that allows for cross-modal prediction and discovery of latent structure in different data modalities. We conduct extensive experiments to demonstrate the effectiveness of our method. ‡
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Revealing Vision-Language Integration in the Brain with Multimodal NetworksVighnesh Subramaniam, Colin Conwell, Christopher Wang, Gabriel Kreiman 等ICML 2024 · 被引用 19 次
- MultiModal Action Conditioned Video SimulationYichen Li, Antonio TorralbaICCV 2025 · 被引用 2 次
它引用的顶会 Paper13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi 等ICCV 2019 · 被引用 1,437 次
- Neural Articulated Radiance FieldAtsuhiro Noguchi, Xiao Sun, Stephen Lin, Tatsuya HaradaICCV 2021 · 被引用 242 次
- Where2Act: From Pixels to Actions for Articulated 3D ObjectsKaichun Mo, Leonidas J. Guibas, Mustafa Mukadam, Abhinav Gupta 等ICCV 2021 · 被引用 240 次
相关 Paper
- TACO: Benchmarking Generalizable Bimanual Tool-ACtion-Object UnderstandingYun Liu, Haolin Yang, Xu Si, Ling Liu 等CVPR 2024 · 被引用 12 次
- VTDexManip: A Dataset and Benchmark for Visual-tactile Pretraining and Dexterous Manipulation with Reinforcement LearningQingtao Liu, Yu Cui, Zhengnan Sun, Gaofeng Li 等ICLR 2025
- AdaManip: Adaptive Articulated Object Manipulation Environments and Policy LearningYuanfei Wang, Xiaojie Zhang, Ruihai Wu, Yu Li 等ICLR 2025
- Learning Object-Centric Motion Priors from Human for Robotic Dexterous ManipulationZhengdong Hong, Guofeng ZhangAAAI 2026
- Cross-Hand Latent Representation for Vision-Language-Action ModelsGuangqi Jiang, Yutong Liang, Jianglong Ye, Jia-Yang Huang 等CVPR 2026 · 被引用 14 次
