ObjectFolder 2.0: A Multisensory Object Dataset for Sim2Real Transfer
Ruohan Gao, Zilin Si, Yen-Yu Chang, Samuel Clarke, Jeannette Bohg, Li Fei-Fei, Wenzhen Yuan, Jiajun Wu
摘要
Objects play a crucial role in our everyday activities. Though multisensory object-centric learning has shown great potential lately, the modeling of objects in prior work is rather unrealistic. OBJECTFOLDER 1.0 is a recent dataset that introduces 100 virtualized objects with visual, acoustic, and tactile sensory data. However, the dataset is small in scale and the multisensory data is of limited quality, hampering generalization to real-world scenarios. We present OBJECTFOLDER 2.0, a large-scale, multisensory dataset of common household objects in the form of implicit neural representations that significantly enhances OBJECTFOLDER 1.0 in three aspects. First, our dataset is 10 times larger in the amount of objects and orders of magnitude faster in rendering time. Second, we significantly improve the multisensory rendering quality for all three modalities. Third, we show that models learned from virtual objects in our dataset successfully transfer to their real-world counterparts in three challenging tasks: object scale estimation, contact localization, and shape reconstruction. OBJECTFOLDER 2.0 offers a new path and testbed for multisensory learning in computer vision and robotics. The dataset is available at https://github . com/rhgao/ObjectFolder.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper34
- A Touch, Vision, and Language Dataset for Multimodal AlignmentLetian Fu, Gaurav Datta, Huang Huang, William Chung-Ho Panitch 等ICML 2024 · 被引用 89 次
- Touch in the Wild: Learning Fine-Grained Manipulation with a Portable Visuo-Tactile GripperXinyue Zhu, Binghao Huang, Yunzhu LiNeurIPS 2025 · 被引用 62 次
- Binding Touch to Everything: Learning Unified Multimodal Tactile RepresentationsFengyu Yang, Chao Feng, Ziyang Chen, Hyoungseob Park 等CVPR 2024 · 被引用 47 次
- Generating Visual Scenes from TouchFengyu Yang, Jiacheng Zhang, Andrew OwensICCV 2023 · 被引用 39 次
- AnyTouch 2: General Optical Tactile Representation Learning For Dynamic Tactile PerceptionRuoxuan Feng, Yuxuan Zhou, Siyu Mei, Dongzhan Zhou 等ICLR 2026 · 被引用 25 次
它引用的顶会 Paper14
- Implicit Neural Representations with Periodic Activation FunctionsVincent Sitzmann, Julien N. P. Martel, Alexander W. Bergman, David B. Lindell 等NeurIPS 2020 · 被引用 4,008 次
- Neural Sparse Voxel FieldsLingjie Liu, Jiatao Gu, Kyaw Zaw Lin, Tat-Seng Chua 等NeurIPS 2020 · 被引用 1,535 次
- PlenOctrees for Real-time Rendering of Neural Radiance FieldsAlex Yu, Ruilong Li, Matthew Tancik, Hao Li 等ICCV 2021 · 被引用 1,284 次
- KiloNeRF: Speeding up Neural Radiance Fields with Thousands of Tiny MLPsChristian Reiser, Songyou Peng, Yiyi Liao, Andreas GeigerICCV 2021 · 被引用 963 次
- Baking Neural Radiance Fields for Real-Time View SynthesisPeter Hedman, Pratul P. Srinivasan, Ben Mildenhall, Jonathan T. Barron 等ICCV 2021 · 被引用 636 次
相关 Paper
- The Object Folder Benchmark : Multisensory Learning with Neural and Real ObjectsRuohan Gao, Yiming Dou, Hao Li, Tanmay Agarwal 等CVPR 2023
- X-Capture: An Open-Source Portable Device for Multi-Sensory LearningSamuel Clarke, Suzannah Wistreich, Yanjie Ze, Jiajun WuICCV 2025 · 被引用 1 次
- REALIMPACT: A Dataset of Impact Sound Fields for Real ObjectsSamuel Clarke, Ruohan Gao, Mason L. Wang, Mark Rau 等CVPR 2023
- Finding Fallen Objects Via Asynchronous Audio-Visual IntegrationChuang Gan, Yi Gu, Siyuan Zhou, Jeremy Schwartz 等CVPR 2022 · 被引用 13 次
- 3D Shape Reconstruction from Vision and TouchEdward J. Smith, Roberto Calandra, Adriana Romero, Georgia Gkioxari 等NeurIPS 2020 · 被引用 90 次
