Omnidata: A Scalable Pipeline for Making Multi-Task Mid-Level Vision Datasets from 3D Scans
Ainaz Eftekhar, Alexander Sax, Jitendra Malik, Amir Zamir
摘要
This paper introduces a pipeline to parametrically sample and render static multi-task vision datasets from comprehensive 3D scans from the real-world. In addition to enabling interesting lines of research, we show the tooling and generated data suffice to train robust vision models. Familiar architectures trained on a generated starter dataset reached state-of-the-art performance on multiple common vision tasks and benchmarks, despite having seen no benchmark or non-pipeline data. The depth estimation network outperforms MiDaS and the surface normal estimation network is the first to achieve human-level performance for in-the-wild surface normal estimation—at least according to one metric on the OASIS benchmark. The Dockerized pipeline with CLI, the (mostly python) code, PyTorch dataloaders for the generated data, the generated starter dataset, download scripts and other utilities are all available .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper158
- MonoSDF: Exploring Monocular Geometric Cues for Neural Implicit Surface ReconstructionZehao Yu, Songyou Peng, Michael Niemeyer, Torsten Sattler 等NeurIPS 2022 · 被引用 670 次
- Metric3D: Towards Zero-shot Metric 3D Prediction from A Single ImageWei Yin, Chi Zhang, Hao Chen, Zhipeng Cai 等ICCV 2023 · 被引用 388 次
- ZSON: Zero-Shot Object-Goal Navigation using Multimodal Goal EmbeddingsArjun Majumdar, Gunjan Aggarwal, Bhavika Devnani, Judy Hoffman 等NeurIPS 2022 · 被引用 344 次
- DreamCraft3D: Hierarchical 3D Generation with Bootstrapped Diffusion PriorJingxiang Sun, Bo Zhang, Ruizhi Shao, Lizhen Wang 等ICLR 2024 · 被引用 181 次
- 4M: Massively Multimodal Masked ModelingDavid Mizrahi, Roman Bachmann, Oguzhan Fatih Kar, Teresa Yeo 等NeurIPS 2023 · 被引用 154 次
它引用的顶会 Paper10
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra 等ICCV 2019 · 被引用 1,863 次
- Hypersim: A Photorealistic Synthetic Dataset for Holistic Indoor Scene UnderstandingMike Roberts, Jason Ramapuram, Anurag Ranjan, Atulit Kumar 等ICCV 2021 · 被引用 633 次
相关 Paper
- OASIS: A Large-Scale Dataset for Single Image 3D in the WildWeifeng Chen, Shengyi Qian, David Fan, Noriyuki Kojima 等CVPR 2020
- MPI-Flow: Learning Realistic Optical Flow with Multiplane ImagesYingping Liang, Jiaming Liu, Debing Zhang, Ying FuICCV 2023 · 被引用 12 次
- OmniObject3D: Large-Vocabulary 3D Object Dataset for Realistic Perception, Reconstruction and GenerationTong Wu, Jiarui Zhang, Xiao Fu, Yuxin Wang 等CVPR 2023
- DAViD: Data-Efficient and Accurate Vision Models from Synthetic Data DAViD also references Michelangelo's David - an iconic symbol of anatomical precision-and the David vs. Goliath story, reflecting our small yet powerful dataset and modelsFatemeh Saleh, Sadegh Aliakbarian, Charlie Hewitt, Lohit Petikam 等ICCV 2025 · 被引用 3 次
- Fake it till you make it: face analysis in the wild using synthetic data aloneErroll Wood, Tadas Baltrusaitis, Charlie Hewitt, Sebastian Dziadzio 等ICCV 2021 · 被引用 331 次
