Hypersim: A Photorealistic Synthetic Dataset for Holistic Indoor Scene Understanding
Mike Roberts, Jason Ramapuram, Anurag Ranjan, Atulit Kumar, Miguel Ángel Bautista, Nathan Paczan, Russ Webb, Joshua M. Susskind
摘要
For many fundamental scene understanding tasks, it is difficult or impossible to obtain per-pixel ground truth labels from real images. We address this challenge by introducing Hypersim, a photorealistic synthetic dataset for holistic indoor scene understanding. To create our dataset, we leverage a large repository of synthetic scenes created by professional artists, and we generate 77,400 images of 461 indoor scenes with detailed per-pixel labels and corresponding ground truth geometry. Our dataset: (1) relies exclusively on publicly available 3D assets; (2) includes complete scene geometry, material information, and lighting information for every scene; (3) includes dense per-pixel semantic instance segmentations and complete camera information for every image; and (4) factors every image into diffuse reflectance, diffuse illumination, and a non-diffuse residual term that captures view-dependent lighting effects.We analyze our dataset at the level of scenes, objects, and pixels, and we analyze costs in terms of money, computation time, and annotation effort. Remarkably, we find that it is possible to generate our entire dataset from scratch, for roughly half the cost of training a popular open-source natural language processing model. We also evaluate sim-to-real transfer performance on two real-world scene understanding tasks – semantic segmentation and 3D shape prediction – where we find that pre-training on our dataset significantly improves performance on both tasks, and achieves state-of-the-art performance on the most challenging Pix3D test set. All of our rendered image data, as well as all the code we used to generate our dataset and perform our experiments, is available online.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper243
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao 等NeurIPS 2024 · 被引用 2,305 次
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMsPeter Tong, Ellis Brown, Penghao Wu, Sanghyun Woo 等NeurIPS 2024 · 被引用 1,004 次
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu 等CVPR 2024 · 被引用 847 次
- Depth Anything 3: Recovering the Visual Space from Any ViewsHaotong Lin, Sili Chen, Jun Hao Liew, Donny Y. Chen 等ICLR 2026 · 被引用 720 次
它引用的顶会 Paper5
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra 等ICCV 2019 · 被引用 1,863 次
- Which Tasks Should Be Learned Together in Multi-task Learning?Trevor Standley, Amir Zamir, Dawn Chen, Leonidas J. Guibas 等ICML 2020 · 被引用 651 次
- Neural Inverse Rendering of an Indoor Scene From a Single ImageSoumyadip Sengupta, Jinwei Gu, Kihwan Kim, Guilin Liu 等ICCV 2019 · 被引用 172 次
- Inverse Rendering for Complex Indoor Scenes: Shape, Spatially-Varying Lighting and SVBRDF From a Single ImageZhengqin Li, Mohammad Shafiei, Ravi Ramamoorthi, Kalyan Sunkavalli 等CVPR 2020
- Mesh R-CNNGeorgia Gkioxari, Justin Johnson, Jitendra MalikICCV 2019
相关 Paper
- GeoSynth: A Photorealistic Synthetic Indoor Dataset for Scene UnderstandingBrian Pugh, Davin Chernak, Salma JiddiIEEE VR 2023 · 被引用 7 次
- OpenRooms: An Open Framework for Photorealistic Indoor Scene DatasetsZhengqin Li, Ting-Wei Yu, Shen Sang, Sarah Wang 等CVPR 2021
- 3D Segmentation of Humans in Point Clouds with Synthetic DataAyça Takmaz, Jonas Schult, Irem Kaftan, Mertcan Akçay 等ICCV 2023 · 被引用 31 次
- Indoor Scene Generation from a Collection of Semantic-Segmented Depth ImagesMingjia Yang, Yu-Xiao Guo, Bin Zhou, Xin TongICCV 2021 · 被引用 41 次
- PanoContext-Former: Panoramic Total Scene Understanding with a TransformerYuan Dong, Chuan Fang, Liefeng Bo, Zilong Dong 等CVPR 2024
