Footprints and Free Space From a Single Color Image
Jamie Watson, Michael Firman, Áron Monszpart, Gabriel J. Brostow
Abstract
Understanding the shape of a scene from a single color image is a formidable computer vision task. However, most methods aim to predict the geometry of surfaces that are visible to the camera, which is of limited use when planning paths for robots or augmented reality agents. Such agents can only move when grounded on a traversable surface, which we define as the set of classes which humans can also walk over, such as grass, footpaths and pavement. Models which predict beyond the line of sight often parameterize the scene with voxels or meshes, which can be expensive to use in machine learning frameworks. We introduce a model to predict the geometry of both visible and occluded traversable surfaces, given a single RGB image as input. We learn from stereo video sequences, using camera poses, per-frame depth and semantic segmentation to form training data, which is used to supervise an imageto-image network. We train models from the KITTI driving dataset, the indoor Matterport dataset, and from our own casually captured stereo footage. We find that a surprisingly low bar for spatial coverage of training scenes is required. We validate our algorithm against a range of strong baselines, and include an assessment of our predictions for a path-planning task.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4e60f9ed-04bf-47d6-9d92-e5efd924b33bCited by top-tier papers1
Ask how each one uses itBuilds on3
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 2,416 citations
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 1,467 citations
- Self-Supervised Monocular Depth HintsJamie Watson, Michael Firman, Gabriel J. Brostow, Daniyar TurmukhambetovICCV 2019 · 287 citations
Related papers
- SelfOcc: Self-Supervised Vision-Based 3D Occupancy PredictionYuanhui Huang, Wenzhao Zheng, Borui Zhang, Jie Zhou et al.CVPR 2024
- LaRI: Layered Ray Intersections for Single-view 3D Geometric ReasoningRui Li, Biao Zhang, Zhenyu Li, Federico Tombari et al.ICML 2026
- Dream-to-Recon: Monocular 3D Reconstruction with Diffusion-Depth Distillation from Single ImagesPhilipp Wulff, Felix Wimbauer, Dominik Muhle, Daniel CremersICCV 2025 · 1 citation
- Holistic 3D Human and Scene Mesh Estimation From Single View ImagesZhenzhen Weng, Serena YeungCVPR 2021
- Behind the Scenes: Density Fields for Single View ReconstructionFelix Wimbauer, Nan Yang, Christian Rupprecht, Daniel CremersCVPR 2023
