Shoe Style-Invariant and Ground-Aware Learning for Dense Foot Contact Estimation
Daniel Jung, Kyoung Mu Lee
Abstract
Foot contact plays a critical role in human interaction with the world, and thus exploring foot contact can advance our understanding of human movement and physical interaction. Despite its importance, existing methods often approximate foot contact using a zero-velocity constraint and focus on joint-level contact, failing to capture the detailed interaction between the foot and the world. Dense estimation of foot contact is crucial for accurately modeling this interaction, yet predicting dense foot contact from a single RGB image remains largely underexplored. There are two main challenges for learning dense foot contact estimation. First, shoes exhibit highly diverse appearances, making it difficult for models to generalize across different styles. Second, ground often has a monotonous appearance, making it difficult to extract informative features. To tackle these issues, we present a FEet COntact estimation (FECO) framework that learns dense foot contact with shoe style-invariant and ground-aware learning. To overcome the challenge of shoe appearance diversity, our approach incorporates shoe style adversarial training that enforces shoe style-invariant features for contact estimation. To effectively utilize ground information, we introduce a ground feature extractor that captures ground properties based on spatial context. As a result, our proposed method achieves robust foot contact estimation regardless of shoe appearance and effectively leverages ground information. The codes are available at https://github. com/dqj5182/FECO_RELEASE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 256790bd-d474-4cd4-b014-2631eb1983a9Builds on30
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Domain Generalization with MixStyleKaiyang Zhou, Yongxin Yang, Yu Qiao, Tao XiangICLR 2021 · 986 citations
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu et al.CVPR 2024 · 847 citations
- HuMoR: 3D Human Motion Model for Robust Pose EstimationDavis Rempe, Tolga Birdal, Aaron Hertzmann, Jimei Yang et al.ICCV 2021 · 398 citations
Related papers
- LaserShoes: Low-Cost Ground Surface Detection Using Laser Speckle ImagingZihan Yan, Yuxiaotong Lin, Guanyun Wang, Yu Cai et al.CHI 2023 · 14 citations
- Learning Dense Hand Contact Estimation from Imbalanced DataDaniel Sungho Jung, Kyoung Mu LeeNeurIPS 2025 · 14 citations
- Towards Stable Human Pose Estimation via Cross-View Fusion and Foot StabilizationLi'an Zhuo, Jian Cao, Qi Wang, Bang Zhang et al.CVPR 2023
- Automatic Human Scene Interaction through Contact Estimation and Motion AdaptationMingrui Zhang, Ming Chen, Yan Zhou, Li Chen et al.ACM MM 2023
- FrictShoes: Providing Multilevel Nonuniform Friction Feedback on Shoes in VRChih-An Tsao, Tzu-Chun Wu, Hsin-Ruey Tsai, Tzu-Yun Wei et al.IEEE VR 2022 · 15 citations
