BIGS: Bimanual Category-agnostic Interaction Reconstruction from Monocular Videos via 3D Gaussian Splatting
Jeongwan On, Kyeonghwan Gwak, Gunyoung Kang, Junuk Cha, Soohyun Hwang, Hyein Hwang, Seungryul Baek
Abstract
Novel view video (e.g., egocentric view) Output Novel hand / object / camera pose Animatable hands and objects time Reconstructed hand-object interaction Figure 1. Our approach reconstructs 3D Gaussians of bimanual category-agnostic interactions from a monocular video, where the two hands interact with an unknown object. Even with limited observations, our method reliably builds the 3D Gaussians in this scenario and once 3D Gaussians are built, our method can be used to render new videos with novel poses of hand, object and camera (i.e., view).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- ArtHOI: Taming Foundation Models for Monocular 4D Reconstruction of Hand-Articulated-Object InteractionsZikai Wang, Zhilu Zhang, Yiqing Wang, Hui Li et al.CVPR 2026 · 4 citations
- HandVQA: Diagnosing and Improving Fine-Grained Spatial Reasoning about Hands in Vision-Language ModelsKhalequzzaman Chowdhury Sayem, Mubarrat Chowdhury, Yihalem Yimolal Tiruneh, Muneeb Ahmed Khan et al.CVPR 2026 · 3 citations
- Clay-to-Stone: Phase-wise 3D Gaussian Splatting for Monocular Articulated Hand-Object Manipulation ModelingXingyu Liu, Pengfei Ren, Qi Qi, Haifeng Sun et al.CVPR 2026 · 1 citation
- AGILE: Hand-object Interaction Reconstruction from Video via Agentic GenerationJin-Chuan Shi, Binhong Ye, Tao Liu, Xiaoyang Liu et al.SIGGRAPH 2026
Builds on36
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Pixel Difference Networks for Efficient Edge DetectionZhuo Su, Wenzhe Liu, Zitong Yu, Dewen Hu et al.ICCV 2021 · 488 citations
- DreamFusion: Text-to-3D using 2D DiffusionBen Poole, Ajay Jain, Jonathan T. Barron, Ben MildenhallICLR 2023 · 463 citations
Related papers
- HOLD: Category-Agnostic 3D Reconstruction of Interacting Hands and Objects from VideoZicong Fan, Maria Parelli, Maria Eleni Kadoglou, Xu Chen et al.CVPR 2024
- Diffusion-Guided Reconstruction of Everyday Hand-Object Interaction ClipsYufei Ye, Poorvi Hebbar, Abhinav Gupta, Shubham TulsianiICCV 2023 · 80 citations
- MagicHOI: Leveraging 3D Priors for Accurate Hand-Object Reconstruction from Short Monocular Video ClipsShibo Wang, Haonan He, Maria Parelli, Christoph Gebhardt et al.ICCV 2025 · 2 citations
- Guess The Unseen: Dynamic 3D Scene Reconstruction from Partial 2D GlimpsesInhee Lee, Byungjun Kim, Hanbyul JooCVPR 2024
- Hand-held Object Reconstruction from RGB Video with Dynamic InteractionShijian Jiang, Qi Ye, Rengan Xie, Yuchi Huo et al.CVPR 2025
