Advancing Video Synchronization with Fractional Frame Analysis: Introducing a Novel Dataset and Model
Yuxuan Liu, Haizhou Ai, Junliang Xing, Xuri Li, Xiaoyi Wang, Pin Tao
Abstract
Multiple views play a vital role in 3D pose estimation tasks. Ideally, multi-view 3D pose estimation tasks should directly utilize naturally collected videos for pose estimation. However, due to the constraints of video synchronization, existing methods often use expensive hardware devices to synchronize the initiation of cameras, which restricts most 3D pose collection scenarios to indoor settings. Some recent works learn deep neural networks to align desynchronized datasets derived from synchronized cameras and can only produce frame-level accuracy. For fractional frame video synchronization, this work proposes an Inter-Frame and Intra-Frame Desynchronized Dataset (IFID), which labels fractional time intervals between two video clips. IFID is the first dataset that annotates inter-frame and intra-frame intervals, with a total of 382, 500 video clips annotated, making it the largest dataset to date. We also develop a novel model based on the Transformer architecture, named InSynFormer, for synchronizing inter-frame and intra-frame. Extensive experimental evaluations demonstrate its promising performance. The dataset and source code of the model are available at https: //github.com/yuxuan-cser/InSynFormer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c0e6e483-437c-491f-bb24-8a86a096363eCited by top-tier papers2
- Multi-View 3D Human Pose Estimation with Weakly Synchronized ImagesLing Li, Ruiwen Gu, Chongyang Wang, Junliang Xing et al.AAAI 2025 · 3 citations
- Visual Sync: Multi-Camera Synchronization via Cross-View Object MotionShaowei Liu, David Yifan Yao, Saurabh Gupta, Shenlong WangNeurIPS 2025 · 1 citation
Builds on8
- Optimizing Network Structure for 3D Human Pose EstimationHai Ci, Chunyu Wang, Xiaoxuan Ma, Yizhou WangICCV 2019 · 267 citations
- Direct Multi-view Multi-person 3D Pose EstimationTao Wang, Jianfeng Zhang, Yujun Cai, Shuicheng Yan et al.NeurIPS 2021 · 147 citations
- 3D Human Pose Estimation Using Spatio-Temporal Networks with Explicit Occlusion TrainingYu Cheng, Bo Yang, Bo Wang, Robby T. TanAAAI 2020 · 145 citations
- Graph and Temporal Convolutional Networks for 3D Multi-person Pose Estimation in Monocular VideosYu Cheng, Bo Wang, Bo Yang, Robby T. TanAAAI 2021 · 55 citations
- Novel View Synthesis of Human Interactions from Sparse Multi-view VideosQing Shuai, Chen Geng, Qi Fang, Sida Peng et al.SIGGRAPH 2022 · 43 citations
Related papers
- FreeMan: Towards Benchmarking 3D Human Pose Estimation Under Real-World ConditionsJiong Wang, Fengyu Yang, Bingliang Li, Wenbo Gou et al.CVPR 2024 · 8 citations
- PoseSyn: Synthesizing Diverse 3D Pose Data from In-the-Wild 2D DataChangHee Yang, Hyeonseop Song, Seokhun Choi, Seungwoo Lee et al.ICCV 2025 · 1 citation
- Self-Supervised Human Pose based Multi-Camera Video SynchronizationLiqiang Yin, Ruize Han, Wei Feng, Song WangACM MM 2022 · 7 citations
- Neural Video Depth StabilizerYiran Wang, Min Shi, Jiaqi Li, Zihao Huang et al.ICCV 2023
- 3D Human Pose Estimation with Spatial and Temporal TransformersCe Zheng, Sijie Zhu, Matías Mendieta, Taojiannan Yang et al.ICCV 2021 · 648 citations
