Skating-Mixer: Long-Term Sport Audio-Visual Modeling with MLPs
Jingfei Xia, Mingchen Zhuge, Tiantian Geng, Shun Fan, Yuantai Wei, Zhenyu He, Feng Zheng
摘要
Figure skating scoring is challenging because it requires judging the technical moves of the players as well as their coordination with the background music. Most learning-based methods cannot solve it well for two reasons: 1) each move in figure skating changes quickly, hence simply applying traditional frame sampling will lose a lot of valuable information, especially in 3 to 5 minutes long videos; 2) prior methods rarely considered the critical audio-visual relationship in their models. Due to these reasons, we introduce a novel architecture, named Skating-Mixer. It extends the MLP framework into a multimodal fashion and effectively learns longterm representations through our designed memory recurrent unit (MRU). Aside from the model, we collected a highquality audio-visual FS1000 dataset, which contains over 1000 videos on 8 types of programs with 7 different rating metrics, overtaking other datasets in both quantity and diversity. Experiments show the proposed method achieves SOTAs over all major metrics on the public Fis-V and our FS1000 dataset. In addition, we include an analysis applying our method to the recent competitions in Beijing 2022 Winter Olympic Games, proving our method has strong applicability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- BriMA: Bridged Modality Adaptation for Multi-Modal Continual Action Quality AssessmentKanglei Zhou, Chang Li, Qingyi Pan, Liyuan WangCVPR 2026 · 被引用 3 次
- MCMoE: Completing Missing Modalities with Mixture of Experts for Incomplete Multimodal Action Quality AssessmentHuangbiao Xu, Huanqi Wu, Xiao Ke, Junyi Wu 等AAAI 2026 · 被引用 2 次
- LIMSSR: LLM-Driven Sequence-to-Score Reasoning under Training-Time Incomplete Multimodal ObservationsHuangbiao Xu, huanqi wu, Xiao Ke, Yuxin PengICML 2026 · 被引用 1 次
- TechCoach: Towards Technical-Point-Aware Descriptive Action CoachingYuan-Ming Li, An-Lan Wang, Ling-An Zeng, Kun-Yu Lin 等AAAI 2026
- Language-Guided Audio-Visual Learning for Long-Term Sports AssessmentHuangbiao Xu, Xiao Ke, Huanqi Wu, Rui Xu 等CVPR 2025
它引用的顶会 Paper14
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer 等NeurIPS 2021 · 被引用 3,862 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 被引用 1,550 次
- Attention Bottlenecks for Multimodal FusionArsha Nagrani, Shan Yang, Anurag Arnab, Aren Jansen 等NeurIPS 2021 · 被引用 884 次
相关 Paper
- Learning Long-Range Action Representation by Two-Stream Mamba Pyramid Network for Figure Skating AssessmentFengshun Wang, Qiurui Wang, Peilin ZhaoACM MM 2025 · 被引用 1 次
- FSBench: A Figure Skating Benchmark for Advancing Artistic Sports UnderstandingRong Gao, Xin Liu, Zhuozhao Hu, Bohao Xing 等CVPR 2025
- Localization-assisted Uncertainty Score Disentanglement Network for Action Quality AssessmentYanli Ji, Lingfeng Ye, Huili Huang, Lijing Mao 等ACM MM 2023 · 被引用 25 次
- Long-Term Rhythmic Video SoundtrackerJiashuo Yu, Yaohui Wang, Xinyuan Chen, Xiao Sun 等ICML 2023 · 被引用 24 次
- A Figure Skating Jumping Dataset for Replay-Guided Action Quality AssessmentYanchao Liu, Xina Cheng, Takeshi IkenagaACM MM 2023 · 被引用 16 次
