Skating-Mixer: Long-Term Sport Audio-Visual Modeling with MLPs
Jingfei Xia, Mingchen Zhuge, Tiantian Geng, Shun Fan, Yuantai Wei, Zhenyu He, Feng Zheng
Abstract
Figure skating scoring is challenging because it requires judging the technical moves of the players as well as their coordination with the background music. Most learning-based methods cannot solve it well for two reasons: 1) each move in figure skating changes quickly, hence simply applying traditional frame sampling will lose a lot of valuable information, especially in 3 to 5 minutes long videos; 2) prior methods rarely considered the critical audio-visual relationship in their models. Due to these reasons, we introduce a novel architecture, named Skating-Mixer. It extends the MLP framework into a multimodal fashion and effectively learns longterm representations through our designed memory recurrent unit (MRU). Aside from the model, we collected a highquality audio-visual FS1000 dataset, which contains over 1000 videos on 8 types of programs with 7 different rating metrics, overtaking other datasets in both quantity and diversity. Experiments show the proposed method achieves SOTAs over all major metrics on the public Fis-V and our FS1000 dataset. In addition, we include an analysis applying our method to the recent competitions in Beijing 2022 Winter Olympic Games, proving our method has strong applicability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 112e9401-df55-4b8b-9132-6165401290c4Cited by top-tier papers6
- BriMA: Bridged Modality Adaptation for Multi-Modal Continual Action Quality AssessmentKanglei Zhou, Chang Li, Qingyi Pan, Liyuan WangCVPR 2026 · 3 citations
- MCMoE: Completing Missing Modalities with Mixture of Experts for Incomplete Multimodal Action Quality AssessmentHuangbiao Xu, Huanqi Wu, Xiao Ke, Junyi Wu et al.AAAI 2026 · 2 citations
- LIMSSR: LLM-Driven Sequence-to-Score Reasoning under Training-Time Incomplete Multimodal ObservationsHuangbiao Xu, huanqi wu, Xiao Ke, Yuxin PengICML 2026 · 1 citation
- TechCoach: Towards Technical-Point-Aware Descriptive Action CoachingYuan-Ming Li, An-Lan Wang, Ling-An Zeng, Kun-Yu Lin et al.AAAI 2026
- Language-Guided Audio-Visual Learning for Long-Term Sports AssessmentHuangbiao Xu, Xiao Ke, Huanqi Wu, Rui Xu et al.CVPR 2025
Builds on14
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 1,550 citations
- Attention Bottlenecks for Multimodal FusionArsha Nagrani, Shan Yang, Anurag Arnab, Aren Jansen et al.NeurIPS 2021 · 884 citations
Related papers
- Learning Long-Range Action Representation by Two-Stream Mamba Pyramid Network for Figure Skating AssessmentFengshun Wang, Qiurui Wang, Peilin ZhaoACM MM 2025 · 1 citation
- FSBench: A Figure Skating Benchmark for Advancing Artistic Sports UnderstandingRong Gao, Xin Liu, Zhuozhao Hu, Bohao Xing et al.CVPR 2025
- Localization-assisted Uncertainty Score Disentanglement Network for Action Quality AssessmentYanli Ji, Lingfeng Ye, Huili Huang, Lijing Mao et al.ACM MM 2023 · 25 citations
- Long-Term Rhythmic Video SoundtrackerJiashuo Yu, Yaohui Wang, Xinyuan Chen, Xiao Sun et al.ICML 2023 · 24 citations
- A Figure Skating Jumping Dataset for Replay-Guided Action Quality AssessmentYanchao Liu, Xina Cheng, Takeshi IkenagaACM MM 2023 · 16 citations
