Fine-Grained Similarity Measurement between Educational Videos and Exercises
Xin Wang, Wei Huang, Qi Liu, Yu Yin, Zhenya Huang, Le Wu, Jianhui Ma, Xue Wang
Abstract
In online learning systems, measuring the similarity between educational videos and exercises is a fundamental task with great application potentials. In this paper, we explore to measure the fine-grained similarity by leveraging multimodal information. The problem remains pretty much open due to several domain-specific characteristics. First, unlike general videos, educational videos contain not only graphics but also text and formulas, which have a fixed reading order. Both spatial and temporal information embedded in the frames should be modeled. Second, there are semantic associations between adjacent video segments. The semantic associations will affect the similarity and different exercises usually focus on the related context of different ranges. Third, the fine-grained labeled data for training the model is scarce and costly. To tackle the aforementioned challenges, we propose VENet to measure the similarity at both video-level and segment-level by just exploiting the video-level labeled data. Extensive experimental results on real-world data demonstrate the effectiveness of VENet.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 72b51dc3-41c2-4a39-ac9d-af0a35c6a82cCited by top-tier papers3
- Enhancing Knowledge Tracing via Adversarial TrainingXiaopeng Guo, Zhijie Huang, Jie Gao, Mingyu Shang et al.ACM MM 2021 · 100 citations
- NeurJudge: A Circumstance-aware Neural Framework for Legal Judgment PredictionLinan Yue, Qi Liu, Binbin Jin, Han Wu et al.SIGIR 2021 · 83 citations
- Class Gradient Projection For Continual LearningCheng Chen, Ji Zhang, Jingkuan Song, Lianli GaoACM MM 2022 · 14 citations
Builds on1
Related papers
- Fine-grained Cross-modal Alignment Network for Text-Video RetrievalNing Han, Jingjing Chen, Guangyi Xiao, Hao Zhang et al.ACM MM 2021 · 47 citations
- ViSiL: Fine-Grained Spatio-Temporal Video Similarity LearningGiorgos Kordopatis-Zilos, Symeon Papadopoulos, Ioannis Patras, Yiannis KompatsiarisICCV 2019 · 91 citations
- Tencent-MVSE: A Large-Scale Benchmark Dataset for Multi-Modal Video Similarity EvaluationZhaoyang Zeng, Yongsheng Luo, Zhenhua Liu, Fengyun Rao et al.CVPR 2022 · 5 citations
- Video-Level Multimodal Relation Extraction with Event-Entity Semantic ConsistencyZefan Zhang, Weiqi Zhang, Kailong Suo, Yanhui Li et al.ACM MM 2025
- HANet: Hierarchical Alignment Networks for Video-Text RetrievalPeng Wu, Xiangteng He, Mingqian Tang, Yiliang Lv et al.ACM MM 2021 · 62 citations
