Short Video Ordering via Position Decoding and Successor Prediction
Shiping Ge, Qiang Chen, Zhiwei Jiang, Yafeng Yin, Ziyao Chen, Qing Gu
Abstract
Short video collection is an easy way for users to consume coherent content on various online short video platforms, such as TikTok, YouTube, Douyin, and WeChat Channel. These collections cover a wide range of content, including online courses, TV series, movies, and cartoons. However, short video creators occasionally publish videos in a disorganized manner due to various reasons, such as revisions, secondary creations, deletions, and reissues, which often result in a poor browsing experience for users. Therefore, accurately reordering videos within a collection based on their content coherence is a vital task that can enhance user experience and presents an intriguing research problem in the field of video narrative reasoning. In this work, we curate a dedicated multimodal dataset for this Short Video Ordering (SVO) task and present the performance of some benchmark methods on the dataset. In addition, we further propose an advanced SVO framework with the aid of position decoding and successor prediction. The proposed framework combines both pairwise and listwise ordering paradigms, which can get rid of the issues from both quadratic growth and cascading conflict in the pairwise paradigm, and improve the performance of existing listwise methods. Extensive experiments demonstrate that our method achieves the best performance on our open SVO dataset, and each component of the framework contributes to the final performance. Both the SVO dataset and code will be released at https://github.com/ShipingGe/SVO.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get c40f2be2-4fb8-49c1-91f5-4da63a855be6Cited by top-tier papers1
Ask how each one uses itRelated papers
- A Large-Scale Dataset for Short-Video Topic Peak Prediction and a Large Heterogeneous Graph ModelShangheng Chen, Shengsheng Qian, Quan Fang, Jun Hu et al.ACM MM 2025
- Multimodal Coreference Resolution for Chinese Social Media Dialogues: Dataset and Benchmark ApproachXingyu Li, Chen Gong, Guohong FuACL 2025
- E-VAds: An E-commerce Short Videos Understanding Benchmark for MLLMsXianjie Liu, Yiman Hu, Liang Wu, Ping Hu et al.ICML 2026 · 1 citation
- Mitigating World Biases: A Multimodal Multi-View Debiasing Framework for Fake News Video DetectionZhi Zeng, Minnan Luo, Xiangzheng Kong, Huan Liu et al.ACM MM 2024 · 43 citations
- Spatiotemporal Fine-grained Video Description for Short VideosTe Yang, Jian Jia, Bo Wang, Yanhua Cheng et al.ACM MM 2024 · 1 citation
