HarmonySet: A Comprehensive Dataset for Understanding Video-Music Semantic Alignment and Temporal Synchronization
Zitang Zhou, Ke Mei, Yu Lu, Tianyi Wang, Fengyun Rao
2025Year
3Top-tier citations
Abstract
with detailed information on rhythmic synchronization, emotional alignment, thematic coherence, and cultural relevance. We propose a multi-step human-machine collaborative framework for efficient annotation, combining human insights with machine-generated descriptions to identify key transitions and assess alignment across multiple dimensions. Addition-
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 12e7403e-f27f-4de1-8c0d-61a2feddbe6eCited by top-tier papers3
- AudioX: A Unified Framework for Anything-to-Audio GenerationZeyue Tian, Zhaoyang Liu, Yizhu Jin, Ruibin Yuan et al.ICLR 2026 · 38 citations
- FlexSelect: Flexible Token Selection for Efficient Long Video UnderstandingYunzhu Zhang, Yu Lu, Tianyi Wang, Fengyun Rao et al.NeurIPS 2025 · 22 citations
- Video Echoed in Music: Semantic, Temporal, and Rhythmic Alignment for Video-to-Music GenerationXinyi Tong, Yiran Zhu, Jishang Chen, Chunru Zhan et al.AAAI 2026 · 4 citations
Builds on25
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- MM-Vet: Evaluating Large Multimodal Models for Integrated CapabilitiesWeihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang et al.ICML 2024 · 1,191 citations
- AI Choreographer: Music Conditioned 3D Dance Generation with AIST++Ruilong Li, Shan Yang, David A. Ross, Angjoo KanazawaICCV 2021 · 701 citations
- LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic AlignmentBin Zhu, Bin Lin, Munan Ning, Yang Yan et al.ICLR 2024 · 403 citations
- Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language ModelsMuhammad Maaz, Hanoona Abdul Rasheed, Salman Khan, Fahad KhanACL 2024 · 279 citations
Related papers
- TMD-Bench: A Multi-Level Evaluation Paradigm for Music–Dance Co-GenerationXiaoda Yang, Majun Zhang, Changhao Pan, Nick Huang et al.ICML 2026 · 1 citation
- Instilling an Active Mind in Avatars via Cognitive SimulationJianwen Jiang, Weihong Zeng, Zerong Zheng, Jiaqi Yang et al.ICLR 2026 · 26 citations
- DanceEditor: Towards Iterative Editable Music-Driven Dance Generation with Open-Vocabulary DescriptionsHengyuan Zhang, Zhe Li, Xingqun Qi, Mengze Li et al.ICCV 2025 · 3 citations
- It's Time for Artistic Correspondence in Music and VideoDídac Surís, Carl Vondrick, Bryan C. Russell, Justin SalamonCVPR 2022 · 33 citations
- MotivDance: Fine-Grained Text-Guided Motivation Choreography with Music SynchronizationChenguang Li, Yu-Hui Wen, Liping JingAAAI 2026
