StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval
Shaokun Wang, Weili Guan, Jizhou Han, Jianlong Wu, Yupeng Hu, Liqiang Nie
Abstract
Continual Text-to-Video Retrieval (CTVR) is a challenging multimodal continual learning setting, where models must incrementally learn new semantic categories while maintaining accurate text-video alignment for previously learned ones, thus making it particularly prone to catastrophic forgetting. A key challenge in CTVR is feature drift, which manifests in two forms: intra-modal feature drift caused by continual learning within each modality, and non-cooperative feature drift across modalities that leads to modality misalignment. To mitigate these issues, we propose StructAlign, a structured cross-modal alignment method for CTVR. First, StructAlign introduces a simplex Equiangular Tight Frame (ETF) geometry as a unified geometric prior to mitigate modality misalignment. Building upon this geometric prior, we design a cross-modal ETF alignment loss that aligns text and video features with category-level ETF prototypes, encouraging the learned representations to form an approximate simplex ETF geometry. In addition, to suppress intra-modal feature drift, we design a Cross-modal Relation Preserving loss, which leverages complementary modalities to preserve cross-modal similarity relations, providing stable relational supervision for feature updates. By jointly addressing non-cooperative feature drift across modalities and intra-modal feature drift, StructAlign effectively alleviates catastrophic forgetting in CTVR. Extensive experiments on benchmark datasets demonstrate that our method shows competitive advantages over state-of-the-art continual retrieval approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0f4013f4-6786-4be3-8b92-33bf0247cbb7Cited by top-tier papers1
Ask how each one uses itBuilds on39
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 1,550 citations
- X-Pool: Cross-Modal Language-Video Attention for Text-Video RetrievalSatya Krishna Gorti, Noël Vouitsis, Junwei Ma, Keyvan Golestan et al.CVPR 2022 · 190 citations
- TACo: Token-aware Cascade Contrastive Learning for Video-Text AlignmentJianwei Yang, Yonatan Bisk, Jianfeng GaoICCV 2021 · 159 citations
- TeachText: CrossModal Generalized Distillation for Text-Video RetrievalIoana Croitoru, Simion-Vlad Bogolin, Marius Leordeanu, Hailin Jin et al.ICCV 2021 · 147 citations
Related papers
- Continual Text-to-Video Retrieval with Frame Fusion and Task-Aware RoutingZecheng Zhao, Zhi Chen, Zi Huang, Shazia Sadiq et al.SIGIR 2025 · 6 citations
- Boosting Multi-Modal Alignment: Geometric Feature Separation for Class Incremental LearningGuoqiang Liang, Chuan Qin, De Cheng, Shizhou Zhang et al.ACM MM 2025
- Continual Vision-Language Representation Learning with Off-Diagonal InformationZixuan Ni, Longhui Wei, Siliang Tang, Yueting Zhuang et al.ICML 2023 · 40 citations
- GOAL: Geometrically Optimal Alignment for Continual Generalized Category DiscoveryJizhou Han, Chenhao Ding, Songlin Dong, Yuhang He et al.AAAI 2026 · 2 citations
- From Selection to Scheduling: Federated Geometry-Aware Correction Makes Exemplar Replay Work Better under Continual Dynamic HeterogeneityZhuang Qi, Ying-Peng Tang, Lei Meng, Guoqing Chao et al.CVPR 2026
