Scaling Action Detection: AdaTAD++ with Transformer-Enhanced Temporal-Spatial Adaptation
Tanay Agrawal, Abid Ali, Antitza Dantcheva, François Brémond
Abstract
Temporal Action Detection (TAD) is essential for analyzing long-form videos by identifying and segmenting actions within untrimmed sequences. While recent innovations like Temporal Informative Adapters (TIA) have improved resolution, memory constraints still limit large video processing. To address this, we introduce AdaTAD++, an enhanced framework that decouples temporal and spatial processing within adapters, organizing them into independently trainable modules. Our novel two-step training strategy first optimizes for high temporal and low spatial resolution, then vice versa, allowing the model to utilize both high spatial and temporal resolutions during inference, while maintaining training efficiency. Additionally, we incorporate a more sophisticated temporal module capable of capturing long-range dependencies more effectively than previous methods. Experiments on benchmark datasets, including ActivityNet-1.3, THUMOS14, and EPIC-Kitchens 100, demonstrate that AdaTAD++ achieves state-of-the-art performance. We also explore various adapter configurations, discussing their trade-offs regarding resource constraints and performance, providing valuable insights into their optimal application.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d55db3e5-1f95-445c-a648-89e520361914Builds on16
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- AdaptFormer: Adapting Vision Transformers for Scalable Visual RecognitionShoufa Chen, Chongjian Ge, Zhan Tong, Jiangliu Wang et al.NeurIPS 2022 · 1,291 citations
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang et al.ICLR 2022 · 1,218 citations
- LongLoRA: Efficient Fine-tuning of Long-Context Large Language ModelsYukang Chen, Shengju Qian, Haotian Tang, Xin Lai et al.ICLR 2024 · 254 citations
- Fast Learning of Temporal Action Proposal via Dense Boundary GeneratorChuming Lin, Jian Li, Yabiao Wang, Ying Tai et al.AAAI 2020 · 226 citations
Related papers
- End-to-End Temporal Action Detection with 1B Parameters Across 1000 FramesShuming Liu, Chen-Lin Zhang, Chen Zhao, Bernard GhanemCVPR 2024 · 35 citations
- Post-Processing Temporal Action DetectionSauradip Nag, Xiatian Zhu, Yi-Zhe Song, Tao XiangCVPR 2023
- Temporal Action Localization with Cross Layer Task Decoupling and RefinementQiang Li, Di Liu, Jun Kong, Sen Li et al.AAAI 2025 · 3 citations
- RefineTAD: Learning Proposal-free Refinement for Temporal Action DetectionYue Feng, Zhengye Zhang, Rong Quan, Limin Wang et al.ACM MM 2023 · 8 citations
- Low-Fidelity Video Encoder Optimization for Temporal Action LocalizationMengmeng Xu, Juan-Manuel Pérez-Rúa, Xiatian Zhu, Bernard Ghanem et al.NeurIPS 2021 · 30 citations
