Is Multiple Object Tracking a Matter of Specialization?
Gianluca Mancusi, Mattia Bernardi, Aniello Panariello, Angelo Porrello, Rita Cucchiara, Simone Calderara
Abstract
End-to-end transformer-based trackers have achieved remarkable performance on most human-related datasets. However, training these trackers in heterogeneous scenarios poses significant challenges, including negative interference - where the model learns conflicting scene-specific parameters - and limited domain generalization, which often necessitates expensive fine-tuning to adapt the models to new domains. In response to these challenges, we introduce Parameter-efficient Scenario-specific Tracking Architecture (PASTA), a novel framework that combines Parameter-Efficient Fine-Tuning (PEFT) and Modular Deep Learning (MDL). Specifically, we define key scenario attributes (e.g, camera-viewpoint, lighting condition) and train specialized PEFT modules for each attribute. These expert modules are combined in parameter space, enabling systematic generalization to new domains without increasing inference time. Extensive experiments on MOTSynth, along with zero-shot evaluations on MOT17 and PersonPath22 demonstrate that a neural tracker built from carefully selected modules surpasses its monolithic counterpart. We release models and code.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Accurate and Efficient Low-Rank Model Merging in Core SpaceAniello Panariello, Daniel Marczak, Simone Magistri, Angelo Porrello et al.NeurIPS 2025 · 32 citations
- DitHub: A Modular Framework for Incremental Open-Vocabulary Object DetectionChiara Cappellino, Gianluca Mancusi, Matteo Mosconi, Angelo Porrello et al.NeurIPS 2025 · 4 citations
Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context LearningHaokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta et al.NeurIPS 2022 · 1,483 citations
- Tracking Without Bells and WhistlesPhilipp Bergmann, Tim Meinhardt, Laura Leal-TaixéICCV 2019 · 1,030 citations
- TrackFormer: Multi-Object Tracking with TransformersTim Meinhardt, Alexander Kirillov, Laura Leal-Taixé, Christoph FeichtenhoferCVPR 2022 · 927 citations
Related papers
- SEATrack: Simple, Efficient, and Adaptive Multimodal TrackerJunbin Su, Ziteng Xue, Shihui Zhang, Kun Chen et al.CVPR 2026 · 3 citations
- TADFormer: Task-Adaptive Dynamic TransFormer for Efficient Multi-Task LearningSeungmin Baek, Soyul Lee, Hayeon Jo, Hyesong Choi et al.CVPR 2025
- TDSS: Task Dynamic-Synergistic Skill Adaptation for Boosting Efficient and Scalable Multi-Task Learning in Dense Visual PredictionHaiming Yao, Qiyu Chen, Wei Luo, Zheng Zhang et al.AAAI 2026
- PEANuT: Parameter-Efficient Adaptation with Weight-aware Neural TweakersYibo Zhong, Haoxiang Jiang, Lincan Li, Ryumei Nakada et al.KDD 2026 · 7 citations
- MoDULA: Mixture of Domain-Specific and Universal LoRA for Multi-Task LearningYufei Ma, Zihan Liang, Huangyu Dai, Ben Chen et al.EMNLP 2024 · 4 citations
