JTD-UAV: MLLM-Enhanced Joint Tracking and Description Framework for Anti-UAV Systems
Yifan Wang, Jian Zhao, Zhaoxin Fan, Xin Zhang, Xuecheng Wu, Yudian Zhang, Lei Jin, Xinyue Li, Gang Wang, Mengxi Jia, Ping Hu, Zheng Zhu, Xuelong Li
摘要
Unmanned Aerial Vehicles (UAVs) are widely adopted across various fields, yet they raise significant privacy and safety concerns, demanding robust monitoring solutions. Existing anti-UAV methods primarily focus on position tracking but fail to capture UAV behavior and intent. To address this, we introduce a novel task-UAV Tracking and Intent Understanding (UTIU)-which aims to track UAVs while inferring and describing their motion states and intent for a more comprehensive monitoring approach. To tackle the task, we propose JTD-UAV, the first joint tracking, and intent description framework based on large language models. Our dual-branch architecture integrates UAV tracking with Visual Question Answering (VQA), allowing simultaneous localization and behavior description. To benchmark this task, we introduce the TDUAV dataset, the largest dataset for joint UAV tracking and intent understanding, featuring 1,328 challenging video sequences, over 163K annotated thermal frames, and 3K VQA pairs. Our benchmark demonstrates the effectiveness of JTD-UAV.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Tracking Tiny Drones Against Clutter: Large-Scale Infrared Benchmark with Motion-Centric Adaptive AlgorithmJiahao Zhang, Zongli Jiang, Jinli Zhang, Yixin Wei 等ICCV 2025 · 被引用 3 次
- Diffuse to Detect: Bi-Level Sample Rebalancing with Pseudo-Label Diffusion for Point-Supervised Infrared Small-Target DetectionZhu Liu, Yuanhang Yao, Ping Qian, Zihang Chen 等ICML 2026
它引用的顶会 Paper28
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- YOLOv10: Real-Time End-to-End Object DetectionAo Wang, Hui Chen, Lihao Liu, Kai Chen 等NeurIPS 2024 · 被引用 6,113 次
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch 等ICML 2023 · 被引用 2,601 次
相关 Paper
- AerialMind: Towards Referring Multi-Object Tracking in UAV ScenariosChenglizhao Chen, Shaofeng Liang, Runwei Guan, Xiaolou Sun 等AAAI 2026 · 被引用 2 次
- Towards Autonomous UAV Visual Object Search in City Space: Benchmark and Agentic MethodologyYatai Ji, Zhengqiu Zhu, Yong Zhao, Beidan Liu 等AAAI 2026 · 被引用 8 次
- MOR-UAV: A Benchmark Dataset and Baselines for Moving Object Recognition in UAV VideosMurari Mandal, Lav Kush Kumar, Santosh Kumar VipparthiACM MM 2020 · 被引用 58 次
- AeroDuo: Aerial Duo for UAV-based Vision and Language NavigationRuipu Wu, Yige Zhang, Jinyu Chen, Linjiang Huang 等ACM MM 2025 · 被引用 6 次
- Towards Realistic UAV Vision-Language Navigation: Platform, Benchmark, and MethodologyXiangyu Wang, Donglin Yang, Ziqin Wang, Hohin Kwan 等ICLR 2025
