Lune

CVPR2026顶会

DyFCLT: Dynamic Frequency-Decoupled Cross-Modal Learning Transformer for Multimodal Tiny Object Detection

Chaolang Li, Pengwen Dai, Jingyu Li, Siyuan Yao, Yuchen Jiang, Zhuoran Zheng

出版方
2026年份

摘要

Multimodal tiny object detection plays a critical role in real-world applications. However, detecting tiny objects remains challenging due to environmental complexities. While recent methods leverage spatial multi-scale representations or frequency-domain enhancements, most focus solely on visible images and overlook complementary multimodal frequency cues. This paper explores how to effectively harness cross-modal frequency information for infrared–visible tiny object detection. Through frequency characteristic analysis, we observe that tiny objects exhibit rich mid- and high-frequency energy across both modalities, motivating the design of a Dynamic Frequency-decoupled Cross-modal Learning Transformer (DyFCLT). Our approach introduces a Dynamic Frequency-Band Decoupled Cross-Modal Attention (DFCA) mechanism to extract and interact frequency components across modalities. To suppress noise while enhancing foreground signals, a Selective Smoothing Enhancement (SSE) strategy is proposed, which smoothes background interference and guides multi-scale feature fusion. DFCA and SSE collaborate to achieve synergistic enrichment and refinement of cross-modal features. Extensive experiments on two tiny-object benchmarks and one general-scale benchmark demonstrate that DyFCLT sets new state-of-the-art results, outperforming prior leading methods by significant margins and exhibiting strong generalization across scales and scenarios.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper26

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖