Dynamic MLP for Fine-Grained Image Classification by Leveraging Geographical and Temporal Information
Lingfeng Yang, Xiang Li, Renjie Song, Borui Zhao, Juntian Tao, Shihao Zhou, Jiajun Liang, Jian Yang
摘要
Fine-grained image classification is a challenging computer vision task where various species share similar visual appearances, resulting in misclassification if merely based on visual clues. Therefore, it is helpful to leverage additional information, e.g., the locations and dates for data shooting, which can be easily accessible but rarely exploited. In this paper, we first demonstrate that existing multimodal methods fuse multiple features only on a single dimension, which essentially has insufficient help in feature discrimination. To fully explore the potential of multimodal information, we propose a dynamic MLP on top of the image representation, which interacts with multimodal features at a higher and broader dimension. The dynamic MLP is an efficient structure parameterized by the learned embeddings of variable locations and dates. It can be regarded as an adaptive nonlinear projection for generating more discriminative image representations in visual tasks. To our best knowledge, it is the first attempt to explore the idea of dynamic networks to exploit multimodal information in fine-grained image classification tasks. Extensive experiments demonstrate the effectiveness of our method. The t-SNE algorithm visually indicates that our technique improves the recognizability of image representations that are visually similar but with different categories. Furthermore, among published works across multiple fine-grained datasets, dynamic MLP consistently achieves SOTA results <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> https://paperswithcode.com/dataset/inaturalist and takes third place in the iNaturalist challenge at FGVC8 <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup> <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup> https://www.kaggle.com/c/inaturalist-2021/leaderboard. Code is available at httpsr//glthub.com/megvii-research/DynamicMLPForFinegrained.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- CSP: Self-Supervised Contrastive Spatial Pre-Training for Geospatial-Visual RepresentationsGengchen Mai, Ni Lao, Yutong He, Jiaming Song 等ICML 2023 · 被引用 103 次
- Spatial Implicit Neural Representations for Global-Scale Species MappingElijah Cole, Grant Van Horn, Christian Lange, Alexander Shepard 等ICML 2023 · 被引用 70 次
- MapFormer: Boosting Change Detection by Using Pre-change InformationMaximilian Bernhard, Niklas Strauß, Matthias SchubertICCV 2023 · 被引用 15 次
- Fusion Meets Diverse Conditions: A High-Diversity Benchmark and Baseline for UAV-Based Multimodal Object Detection with Condition CuesChen Chen, Kangcheng Bin, Ting Hu, Jiahao Qi 等ICCV 2025 · 被引用 8 次
- LocDiff: Identifying Locations on Earth by Diffusing in the Hilbert SpaceZhangyu Wang, Zeping Liu, Jielu Zhang, Zhongliang Zhou 等NeurIPS 2025 · 被引用 7 次
它引用的顶会 Paper13
- SOLOv2: Dynamic and Fast Instance SegmentationXinlong Wang, Rufeng Zhang, Tao Kong, Lei Li 等NeurIPS 2020 · 被引用 1,193 次
- TransFG: A Transformer Architecture for Fine-Grained RecognitionJu He, Jieneng Chen, Shuai Liu, Adam Kortylewski 等AAAI 2022 · 被引用 529 次
- Presence-Only Geographical Priors for Fine-Grained Image ClassificationOisin Mac Aodha, Elijah Cole, Pietro PeronaICCV 2019 · 被引用 206 次
- Channel Interaction Networks for Fine-Grained Image CategorizationYu Gao, Xintong Han, Xun Wang, Weilin Huang 等AAAI 2020 · 被引用 177 次
- Multi-Scale Representation Learning for Spatial Feature Distributions using Grid CellsGengchen Mai, Krzysztof Janowicz, Bo Yan, Rui Zhu 等ICLR 2020 · 被引用 161 次
相关 Paper
- Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought ReasoningHulingxiao He, Zijun Geng, Yuxin PengICLR 2026 · 被引用 12 次
- Delving into Multimodal Prompting for Fine-Grained Visual ClassificationXin Jiang, Hao Tang, Junyao Gao, Xiaoyu Du 等AAAI 2024 · 被引用 71 次
- Dynamic Position-aware Network for Fine-grained Image RecognitionShijie Wang, Haojie Li, Zhihui Wang, Wanli OuyangAAAI 2021 · 被引用 36 次
- Grafit: Learning fine-grained image representations with coarse labelsHugo Touvron, Alexandre Sablayrolles, Matthijs Douze, Matthieu Cord 等ICCV 2021 · 被引用 79 次
- DynaMixer: A Vision MLP Architecture with Dynamic MixingZiyu Wang, Wenhao Jiang, Yiming Zhu, Li Yuan 等ICML 2022 · 被引用 55 次
