Dynamic MLP for Fine-Grained Image Classification by Leveraging Geographical and Temporal Information
Lingfeng Yang, Xiang Li, Renjie Song, Borui Zhao, Juntian Tao, Shihao Zhou, Jiajun Liang, Jian Yang
Abstract
Fine-grained image classification is a challenging computer vision task where various species share similar visual appearances, resulting in misclassification if merely based on visual clues. Therefore, it is helpful to leverage additional information, e.g., the locations and dates for data shooting, which can be easily accessible but rarely exploited. In this paper, we first demonstrate that existing multimodal methods fuse multiple features only on a single dimension, which essentially has insufficient help in feature discrimination. To fully explore the potential of multimodal information, we propose a dynamic MLP on top of the image representation, which interacts with multimodal features at a higher and broader dimension. The dynamic MLP is an efficient structure parameterized by the learned embeddings of variable locations and dates. It can be regarded as an adaptive nonlinear projection for generating more discriminative image representations in visual tasks. To our best knowledge, it is the first attempt to explore the idea of dynamic networks to exploit multimodal information in fine-grained image classification tasks. Extensive experiments demonstrate the effectiveness of our method. The t-SNE algorithm visually indicates that our technique improves the recognizability of image representations that are visually similar but with different categories. Furthermore, among published works across multiple fine-grained datasets, dynamic MLP consistently achieves SOTA results <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> https://paperswithcode.com/dataset/inaturalist and takes third place in the iNaturalist challenge at FGVC8 <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup> <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup> https://www.kaggle.com/c/inaturalist-2021/leaderboard. Code is available at httpsr//glthub.com/megvii-research/DynamicMLPForFinegrained.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ad68ca95-72c5-43f1-8948-e1fe73b7572eCited by top-tier papers8
- CSP: Self-Supervised Contrastive Spatial Pre-Training for Geospatial-Visual RepresentationsGengchen Mai, Ni Lao, Yutong He, Jiaming Song et al.ICML 2023 · 103 citations
- Spatial Implicit Neural Representations for Global-Scale Species MappingElijah Cole, Grant Van Horn, Christian Lange, Alexander Shepard et al.ICML 2023 · 70 citations
- MapFormer: Boosting Change Detection by Using Pre-change InformationMaximilian Bernhard, Niklas Strauß, Matthias SchubertICCV 2023 · 15 citations
- Fusion Meets Diverse Conditions: A High-Diversity Benchmark and Baseline for UAV-Based Multimodal Object Detection with Condition CuesChen Chen, Kangcheng Bin, Ting Hu, Jiahao Qi et al.ICCV 2025 · 8 citations
- LocDiff: Identifying Locations on Earth by Diffusing in the Hilbert SpaceZhangyu Wang, Zeping Liu, Jielu Zhang, Zhongliang Zhou et al.NeurIPS 2025 · 7 citations
Builds on13
- SOLOv2: Dynamic and Fast Instance SegmentationXinlong Wang, Rufeng Zhang, Tao Kong, Lei Li et al.NeurIPS 2020 · 1,193 citations
- TransFG: A Transformer Architecture for Fine-Grained RecognitionJu He, Jieneng Chen, Shuai Liu, Adam Kortylewski et al.AAAI 2022 · 529 citations
- Presence-Only Geographical Priors for Fine-Grained Image ClassificationOisin Mac Aodha, Elijah Cole, Pietro PeronaICCV 2019 · 206 citations
- Channel Interaction Networks for Fine-Grained Image CategorizationYu Gao, Xintong Han, Xun Wang, Weilin Huang et al.AAAI 2020 · 177 citations
- Multi-Scale Representation Learning for Spatial Feature Distributions using Grid CellsGengchen Mai, Krzysztof Janowicz, Bo Yan, Rui Zhu et al.ICLR 2020 · 161 citations
Related papers
- Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought ReasoningHulingxiao He, Zijun Geng, Yuxin PengICLR 2026 · 12 citations
- Delving into Multimodal Prompting for Fine-Grained Visual ClassificationXin Jiang, Hao Tang, Junyao Gao, Xiaoyu Du et al.AAAI 2024 · 71 citations
- Dynamic Position-aware Network for Fine-grained Image RecognitionShijie Wang, Haojie Li, Zhihui Wang, Wanli OuyangAAAI 2021 · 36 citations
- Grafit: Learning fine-grained image representations with coarse labelsHugo Touvron, Alexandre Sablayrolles, Matthijs Douze, Matthieu Cord et al.ICCV 2021 · 79 citations
- DynaMixer: A Vision MLP Architecture with Dynamic MixingZiyu Wang, Wenhao Jiang, Yiming Zhu, Li Yuan et al.ICML 2022 · 55 citations
