RealBirdID: Benchmarking Bird Species Identification in the Era of MLLMs
Logan Lawrence, Oindrila Saha, Rangel Daroya, Mustafa Chasmai, Wuao Liu, Max Hamilton, Aaron Sun, Seoyun Jeong, Fabien Delattre, Subhransu Maji, Grant Horn
Abstract
Fine-grained bird species identification in the wild is frequently unanswerable from a single image: key cues may be non-visual (vocalization, range, season), or obscured due to occlusion, camera angle, or low resolution. Yet today’s multimodal systems are typically judged on answerable, in-schema cases, encouraging confident guesses rather than principled abstention. We propose the RealBirdID benchmark: given an image of a bird, a system should either answer with a species or abstain with a concrete, evidence-based rational (e.g., “requires vocalization,” “out of range,” “view obstructed”). For each genus, the dataset includes a validation split composed of curated unanswerable examples with labeled rationales, paired with a companion set of clearly answerable instances. We find that (1) the species identification on the answerable set is challenging for a variety of open-source and proprietary models ( accuracy including GPT-5 and Gemini-2.5 Pro), (2) models with greater classification ability are not necessarily more calibrated to abstain from unanswerable examples, and (3) that MLLMs generally fail at providing correct reasons even when they do abstain. RealBirdID establishes a focused target for abstention-aware fine-grained recognition and a recipe for measuring progress.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7baad3ea-3f1f-4e40-9385-b5065edbfa90Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
Related papers
- Med-CMR: A Fine-Grained Benchmark Integrating Visual Evidence and Clinical Logic for Medical Complex Multimodal ReasoningHaozhen Gong, Xiaozhong Ji, Yuansen Liu, Wenbin Wu et al.CVPR 2026 · 15 citations
- Fine-Grained Multi Image Object Hallucination BenchmarkJoonki Min, Chaeyun Kim, Hyungwook Choi, Yejin Kim et al.CVPR 2026 · 1 citation
- When Seeing Is not Enough: Revealing the Limits of Active Reasoning in MLLMsHongcheng Liu, Pingjie Wang, Yuhao Wang, Siqu Ou et al.ACL 2026
- RefineBench: Evaluating Refinement Capability of Language Models via ChecklistsYoung-Jun Lee, Seungone Kim, Byung-Kwan Lee, Minkyeong Moon et al.ICLR 2026 · 13 citations
- GeoRC: A Benchmark for Geolocation Reasoning ChainsMohit Talreja, Joshua Diao, Jim James, Radu Casapu et al.ACL 2026 · 1 citation
