Not All Birds Look The Same: Identity-Preserving Generation For Birds
Aaron Sun, Oindrila Saha, Subhransu Maji
Abstract
Since the advent of controllable image generation, increasingly rich modes of control have enabled greater customization and accessibility for everyday users.Zero-shot, identity-preserving models such as Insert Anything and OminiControl now support applications like virtual try-on without requiring additional fine-tuning.While these models may be fitting for humans and rigid everyday objects, they still have limitations for non-rigid or fine-grained categories. These domains often lack accessible, high-quality data—especially videos or multi-view observations of the same subject—making them difficult both to evaluate and to improve upon. Yet, such domains are essential for moving beyond content creation toward applications that demand accuracy and fine detail.Birds are an excellent domain for this task: they exhibit high diversity, require fine-grained cues for identification, and come in a wide variety of poses. We introduce the NABirds Look-Alikes (NABLA) dataset, consisting of 4,759 expert-curated image pairs. Together with 1,073 pairs collected from multi-image observations on iNaturalist and a small set of videos, this forms a benchmark for evaluating identity-preserving generation of birds.We show that state-of-the-art baselines fail to maintain identity on this dataset, and we demonstrate that training on images grouped by species, age, and sex---used as a proxy for identity---substantially improves performance on both seen and unseen species.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on20
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- 4DComplete: Non-Rigid Motion Estimation Beyond the Observable SurfaceYang Li, Hikari Takehara, Takafumi Taketomi, Bo Zheng et al.ICCV 2021 · 160 citations
Related papers
- WithAnyone: Toward Controllable and ID Consistent Image GenerationHengyuan Xu, Wei Cheng, Peng Xing, Yixiao Fang et al.ICLR 2026 · 12 citations
- Garments2Look: A Multi-Reference Dataset for High-Fidelity Outfit-Level Virtual Try-On with Clothing and AccessoriesJunyao Hu, Zhongwei Cheng, Waikeung Wong, Xingxing ZouCVPR 2026 · 4 citations
- RealBirdID: Benchmarking Bird Species Identification in the Era of MLLMsLogan Lawrence, Oindrila Saha, Rangel Daroya, Mustafa Chasmai et al.CVPR 2026
- Inpaint-Anywhere: Zero-Shot Multi-Identity Inpainting with Efficient Diffusion TransformerJunsheng Luan, Lei Zhao, Wei XingAAAI 2026
- CustAny: Customizing Anything from A Single ExampleLingjie Kong, Kai Wu, Chengming Xu, Xiaobin Hu et al.CVPR 2025
