Learning Canonical 3D Object Representation for Fine-Grained Recognition
Sunghun Joung, Seungryong Kim, Minsu Kim, Ig-Jae Kim, Kwanghoon Sohn
摘要
We propose a novel framework for fine-grained object recognition that learns to recover object variation in 3D space from a single image, trained on an image collection without using any ground-truth 3D annotation. We accomplish this by representing an object as a composition of 3D shape and its appearance, while eliminating the effect of camera viewpoint, in a canonical configuration. Unlike conventional methods modeling spatial variation in 2D images only, our method is capable of reconfiguring the appearance feature in a canonical 3D space, thus enabling the subsequent object classifier to be invariant under 3D geometric variation. Our representation also allows us to go beyond existing methods, by incorporating 3D shape variation as an additional cue for object recognition. To learn the model without ground-truth 3D annotation, we deploy a differentiable renderer in an analysis-by-synthesis frame- work. By incorporating 3D shape and appearance jointly in a deep representation, our method learns the discriminative representation of the object and achieves competitive performance on fine-grained image recognition and vehicle re-identification. We also demonstrate that the performance of 3D shape reconstruction is improved by learning fine-grained shape deformation in a boosting manner.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- SIM-Trans: Structure Information Modeling Transformer for Fine-grained Visual CategorizationHongbo Sun, Xiangteng He, Yuxin PengACM MM 2022 · 被引用 128 次
- Not All Birds Look The Same: Identity-Preserving Generation For BirdsAaron Sun, Oindrila Saha, Subhransu MajiCVPR 2026
它引用的顶会 Paper14
- Soft Rasterizer: A Differentiable Renderer for Image-Based 3D ReasoningShichen Liu, Weikai Chen, Tianye Li, Hao LiICCV 2019 · 被引用 789 次
- Pixel2Mesh++: Multi-View 3D Mesh Generation via DeformationChao Wen, Yinda Zhang, Zhuwen Li, Yanwei FuICCV 2019 · 被引用 279 次
- A Dual-Path Model With Adaptive Attention for Vehicle Re-IdentificationPirazh Khorramshahi, Amit Kumar, Neehar Peri, Sai Saketh Rambhatla 等ICCV 2019 · 被引用 236 次
- Cross-X Learning for Fine-Grained Visual CategorizationWei Luo, Xitong Yang, Xianjie Mo, Yuheng Lu 等ICCV 2019 · 被引用 234 次
- Selective Sparse Sampling for Fine-Grained Image RecognitionYao Ding, Yanzhao Zhou, Yi Zhu, Qixiang Ye 等ICCV 2019 · 被引用 227 次
相关 Paper
- Shelf-Supervised Mesh Prediction in the WildYufei Ye, Shubham Tulsiani, Abhinav GuptaCVPR 2021
- Canonical 3D Deformer Maps: Unifying parametric and non-parametric methods for dense weakly-supervised category reconstructionDavid Novotný, Roman Shapovalov, Andrea VedaldiNeurIPS 2020 · 被引用 9 次
- Disentangled3D: Learning a 3D Generative Model with Disentangled Geometry and Appearance from Monocular ImagesAyush Tewari, Mallikarjun B. R., Xingang Pan, Ohad Fried 等CVPR 2022 · 被引用 35 次
- Learning Partonomic 3D Reconstruction from Image CollectionsXiaoqian Ruan, Pei Yu, Dian Jia, Hyeonjeong Park 等CVPR 2025
- AutoRF: Learning 3D Object Radiance Fields from Single View ObservationsNorman Müller, Andrea Simonelli, Lorenzo Porzi, Samuel Rota Bulò 等CVPR 2022 · 被引用 45 次
