A Theory-Inspired Framework for Few-Shot Cross-Modal Sketch Person Re-Identification
Yunpeng Gong, Yongjie Hou, Jiangming Shi, Kim Long Diep, Min Jiang
Abstract
Sketch-based person re-identification aims to match handdrawn sketches with RGB surveillance images, but remains challenging due to severe modality gaps and limited labeled data. To address this, we propose KTCAA, a theoretically inspired framework for few-shot cross-modal generalization. Drawing on generalization bounds, we identify two key factors affecting target risk: (1) domain discrepancy, reflecting the alignment difficulty between source and target distributions; and (2) perturbation invariance, measuring the model's robustness to modality shifts. Accordingly, we design: (1) Alignment Augmentation (AA), which applies localized sketch-style transformations to simulate target distributions and guide progressive alignment; and (2) Knowledge Transfer Catalyst (KTC), which enhances perturbation invariance by introducing worst-case modality perturbations and enforcing consistency. These modules are jointly optimized within a meta-learning paradigm that transfers alignment knowledge from data-abundant RGB domains to sketch scenarios. Experiments on multiple benchmarks show that KTCAA achieves state-of-theart performance, particularly under data-scarce conditions. The code will be available at https://github.com/ finger-monkey/REID_KTCAA .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9c13043d-5e7b-4fcb-87df-cdd2f88821dcCited by top-tier papers4
- Statistical Characteristic-Guided Denoising for Rapid High-Resolution Transmission Electron Microscopy ImagingHesong Li, Ziqi Wu, Ruiwen Shao, Ying FuCVPR 2026 · 4 citations
- VisionLaw: Inferring Interpretable Intrinsic Dynamics from Visual Observations via Bilevel OptimizationJiajing Lin, Shu Jiang, Qingyuan Zeng, Zhenzhong Wang et al.ICLR 2026 · 4 citations
- Re-evaluating Continual VQA: Toward Fair and Robust Evaluation for Multimodal Continual LearningZijian Gao, Zicheng Sun, Xingxing Zhang, Kele Xu et al.CVPR 2026
- Correspondence Cognitive Learning for Multi-Modal Object Re-IdentificationChao Su, Shuying Li, Ruitao Pu, Dezhong Peng et al.ICML 2026
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Learning with Twin Noisy Labels for Visible-Infrared Person Re-IdentificationMouxing Yang, Zhenyu Huang, Peng Hu, Taihao Li et al.CVPR 2022 · 248 citations
- Augmented Dual-Contrastive Aggregation Learning for Unsupervised Visible-Infrared Person Re-IdentificationBin Yang, Mang Ye, Jun Chen, Zesen WuACM MM 2022 · 113 citations
- Cross-Modality Perturbation Synergy Attack for Person Re-identificationYunpeng Gong, Zhun Zhong, Yansong Qu, Zhiming Luo et al.NeurIPS 2024 · 67 citations
- Shallow-Deep Collaborative Learning for Unsupervised Visible-Infrared Person Re-IdentificationBin Yang, Jun Chen, Mang YeCVPR 2024 · 52 citations
Related papers
- Cross-Category Subjectivity Generalization for Style-Adaptive Sketch Re-IDZechao Hu, Zhengwei Yang, Hao Li, Zheng Wang et al.ICCV 2025 · 1 citation
- SG-FSL: Cross-Domain Few-Shot Learning with Style-Decoupled Augmentation and Gradient-Conflict AdjustmentYunyu Zou, Yishu Liu, Jun Liang, Bingzhi ChenACM MM 2025
- Doodle Your Keypoints: Sketch-Based Few-Shot Keypoint DetectionSubhajit Maity, Ayan Kumar Bhunia, Subhadeep Koley, Pinaki Nath Chowdhury et al.ICCV 2025
- Meta Distribution Alignment for Generalizable Person Re-IdentificationHao Ni, Jingkuan Song, Xiaopeng Luo, Feng Zheng et al.CVPR 2022 · 77 citations
- PIRN: Prototypical-based Intra-modal Reconstruction with Normality Communication for Multi-modal Anomaly Detection.YITING LI, Xulei Yang, Jing Zhang, Sichao Tian et al.ICLR 2026
