Bipartite Mode Matching for Vision Training Set Search from a Hierarchical Data Server
Yue Yao, Ruining Yang, Tom Gedeon
Abstract
We explore a situation in which the target domain is accessible, but real-time data annotation is not feasible. Instead, we would like to construct an alternative training set from a large-scale data server so that a competitive model can be obtained. For this problem, because the target domain usually exhibits distinct modes (i.e., semantic clusters representing data distribution), if the training set does not contain these target modes, the model performance would be compromised. While prior existing works improve algorithms iteratively, our research explores the often-overlooked potential of optimizing the structure of the data server. Inspired by the hierarchical nature of web search engines, we introduce a hierarchical data server, together with a bipartite mode matching algorithm (BMM) to align source and target modes. For each target mode, we look in the server data tree for the best mode match, which might be large or small in size. Through bipartite matching, we aim for all target modes to be optimally matched with source modes in a one-on-one fashion. Compared with existing training set search algorithms, we show that the matched server modes constitute training sets that have consistently smaller domain gaps with the target domain across object re-identification (re-ID) and detection tasks. Consequently, models trained on our searched training sets have higher accuracy than those trained otherwise. BMM allows data-centric unsupervised domain adaptation (UDA) orthogonal to existing model-centric UDA methods. By combining the BMM with existing UDA methods like pseudo-labeling, further improvement is observed.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a567e496-ea7e-4e64-a9f8-e37922dfd3c5Builds on11
- Objects365: A Large-Scale, High-Quality Dataset for Object DetectionShuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng et al.ICCV 2019 · 1,018 citations
- Mutual Mean-Teaching: Pseudo Label Refinery for Unsupervised Domain Adaptation on Person Re-identificationYixiao Ge, Dapeng Chen, Hongsheng LiICLR 2020 · 651 citations
- Cross-Domain Adaptive Teacher for Object DetectionYu-Jhe Li, Xiaoliang Dai, Chih-Yao Ma, Yen-Cheng Liu et al.CVPR 2022 · 215 citations
- Too Large; Data Reduction for Vision-Language Pre-TrainingAlex Jinpeng Wang, Kevin Qinghong Lin, David Junhao Zhang, Stan Weixian Lei et al.ICCV 2023 · 35 citations
- Alice Benchmarks: Connecting Real World Re-Identification with the SyntheticXiaoxiao Sun, Yue Yao, Shengjin Wang, Hongdong Li et al.ICLR 2024 · 6 citations
Related papers
- Large-scale Training Data Search for Object Re-identificationYue Yao, Tom Gedeon, Liang ZhengCVPR 2023
- Black-box Unsupervised Domain Adaptation with Bi-directional Atkinson-Shiffrin MemoryJingyi Zhang, Jiaxing Huang, Xueying Jiang, Shijian LuICCV 2023 · 24 citations
- Dual Bipartite Graph Learning: A General Approach for Domain Adaptive Object DetectionChaoqi Chen, Jiongcheng Li, Zebiao Zheng, Yue Huang et al.ICCV 2021 · 65 citations
- Category Dictionary Guided Unsupervised Domain Adaptation for Object DetectionShuai Li, Jianqiang Huang, Xian-Sheng Hua, Lei ZhangAAAI 2021 · 47 citations
- Moment Matching for Multi-Source Domain AdaptationXingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang et al.ICCV 2019 · 2,239 citations
