Partially Fake it Till you Make It: Mixing Real and Fake Thermal Images for Improved Object Detection
Francesco Bongini, Lorenzo Berlincioni, Marco Bertini, Alberto Del Bimbo
Abstract
In this paper we propose a novel data augmentation approach for visual content domains that have scarce training datasets, compositing synthetic 3D objects within real scenes. We show the performance of the proposed system in the context of object detection in thermal videos, a domain where i) training datasets are very limited compared to visible spectrum datasets and ii) creating full realistic synthetic scenes is extremely cumbersome and expensive due to the difficulty in modeling the thermal properties of the materials of the scene. We compare different augmentation strategies, including state of the art approaches obtained through RL techniques, the injection of simulated data and the employment of a generative model, and study how to best combine our proposed augmentation with these other techniques. Experimental results demonstrate the effectiveness of our approach, and our single-modality detector achieves state-of-the-art results on the FLIR ADAS dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 69a2edc1-5223-4e96-89fb-0cbfaca139b9Cited by top-tier papers3
- Data Generation Scheme for Thermal Modality with Edge-Guided Adversarial Conditional Diffusion ModelGuoqing Zhu, Honghu Pan, Qiang Wang, Chao Tian et al.ACM MM 2024 · 7 citations
- Multimodal Decomposed Distillation with Instance Alignment and Uncertainty Compensation for Thermal Object DetectionYanfeng Liu, Lefei ZhangACM MM 2025 · 2 citations
- Pseudo Visible Feature Fine-Grained Fusion for Thermal Object DetectionTing Li, Mao Ye, Tianwen Wu, Nianxin Li et al.CVPR 2025
Builds on2
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- Weakly Aligned Cross-Modal Learning for Multispectral Pedestrian DetectionLu Zhang, Xiangyu Zhu, Xiangyu Chen, Xu Yang et al.ICCV 2019 · 209 citations
Related papers
- 3D Copy-Paste: Physically Plausible Object Insertion for Monocular 3D DetectionYunhao Ge, Hong-Xing Yu, Cheng Zhao, Yuliang Guo et al.NeurIPS 2023 · 20 citations
- D3T: Distinctive Dual-Domain Teacher Zigzagging Across RGB-Thermal Gap for Domain-Adaptive Object DetectionDinh Phat Do, Taehoon Kim, Jaemin Na, Jiwon Kim et al.CVPR 2024
- TIRDet: Mono-Modality Thermal InfraRed Object Detection Based on Prior Thermal-To-Visible TranslationZeyu Wang, Fabien Colonnier, Jinghong Zheng, Jyotibdha Acharya et al.ACM MM 2023 · 28 citations
- CDUPatch: Color-Driven Universal Adversarial Patch Attack for Dual-Modal Visible-Infrared DetectorsJiahuan Long, Wen Yao, Tingsong Jiang, Jiacheng Hou et al.ACM MM 2025 · 7 citations
- Vi2ACT: Video-enhanced Cross-modal Co-learning with Representation Conditional Discriminator for Few-shot Human Activity RecognitionKang Xia, Wenzhong Li, Yimiao Shao, Sanglu LuACM MM 2024 · 1 citation
