SAMDiff-TTR: SAM2-Augmented Diffusion with Test-time Refinement for Small-Sample Infrared-to-Visible Image Generation
Main Article Content
Abstract
Infrared-to-visible image generation provides an effective solution for augmenting visible-spectrum samples in small-sample scenarios and offers important support for cross-spectral information representation in electromagnetic sensing systems, thereby benefiting downstream visual perception tasks such as object detection in autonomous navigation and unmanned aerial vehicle (UAV) applications. However, existing methods often suffer from insufficient semantic consistency and weak generalization when only limited paired cross-modal data are available. To address this issue, we propose SAMDiff-TTR, a novel infrared-to-visible image generation framework that combines a diffusion-based generative model with semantic guidance from a pre-trained SAM2 model and adaptive test-time refinement (TTR). Specifically, a SAMSeg pre-training stage with a Multi-scale Feature Fusion Module (MFFM) is introduced to extract multi-scale semantic features from infrared images for visible-image generation. A denoising U-Net equipped with ResNet Attention (ResAtt) blocks is then employed to synthesize high-quality visible images. Furthermore, a Boundary- and Appearance-aware Test-Time Refinement (BATTR) strategy is developed to enhance robustness under limited-data conditions by jointly enforcing boundary preservation, semantic alignment, and visible-domain appearance constraints during inference without requiring paired visible references. Experiments on the DroneVehicle dataset demonstrate that SAMDiff-TTR achieves competitive image quality and consistently improves downstream object detection under small-sample settings, providing an effective cross-spectral data augmentation framework for infrared electromagnetic imaging and intelligent visual perception.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
Y. Xi, Y. Zhang, Y. Jiang, et al., “Infrared Small Target Detection Based on Entropy Variation Weighted Local Contrast Measure [J],” Remote Sensing, vol. 17, no. 8, 2025.
C. Zhao, W. Wang, Y. Yan, et al., “A Re-Identification Framework for Visible and Thermal-Infrared Aerial Remote Sensing Images with Large Differences of Elevation Angles [J],” Remote Sensing, vol. 17, no. 11, 2025.
N. Bhat, N. Saggu, Pragati, et al., “Generating Visible Spectrum Images from Thermal Infrared using Conditional Generative Adversarial Networks; proceedings of the 2020 5th International Conference on Communication and Electronics Systems (ICCES), F, 2020 [C].”
F. Luo, Y. Li, G. Zeng, et al., “Thermal Infrared Image Colorization for Nighttime Driving Scenes With Top-Down Guided Attention [J],” IEEE Transactions on Intelligent Transportation Systems, vol. 23, no. 9, pp. 15808–15823, 2022.
A. Anoosheh, T. Sattler, R. Timofte, et al., “Night-to-day image translation for retrieval-based localization; proceedings of the 2019 International conference on robotics and automation (ICRA), F, 2019 [C].”
J.-Y. Zhu, T. Park, P. Isola, et al., “Unpaired image-to-image translation using cycle-consistent adversarial networks; proceedings of the Proceedings of the IEEE international conference on computer vision, F, 2017 [C].”
F.-Y. Luo, Y.-J. Cao, K.-F. Yang, et al., “Memory-Guided Collaborative Attention for Nighttime Thermal Infrared Image Colorization of Traffic Scenes [J],” IEEE Transactions on Intelligent Transportation Systems, 2024.
J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, et al., “Deep Unsupervised Learning using Nonequilibrium Thermodynamics; proceedings of the Proceedings of the 32nd International Conference on Machine Learning, Lille, France, F 07–09 Jul, 2015 [C],” PMLR.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models; proceedings of the Proceedings of the 34th International Conference on Neural Information Processing Systems, Red Hook, NY, USA, F, 2020 [C],” Curran Associates Inc.
X. Zhang, Y. Li, F. Li, et al., “Ship-Go: SAR Ship Images Inpainting via instance-to-image Generative Diffusion Models [J],” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 207, pp. 203–217, 2024.
J. Song, H. Xu, G. Jiang, et al., “Frequency domain-based latent diffusion model for underwater image enhancement [J],” Pattern Recognition, vol. 160, p. 111198, 2025.
N. Ravi, V. Gabeur, Y.-T. Hu, et al., “SAM 2: Segment Anything in Images and Videos, F, 2024 [C].”
S. Cui, Y. Li, J. Li, et al., “Continual Test-Time Adaptation for Single Image Defocus Deblurring via Causal Siamese Networks [J],” Int J Comput Vision, vol. 133, no. 7, pp. 4134–4157, 2025.
Y. Sun, B. Cao, P. Zhu, et al., “Drone-Based RGB-Infrared Cross-Modality Vehicle Detection Via Uncertainty-Aware Learning [J],” IEEE Transactions on Circuits and Systems for Video Technology, vol. 32, no. 10, pp. 6700–6713, 2022.
Z. Wang, A. C. Bovik, H. R. Sheikh, et al., “Image quality assessment: from error visibility to structural similarity [J],” IEEE Transactions on Image Processing, vol. 13, no. 4, pp. 600–612, 2004.
R. Zhang, P. Isola, A. A. Efros, et al., “The unreasonable effectiveness of deep features as a perceptual metric; proceedings of the Proceedings of the IEEE conference on computer vision and pattern recognition, F, 2018 [C].”
M. Heusel, H. Ramsauer, T. Unterthiner, et al., “Gans trained by a two time-scale update rule converge to a local nash equilibrium [J],” Advances in neural information processing systems, vol. 30, 2017.
M. Everingham, S. M. A. Eslami, L. Van Gool, et al., “The pascal visual object classes challenge: A retrospective [J],” International journal of computer vision, vol. 111, pp. 98–136, 2015.
P. Isola, J.-Y. Zhu, T. Zhou, et al., “Image-to-Image Translation with Conditional Adversarial Networks; proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), F, 2017 [C].”
N. G. Nair and V. M. Patel, “T2V-DDPM: Thermal to Visible Face Translation using Denoising Diffusion Probabilistic Models; proceedings of the 2023 IEEE 17th International Conference on Automatic Face and Gesture Recognition (FG), F, 2023 [C].”
O. Özdenizci and R. Legenstein, “Restoring Vision in Adverse Weather Conditions With Patch-Based Denoising Diffusion Models [J],” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 8, pp. 10346–10357, 2023.
D. P. Kingma and J. Ba, “Adam: A Method for Stochastic Optimization [Z],” 2017.
L. Tian, Q. Shen, Z. Deng, et al., “Mask-Guided Cross-Modality Fusion Network for Visible-Infrared Vehicle Detection [J],” IEEE Signal Processing Letters, vol. 32, pp. 1815–1819, 2025.
S. Hu, F. Gao, X. Zhou, et al., “Hybrid Convolutional and Attention Network for Hyperspectral Image Denoising [J],” IEEE Geoscience and Remote Sensing Letters, vol. 21, pp. 1–5, 2024.
X. Qiu, R.-J. Zhu, Y. Chou, et al., “Gated attention coding for training high-performance and efficient spiking neural networks; proceedings of the Proceedings of the Thirty-Eighth AAAI Conference on Artificial Intelligence and Thirty-Sixth Conference on Innovative Applications of Artificial Intelligence and Fourteenth Symposium on Educational Advances in Artificial Intelligence, F, 2024 [C],” AAAI Press.
S. Xu, S. Zheng, W. Xu, et al., “HCF-Net: Hierarchical Context Fusion Network for Infrared Small Object Detection; proceedings of the 2024 IEEE International Conference on Multimedia and Expo (ICME), F, 2024 [C].”