Occlusion-Aware Image and Video Perception for Vulnerable Road User Detection: Methods, Benchmarks, and Deployment Challenges
Main Article Content
Abstract
Occlusion remains one of the main failure points in traffic-scene perception, and the errors it causes do not follow a single, predictable pattern. In a still frame, a camera may capture only a pedestrian’s head or upper torso. In video, a tracker can lose that person for several frames and assign a different identity when the person reappears. Cyclists and riders create a related problem because the body parts that remain visible may overlap visually with cars, trucks, or nearby road users. This review examines how image- and video-based methods handle such incomplete observations. The review covers 96 studies published between 2009 and 2025. These studies investigate visible-to-full-body localization, interactions among neighboring instances, feature completion, transformer-based detection, RGB-thermal fusion, tracking, and the less frequently studied problem of inferring fully hidden pedestrians. Because these tasks produce different outputs, we did not combine their results. Pedestrian and crowd detection, multispectral detection, hidden-target prediction, and multi-object tracking are discussed under their original evaluation protocols. CityPersons miss rate measures frame-level detection; MOT17 HOTA includes association across frames; hidden-pedestrian F1 measures prediction quality. We also recorded runtime, hardware constraints, model size, and sensor requirements when the source papers provided those details. The published evidence is strongest for pedestrians who remain partly visible. Much less evidence is available for cyclists, riders, and fully hidden targets. Many deployment claims are also difficult to assess because the papers do not report results for specific occlusion levels or provide enough detail about computational cost.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
R. M. Silva, G. F. Azevedo, M. V. V. Berto, J. R. Rocha, E. C. Fidelis, M. V. Nogueira, P. H. Lisboa, and T. A. Almeida, “Vulnerable road user detection and safety enhancement: A comprehensive survey,” Expert Systems with Applications, vol. 292, Art. no. 128529, 2025, https://doi.org/10.1016/j.eswa.2025.128529.
J. Janai, F. Guney, A. Behl, and A. Geiger, “Computer vision for autonomous vehicles: Problems, datasets and state of the art,” Foundations and Trends in Computer Graphics and Vision, vol. 12, no. 1-3, pp. 1–308, 2020.
D. Feng, C. Haase-Schuetz, L. Rosenbaum, H. Hertlein, C. Glaeser, F. Timm, W. Wiesbeck, and K. Dietmayer, “Deep multi-modal object detection and semantic segmentation for autonomous driving: Datasets, methods, and challenges,” IEEE Transactions on Intelligent Transportation Systems, vol. 22, no. 3, pp. 1341–1360, 2021.
S. Grigorescu, B. Trasnea, T. Cocias, and G. Macesanu, “A survey of deep learning techniques for autonomous driving,” Journal of Field Robotics, vol. 37, no. 3, pp. 362–386, 2020.
L. Liu, W. Ouyang, X. Wang, P. Fieguth, J. Chen, X. Liu, and M. Pietikainen, “Deep learning for generic object detection: A survey,” International Journal of Computer Vision, vol. 128, pp. 261–318, 2020, https://doi.org/10.1007/s11263-019-01247-4.
I. Hasan, S. Liao, J. Li, S. U. Akram, and L. Shao, “Pedestrian detection: Domain generalization, CNNs, transformers and beyond,” arXiv:2201.03176, 2022, https://doi.org/10.48550/arXiv.2201.03176.
Z.-Q. Zhao, P. Zheng, S.-T. Xu, and X. Wu, “Object detection with deep learning: A review,” IEEE Transactions on Neural Networks and Learning Systems, vol. 30, no. 11, pp. 3212–3232, 2019.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster R-CNN: Towards real-time object detection with region proposal networks,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 39, no. 6, pp. 1137–1149, 2017, https://doi.org/10.1109/TPAMI.2016.2577031.
T.-Y. Lin, P. Dollar, R. Girshick, K. He, B. Hariharan, and S. Belongie, “Feature pyramid networks for object detection,” in Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 936– 944, https://doi.org/10.1109/CVPR.2017.106.
T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollar, “Focal loss for dense object detection,” in Proc. IEEE International Conference on Computer Vision (ICCV), 2017, pp. 2999–3007, https://doi.org/10.1109/ICCV.2017.324.
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in Proc. European Conference on Computer Vision (ECCV), 2020, pp. 213–229, https://doi.org/10.1007/978-3-030-58452-8_13.
S. Zhang, R. Benenson, and B. Schiele, “CityPersons: A diverse dataset for pedestrian detection,” in Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 4457–4465, https://doi.org/10.1109/CVPR.2017.474.
S. Zhang, J. Yang, and B. Schiele, “Occlusion-aware R-CNN: Detecting pedestrians in a crowd,” in Proc. European Conference on Computer Vision (ECCV), 2018, pp. 657–674, https://doi.org/10.1007/978-3-030-01219-9_39.
Y. Pang, J. Xie, M. H. Khan, R. M. Anwer, F. S. Khan, and L. Shao, “Mask-guided attention network for occluded pedestrian detection,” in Proc. IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 4966–4974, https://doi.org/10.1109/ICCV.2019.00507.
S. Shao, Z. Zhao, B. Li, T. Xiao, G. Yu, X. Zhang, and J. Sun, “CrowdHuman: A benchmark for detecting human in a crowd,” arXiv:1805.00123, 2018.
S. Hwang, J. Park, N. Kim, Y. Choi, and I. S. Kweon, “Multispectral pedestrian detection: Benchmark dataset and baseline,” in Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015, pp. 1037–1045, https://doi.org/10.1109/CVPR.2015.7298706.
A. Gonzalez, Z. Fang, Y. Socarras, J. Serrat, D. Vazquez, J. Xu, and A. M. Lopez, “Pedestrian detection at day/night time with visible and FIR cameras: A comparison,” Sensors, vol. 16, no. 6, Art. no. 820, 2016.
X. Jia et al., “LLVIP: A visible-infrared paired dataset for low-light vision,” in Proc. IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), 2021.
A. Milan, L. Leal-Taixe, I. Reid, S. Roth, and K. Schindler, “MOT16: A benchmark for multi-object tracking,” arXiv:1603.00831, 2016.
P. Dendorfer et al., “MOT20: A benchmark for multi object tracking in crowded scenes,” arXiv:2003.09003, 2020.
P. Sun et al., “DanceTrack: Multi-object tracking in uniform appearance and diverse motion,” in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
K. Bernardin and R. Stiefelhagen, “Evaluating multiple object tracking performance: The CLEAR MOT metrics,” EURASIP Journal on Image and Video Processing, vol. 2008, Art. no. 246309, 2008.
E. Ristani, F. Solera, R. Zou, R. Cucchiara, and C. Tomasi, “Performance measures and a data set for multi-target, multi-camera tracking,” in Proc. European Conference on Computer Vision Workshops (ECCVW), 2016, pp. 17–35.
J. Luiten et al., “HOTA: A higher order metric for evaluating multi-object tracking,” International Journal of Computer Vision, vol. 129, pp. 548– 578, 2021.
K. Saleh, S. Szenasi, and Z. Vamossy, “Occlusion handling in generic object detection: A review,” in Proc. IEEE 19th World Symposium on Applied Machine Intelligence and Informatics (SAMI), 2021, pp. 477– 484, https://doi.org/10.1109/SAMI50585.2021.9378657.
J. Li et al., “Occlusion handling and multi-scale pedestrian detection based on deep learning: A review,” IEEE Access, 2022.
Z. Guan, Z. Wang, G. Zhang, L. Li, M. Zhang, Z. Shi, and N. Jiang, “Multi-object tracking review: Retrospective and emerging trend,” Artificial Intelligence Review, vol. 58, Art. no. 235, 2025, https://doi.org/10.1007/s10462-025-11212-y.
W. Luo, J. Xing, A. Milan, X. Zhang, W. Liu, and T.-K. Kim, “Multiple object tracking: A literature review,” Artificial Intelligence, vol. 293, Art. no. 103448, 2021.
G. Ciaparrone, F. L. Sanchez, S. Tabik, L. Troiano, R. Tagliaferri, and F. Herrera, “Deep learning in video multi-object tracking: A survey,” Neurocomputing, vol. 381, pp. 61–88, 2020.
A. Brunetti, D. Buongiorno, G. F. Trotta, and V. Bevilacqua, “Computer vision and deep learning techniques for pedestrian detection and tracking: A survey,” Neurocomputing, vol. 300, pp. 17–33, 2018.
M. Ye, J. Shen, G. Lin, T. Xiang, L. Shao, and S. C. H. Hoi, “Deep learning for person re-identification: A survey and outlook,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, no. 6, pp. 2872–2893, 2022.
W. Liu et al., “VLPD: Context-aware pedestrian detection via vision-language semantic self-supervision,” in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023.
Y. Zhang, T. Wang, and X. Zhang, “MOTRv2: Bootstrapping end-to-end multi-object tracking by pretrained object detectors,” in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 22056–22065, https://doi.org/10.1109/CVPR52729.2023.02112.
R. Gao and L. Wang, “MeMOTR: Long-term memory-augmented transformer for multi-object tracking,” in Proc. IEEE/CVF International Conference on Computer Vision (ICCV), 2023.
A. Wang, H. Chen, L. Liu, K. Chen, Z. Lin, J. Han, and G. Ding, “YOLOv10: Real-time end-to-end object detection,” in Advances in Neural Information Processing Systems (NeurIPS), vol. 37, 2024, https://doi.org/10.52202/079017-3429.
Y. Peng, H. Li, P. Wu, Y. Zhang, X. Sun, and F. Wu, “D-FINE: Rede-fine regression task of DETRs as fine-grained distribution refinement,” in Proc. International Conference on Learning Representations (ICLR), 2025. Available: https://openreview.net/forum?id=MFZjrTFE7h.
S. Zhang, M. Ji, Y. Li, and J. Yang, “Imagine the unseen: Occluded pedestrian detection via adversarial feature completion,” arXiv:2405.01311, 2024, https://doi.org/10.48550/arXiv.2405.01311.
A. N. Melo, S. M. Serrano, C. Salinas, and M. A. Sotelo, “Prediction of occluded pedestrians in road scenes using human-like reasoning: Insights from the OccluRoads dataset,” in Proc. IEEE Intelligent Vehicles Symposium (IV), 2025, pp. 385–391, https://doi.org/10.1109/IV64158.2025.11097510.
Y. Zhao et al., “DETRs beat YOLOs on real-time object detection,” in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024, pp. 16965–16974, https://doi.org/10.1109/CVPR52733.2024.01605.
A. Kirillov et al., “Segment anything,” in Proc. IEEE/CVF International Conference on Computer Vision (ICCV), 2023, pp. 4015–4026, https://doi.org/10.1109/ICCV51070.2023.00371.
S. Liu et al., “Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection,” in Proc. European Conference on Computer Vision (ECCV), 2024, pp. 38–55, https://doi.org/10.1007/978-3-031-72970-6_3.
T. Cheng, L. Song, Y. Ge, W. Liu, X. Wang, and Y. Shan, “YOLO-World: Real-time open-vocabulary object detection,” in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024, pp. 16901–16911, https://doi.org/10.1109/CVPR52733.2024.01599.
M. Minderer et al., “Simple open-vocabulary object detection with vision transformers,” in Proc. European Conference on Computer Vision (ECCV), 2022, pp. 728–755.
P. Dollar, C. Wojek, B. Schiele, and P. Perona, “Pedestrian detection: A benchmark,” in Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2009.
P. Dollar, C. Wojek, B. Schiele, and P. Perona, “Pedestrian detection: An evaluation of the state of the art,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 34, no. 4, pp. 743–761, 2012.
S. Tang, Y. Zhou, J. Li, C. Liu, and J. Shi, “Attention-guided sample-based feature enhancement network for crowded pedestrian detection using vision sensors,” Sensors, vol. 24, no. 19, Art. no. 6350, 2024.
Y. Xing, S. Yang, S. Wang, S. Zhang, G. Liang, X. Zhang, and Y. Zhang, “MS-DETR: Multispectral pedestrian detection transformer with loosely coupled fusion and modality-balanced optimization,” IEEE Transactions on Intelligent Transportation Systems, vol. 25, pp. 20628–20642, 2024, https://doi.org/10.1109/TITS.2024.3450584.
J. Gao, Y. Wang, K.-H. Yap, K. Garg, and B. S. Han, “OccluTrack: Rethinking awareness of occlusion for enhancing multiple pedestrian tracking,” IEEE Transactions on Intelligent Transportation Systems, vol. 26, no. 7, pp. 9852–9866, 2025, https://doi.org/10.1109/TITS.2025.3565334.
M. Braun, S. Krebs, F. Flohr, and D. M. Gavrila, “The EuroCity Persons dataset: A novel benchmark for object detection,” arXiv:1805.07193, 2018.
X. Li, F. Flohr, Y. Yang, H. Xiong, M. Braun, S. Pan, K. Li, and D. M. Gavrila, “A new benchmark for vision-based cyclist detection,” in Proc. IEEE Intelligent Vehicles Symposium (IV), 2016.
A. Geiger, P. Lenz, and R. Urtasun, “Are we ready for autonomous driving? The KITTI vision benchmark suite,” in Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2012, pp. 3354– 3361, https://doi.org/10.1109/CVPR.2012.6248074.
M. Cordts et al., “The Cityscapes dataset for semantic urban scene understanding,” in Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 3213–3223.
F. Yu et al., “BDD100K: A diverse driving dataset for heterogeneous multitask learning,” in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 2633–2642, https://doi.org/10.1109/CVPR42600.2020.00271.
H. Caesar et al., “nuScenes: A multimodal dataset for autonomous driving,” in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 11618–11628, https://doi.org/10.1109/CVPR42600.2020.01164.
P. Sun et al., “Scalability in perception for autonomous driving: Waymo Open Dataset,” in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 2443–2451, https://doi.org/10.1109/CVPR42600.2020.00252.
X. Wang et al., “Repulsion loss: Detecting pedestrians in a crowd,” in Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 7774–7783, https://doi.org/10.1109/CVPR.2018.00811.
S. Liu, D. Huang, and Y. Wang, “Adaptive NMS: Refining pedestrian detection in a crowd,” in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 6452–6461, https://doi.org/10.1109/CVPR.2019.00662.
X. Chu, A. Zheng, X. Zhang, and J. Sun, “Detection in crowded scenes: One proposal, multiple predictions,” in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 12211– 12220, https://doi.org/10.1109/CVPR42600.2020.01223.
N. Bodla, B. Singh, R. Chellappa, and L. S. Davis, “Soft-NMS: Improving object detection with one line of code,” in Proc. IEEE International Conference on Computer Vision (ICCV), 2017, pp. 5561–5569.
J. Hosang, R. Benenson, and B. Schiele, “Learning non-maximum suppression,” in Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 4507–4515.
H. Hu, J. Gu, Z. Zhang, J. Dai, and Y. Wei, “Relation networks for object detection,” in Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 3588–3597.
Y. Zhang et al., “ByteTrack: Multi-object tracking by associating every detection box,” in Proc. European Conference on Computer Vision (ECCV), 2022, pp. 1–21, https://doi.org/10.1007/978-3-031-20047-2_1.
F. Zeng et al., “MOTR: End-to-end multiple-object tracking with transformer,” in Proc. European Conference on Computer Vision (ECCV), 2022.
S. Zhang et al., “WiderPerson: A diverse dataset for dense pedestrian detection in the wild,” arXiv:1909.12118, 2019.
M.-F. Chang et al., “Argoverse: 3D tracking and forecasting with rich maps,” in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 8740–8749.
A. Rasouli, I. Kotseruba, T. Kunic, and J. K. Tsotsos, “Are they going to cross? A benchmark dataset and baseline for pedestrian crosswalk behavior,” in Proc. IEEE International Conference on Computer Vision Workshops (ICCVW), 2017.
A. Rasouli, I. Kotseruba, and J. K. Tsotsos, “PIE: A large-scale dataset and models for pedestrian intention estimation and trajectory prediction,” in Proc. IEEE/CVF International Conference on Computer Vision (ICCV), 2019.
Y. Tian, P. Luo, X. Wang, and X. Tang, “DeepParts: Occlusion handling in deep learning based pedestrian detection,” in Proc. IEEE International Conference on Computer Vision (ICCV), 2015, pp. 3258–3266.
C. Zhou and J. Yuan, “Bi-box regression for pedestrian detection and occlusion estimation,” in Proc. European Conference on Computer Vision (ECCV), 2018, pp. 138–154, https://doi.org/10.1007/978-3-030-01246-5_9.
C. Chi et al., “Pedestrian detection via visible-to-full body regression,” arXiv:2104.03106, 2021.
X. Song, B. Chen, P. Li, B. Wang, and H. Zhang, “PRNet++: Learning towards generalized occluded pedestrian detection via progressive refinement network,” Neurocomputing, vol. 482, pp. 98–115, 2022, https://doi.org/10.1016/j.neucom.2022.01.056.
K. N. A. Shastry, J. Chaudhari, D. Thapar, A. Nigam, and C. Arora, “Parts-based attention for highly occluded pedestrian detection with transformers,” in Proc. IEEE International Conference on Image Processing (ICIP), 2023.
J. Kim et al., “DAGN: Deformable attention-guided network for occluded pedestrian detection,” Applied Sciences, vol. 11, no. 13, 2021.
G. Lin, Z. Bao, Z. Huang, Z. Li, W.-S. Zheng, and Y. Chen, “A multi-level relation-aware transformer model for occluded person re-identification,” Neural Networks, vol. 177, Art. no. 106382, 2024.
W. Liu et al., “High-level semantic feature detection: A new perspective for pedestrian detection,” in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019.
Q. Li, Y. Bi, R. Cai, and J. Li, “Occluded pedestrian detection through bi-center prediction in anchor-free network,” Neurocomputing, vol. 507, pp. 199–207, 2022.
X. Zhu et al., “Deformable DETR: Deformable transformers for end-to-end object detection,” in Proc. International Conference on Learning Representations (ICLR), 2021.
J. Zhang, K. Xia, Z. Huang, S. Wang, and R. G. Akindele, “OBhunter: An ensemble spectral-angular based transformer network for occlusion detection,” Expert Systems with Applications, vol. 248, Art. no. 123324, 2024, https://doi.org/10.1016/j.eswa.2024.123324.
P. Cong et al., “STCrowd: A multimodal dataset for pedestrian perception in crowded scenes,” in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
P. Peng, T. Xu, B. Huang, and J. Li, “HAFNet: Hierarchical attentive fusion network for multispectral pedestrian detection,” Remote Sensing, vol. 15, no. 8, Art. no. 2041, 2023, https://doi.org/10.3390/rs15082041.
J. Cao, X. Weng, R. Khirodkar, J. Pang, and K. Kitani, “Observation-centric SORT: Rethinking SORT for robust multi-object tracking,” in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 9686–9696.
Y. Du et al., “StrongSORT: Make DeepSORT great again,” IEEE Transactions on Multimedia, vol. 25, pp. 8725–8737, 2023.
J. Liu, S. Zhang, S. Wang, and D. N. Metaxas, “Multispectral deep neural networks for pedestrian detection,” in Proc. British Machine Vision Conference (BMVC), 2016.
D. Xu, W. Ouyang, E. Ricci, X. Wang, and N. Sebe, “Learning cross-modal deep representations for robust pedestrian detection,” in Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 5363–5371.
L. Zhang, Z. Liu, X. Zhu, Z. Song, X. Yang, Z. Lei, and H. Qiao, “Weakly aligned cross-modal learning for multispectral pedestrian detection,” in Proc. IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 5127–5137.
K. Zhou, L. Chen, and X. Cao, “Improving multispectral pedestrian detection by addressing modality imbalance problems,” in Proc. European Conference on Computer Vision (ECCV), 2020.
A. Bewley, Z. Ge, L. Ott, F. Ramos, and B. Upcroft, “Simple online and realtime tracking,” in Proc. IEEE International Conference on Image Processing (ICIP), 2016, pp. 3464–3468.
N. Wojke, A. Bewley, and D. Paulus, “Simple online and realtime tracking with a deep association metric,” in Proc. IEEE International Conference on Image Processing (ICIP), 2017, pp. 3645–3649.
P. Bergmann, T. Meinhardt, and L. Leal-Taixe, “Tracking without bells and whistles,” in Proc. IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 941–951.
X. Zhou, V. Koltun, and P. Krahenbuhl, “Tracking objects as points,” in Proc. European Conference on Computer Vision (ECCV), 2020, pp. 474– 490.
Y. Zhang, C. Wang, X. Wang, W. Zeng, and W. Liu, “FairMOT: On the fairness of detection and re-identification in multiple object tracking,” International Journal of Computer Vision, vol. 129, pp. 3069–3087, 2021, https://doi.org/10.1007/s11263-021-01513-4.
T. Meinhardt, A. Kirillov, L. Leal-Taixe, and C. Feichtenhofer, “Track-Former: Multi-object tracking with transformers,” in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 8844–8854.
F. Seidenschwarz, G. Braso, V. C. Frias, I. Possegger, and H. Bischof, “Simple cues lead to a strong multi-object tracker,” in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 13813–13823.
J. Redmon and A. Farhadi, “YOLOv3: An incremental improvement,” arXiv:1804.02767, 2018, https://doi.org/10.48550/arXiv.1804.02767.
Z. Ge, S. Liu, F. Wang, Z. Li, and J. Sun, “YOLOX: Exceeding YOLO series in 2021,” arXiv:2107.08430, 2021, https://doi.org/10.48550/arXiv.2107.08430.
C.-Y. Wang, A. Bochkovskiy, and H.-Y. M. Liao, “YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors,” in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 7464–7475.