Detection of Climbing Rod Motion Changes Based on 3D Pose Estimation

Main Article Content

H. P. Zheng
Y. D. Gan
C. Q. Y. Zhang
S. Liang
S. X. Zhao
H. J. Zhang

Abstract

3D human pose estimation has important engineering application value in industrial scenarios such as safety monitoring of power grid pole-climbing operations and standardized action training. Existing skeleton-based human behavior recognition methods still have notable limitations in modeling long-term temporal dependencies, capturing multi-scale dynamic motion patterns, and resolving geometric ambiguities in 2D-to-3D pose lifting. This paper proposes a graph-based network integrating multi-scale temporal convolution with spatial attention, and an interactive spatio-temporal Transformer, and develops an end-to-end web-based recognition system. Experiments on public benchmark datasets fully verify the performance of the proposed methods, which can provide technical support for skeleton-based action recognition and deployment of intelligent perception systems in industrial scenarios.

Downloads

Download data is not yet available.

Article Details

How to Cite
Zheng, H. P., Gan, Y. D., Zhang, C. Q. Y., Liang, S., Zhao, S. X., & Zhang, H. J. (2026). Detection of Climbing Rod Motion Changes Based on 3D Pose Estimation. Advanced Electromagnetics, 15(3), 10028–10038. https://doi.org/10.7716/aem.v15i3.4201
Section
Research Articles

References

H. Wang, M. Chen, X. Xiong, et al., “Behavior Recognition Technology of Power Operator Based on OpenPose and AT-STGCN,” Journal of Sichuan University of Science & Engineering (Natural Science Edition), vol. 36, no. 04, pp. 61–70, 2023.

Z. Cao, G. Hidalgo, T. Simon, et al., “OpenPose: realtime multi-Person 2d pose estimation using part affinity fields,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 43, no. 1, pp. 172–186, 2021.

S. Jin, W. Liu, E. Xie, et al., “Differentiable hierarchical graph grouping for multi-person pose estimation,” in Proceedings of the European Conference on Computer Vision, Glasgow, UK, 2020, pp. 718–734.

G. Pavlakos, X. Zhou, K. G. Derpanis, et al., “Coarse-to-fine volumetric prediction for single-image 3D human pose,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 7025–7034.

D. Mehta, S. Sridhar, O. Sotnychenko, et al., “Vnect: real-time 3d human pose estimation with a single rgb camera,” ACM Transactions on Graphics, vol. 36, no. 4, pp. 1–14, 2017.

X. Sun, B. Xiao, F. Wei, et al., “Integral human pose regression,” in European Conference on Computer Vision (ECCV), 2018, pp. 529–545.

J. Martinez, R. Hossain, J. Romero, et al., “A simple yet effective baseline for 3d human pose estimation,” in IEEE International Conference on Computer Vision (ICCV), 2017, pp. 2659–2668.

C.-H. Chen and D. Ramanan, “3d human pose estimation = 2d pose estimation + matching,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 7035–7043.

H. Yang, H. Liu, Y. Zhang, et al., “Parallel Multi-scale Spatio-temporal Graph Convolutional Network for 3D Human Pose Estimation,” Journal of Software, vol. 36, no. 5, pp. 2151–2166, 2024.

D. Tome, C. Russell, and L. Agapito, “Lifting from the deep: convolutional 3d pose estimation from a single image,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 2500–2509.

C.-H. Chen, A. Tyagi, A. Agrawal, et al., “Unsupervised 3d pose estimation with geometric self-supervision,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 5707–5717.

B. X. Nie, P. Wei, and S. C. Zhu, “Monocular 3d human pose estimation by predicting depth on joints,” in IEEE International Conference on Computer Vision (ICCV), 2017, pp. 3447–3455.

L. Zhao, X. Peng, Y. Tian, et al., “Semantic graph convolutional networks for 3d human pose regression,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 3420–3430.

D. Zhao and M. Zhi, “Spatial multiple-temporal graph convolutional neural network for human action recognition,” Journal of Frontiers of Computer Science & Technology, vol. 17, no. 3, p. 719, 2023.

J. Liu, G. Wang, K. Hu, et al., “Global context-aware attention LSTM networks for 3D action recognition,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 3671–3680.

Y. Wang, M. Bao, and X. Liu, “3D Human Pose Estimation Based on Transformer and Graph Convolutional Network,” Chinese Journal of Sensors and Actuators, vol. 38, no. 9, pp. 1624–1630, 2025.

R. Zhao, K. Wang, H. Su, et al., “Bayesian graph convolution LSTM for skeleton based action recognition,” in IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 6881–6891.

C. Si, W. Chen, W. Wang, et al., “An attention enhanced graph convolutional LSTM network for skeleton-based action recognition,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 1227– 1236.

T. S. Kim and A. Reiter, “Interpretable 3D human action analysis with temporal convolutional networks,” in IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2017, pp. 1623–1631.

P. Wang, W. Li, C. Li, et al., “Action recognition based on joint trajectory maps with convolutional neural networks,” Knowledge-Based Systems, vol. 158, pp. 43–53, 2018.

F. Yang, Y. Wu, et al., “Make skeleton-based action recognition model smaller, faster and better,” in ACM International Conference on Multimedia in Asia, 2019, pp. 1–6.

L. Shi, Y. Zhang, J. Cheng, et al., “Two-stream adaptive graph convolutional networks for skeleton-based action recognition,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 12026–12035.

Z. Liu, H. Zhang, Z. Chen, et al., “Disentangling and unifying graph convolutions for skeleton-based action recognition,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 143–152.

F. Ye, S. Pu, Q. Zhong, et al., “Dynamic GCN: context-enriched topology learning for skeleton-based action recognition,” in ACM International Conference on Multimedia, 2020, pp. 55–63.

W. Peng, J. Shi, and G. Zhao, “Spatial temporal graph deconvolutional network for skeleton-based human action recognition,” IEEE signal processing letters, vol. 28, pp. 244–248, 2021.

Z. Chen, S. Li, B. Yang, et al., “Multi-scale spatial temporal graph convolutional network for skeleton-based action recognition,” Proceedings of the AAAI conference on artificial intelligence, vol. 35, no. 2, pp. 1113–1122, 2021.

J. Lee, M. Lee, D. Lee, et al., “Hierarchically decomposed graph convolutional networks for skeleton-based action recognition,” in IEEE/CVF International Conference on Computer Vision (ICCV), 2023, pp. 10444–10453.