Application of Video Swin Transformer in Modern Dance Choreography Style Generation and Innovation Assistance
Main Article Content
Abstract
Given the complex movements, lack of fixed structure, and highly variable emotional expressions of modern dance, existing models struggle to effectively identify and capture diverse stylistic features. Similar spatiotemporal representation challenges are also encountered in intelligent electromagnetic information processing and signal understanding tasks. This paper proposes a modern dance style modeling method based on the Video Swin Transformer. First, modern dance video data are collected, and skeletal information is extracted to achieve strict spatiotemporal synchronization and spatial alignment between RGB video frames and corresponding human skeletal keypoint sequences. Then, hierarchical spatiotemporal dance features are extracted using the Video Swin Transformer architecture, and a style attention mechanism is designed to construct interpretable style encoding vectors. Subsequently, visual dimensionality reduction is applied to the low-dimensional style space to analyze the distribution of dancer movement styles and establish style mappings. Finally, combined with a generative decoder and a style control interface, the style vectors are used for dance movement reconstruction and creative generation, completing a full cycle from analysis to innovation assistance. Experimental results demonstrate that the proposed model achieves an average classification accuracy of 80.5% across different styles, while the optimal Euclidean distance for style space separation reaches 1.42 with an overlap rate of only 3.8%, validating its effectiveness in style control and generation and providing useful insights for spatiotemporal representation learning in intelligent signal processing.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
Q. Zhou, M. Li, Q. Zeng, A. Aristidou, X. Zhang, L. Chen, et al., “Let’s all dance: Enhancing amateur dance motions,” Computational Visual Media, vol. 9, no. 3, pp. 531-550, 2023, doi: 10.1007/s41095-022-0292-6.
Y. Liang and F. Pan, “Interactive experience design of traditional dance in new media era based on action detection,” Computer-Aided Design and Applications, vol. 21, no. S7, pp. 241-255, 2024, doi: 10.14733/cadaps.2024.S7.241-255.
T. Kim, “Multi-modal beat alignment transformers for dance quality assessment framework,” Journal of Multimedia Information System, vol. 11, no. 2, pp. 149-156, 2024, doi: 10.33851/JMIS.2024.11.2.149.
W. Zhuang, C. Wang, J. Chai, Y. Wang, M. Shao, S. Xia, et al., “Mu-sic2dance: Dancenet for music-driven dance generation,” ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM), vol. 18, no. 2, pp. 1-21, 2022, doi: 10.1145/3485664.
X. Ju, “The Application of Deep Learning in Dance Movement Design,” International Journal of Computational Intelligence Systems, vol. 18, no. 1, pp. 1-23, 2025, doi: 10.1007/s44196-025-00907-3.
H. Jiang and Y. Yan, “Sensor based dance coherent action generation model using deep learning framework,” Scalable Computing: Practice and Experience, vol. 25, no. 2, pp. 1073-1090, 2024, doi: 10.12694/scpe.v25i2.2648.
M. Zhang, Z. Cai, L. Pan, F. Hong, X. Guo, L. Yang, et al., “Motiondiffuse: Text-driven human motion generation with diffusion model,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 6, pp. 4115-4128, 2024, doi: 10.48550/arXiv.2208.15001.
W. Anni and H. Jiao, “Investigating the implementation and innovative integration of virtual simulation technologies in dance pedagogy: a case study on the Chinese Nongle dance ritual virtual experiment,” Research in Dance Education, vol. 26, no. 2, pp. 277-289, 2025, doi: 10.1080/14647893.2025.2466248.
L. Ting and Y. Yi, “Choreography and Creation of Dance Form Integrating Biological Morphology Technology,” Journal of Commercial Biotechnology, vol. 30, no. 1, pp. 122-132, 2025, doi: 10.5912/jcb2493.
Z. Zhou, Y. Huo, G. Huang, A. Zeng, X. Chen, L. Huang, et al., “Qean: quaternion-enhanced attention network for visual dance generation,” The Visual Computer, vol. 41, no. 2, pp. 961-973, 2025, doi: 10.1007/s00371-024-03376-5.
M. R. Nogueira, P. Menezes, and J. Maçãs de Carvalho, “Exploring the impact of machine learning on dance performance: a systematic review,” International Journal of Performance Arts and Digital Media, vol. 20, no. 1, pp. 60-109, 2024, doi: 10.1080/14794713.2024.2338927.
A. Leh, D. Endres, and M. Hegele, “A primitive-based representation of dance: modulations by experience and perceptual validity,” Journal of Neurophysiology, vol. 130, no. 5, pp. 1214-1225, 2023, doi: 10.1152/jn.00161.2023.
A. Dewi S and D. da Ary, “Development of e-module material on recognizing the environment with dance,” Jurnal Penelitian Pendidikan IPA, vol. 10, no. 9, pp. 6835-6842, 2024, doi: 10.29303/jppipa.v10i9.8495.
N. Zhang, “A deep learning-based method for identifying differences in folk dance movements,” Journal of Computational Methods in Sciences and Engineering, vol. 25, no. 3, pp. 2334-2346, 2025, doi: 10.1177/14727978241312996.
S. S. A. Qızı, “EVOLUTION OF DANCE STYLES: FROM TRADITIONAL TO MODERN TRENDS,” Eurasian Journal of Social Sciences, Philosophy and Culture, vol. 4, no. 10, pp. 32-35, 2024, doi: 10.5281/zenodo.13907404.
N. C. Demian, “Reframing of contemporary dance,” Învăţământ, Cercetare, Creaţie, vol. 8, no. 1, pp. 65-70, 2022.
S. Alexanderson, R. Nagy, J. Beskow, and E. G. Henter, “Listen, denoise, action! audio-driven motion synthesis with diffusion models,” ACM Transactions on Graphics (TOG), vol. 42, no. 4, pp. 1-20, 2023, doi: 10.48550/arXiv.2211.09707.
F. Nazarieh, Z. Feng, M. Awais, W. Wang, and J. Kittler, “A survey of cross-modal visual content generation,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 34, no. 8, pp. 6814-6832, 2024, doi: 10.1109/TCSVT.2024.3351601.
T. Huang, X. Ben, C. Gong, B. Zhang, R. Yan, and Q. Wu, “Enhanced spatial-temporal salience for cross-view gait recognition,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 32, no. 10, pp. 6967-6980, 2022, doi: 10.1109/TCSVT.2022.3175959.
H. M. Hsu, Y. Wang, C. Y. Yang, J. Huang, U. L. H. Thuc, and K. Tim, “Learning temporal attention based keypoint-guided embedding for gait recognition,” IEEE Journal of Selected Topics in Signal Processing, vol. 17, no. 3, pp. 689-698, 2023, doi: 10.1109/JSTSP.2023.3271827.
S. Rahmah, A. H. Saragih, and E. Napitupulu, “Impact of Dance Education Learning Model Development (Icosrie) to Improve College Students’ cognitive Skills,” Russian law journal, vol. 11, no. 3, pp. 2406-2424, 2023, doi: 10.52783/rlj.v11i3.2097.