Dance and Martial Arts Fusion Motion Generation Technology Based on Multimodal Fusion and Transformer
Main Article Content
Abstract
This paper presents a multimodal motion generation framework integrating high-frequency textile strain sensing (TSS) arrays with a Spatially Modulated Co-Attention (SMCA) Transformer for real-time dance and martial arts fusion. A contrastive multimodal alignment (CMA) mechanism is employed to achieve robust cross-modal feature integration among strain sensor signals, audio rhythms, and semantic embeddings. The SMCA architecture decouples and fuses spatiotemporal dual-stream features, while differentiable biomechanical constraints—including bone length consistency and zero-moment-point balance—ensure physical plausibility of the generated motion trajectories. Experimental validation demonstrates a motion reconstruction accuracy (MRA) of 0.942–0.965, a bone length consistency loss reduced to 0.15, and a high-efficiency frame processing rate (FPR) of 124.6 fps. By framing the framework as a sensor-array-based signal acquisition and propagation system with physics-constrained output, the study provides an engineering-oriented methodology for real-time generation, spatiotemporal signal alignment, and wave-propagation-inspired trajectory analysis. The approach offers potential applications in immersive human-computer interaction, virtual performance systems, and digital cultural heritage preservation.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
L. Yin and X. Sun, “Textile-based sensors for human motion sensing: Recent developments and future perspectives,” Nanocomposites, vol. 11, no. 1, pp. 79-98, 2025, doi: 10.1080/20550324.2025.2477392.
R. Tchantchane, H. Zhou, S. Zhang, and G. Alici, “A review of hand gesture recognition systems based on noninvasive wearable sensors,” Advanced Intelligent Systems, vol. 5, no. 10, pp. 2300207-2300220, 2023, doi: 10.1002/aisy.202300207.
I. Loi, E. I. Zacharaki, and K. Moustakas, “Machine learning approaches for 3d motion synthesis and musculoskeletal dynamics estimation: A survey,” IEEE Transactions on Visualization and Computer Graphics, vol. 30, no. 8, pp. 5810-5829, 2023, doi: 10.1109/TVCG.2023.3308753.
K. Kritsis, “Computational models of multimodal interaction for music generation and music information retrieval,” no.1, pp. 243-256,2023, doi:10.12681/eadd/56523.
H. Yao, Z. Song, Y. Zhou, T. Ao, B. Chen, and L. Liu, “Moconvq: Unified physics-based motion control via scalable discrete representations,” ACM Transactions on Graphics (TOG), vol. 43, no. 4, pp. 1-21, 2024, doi: 10.1145/3658137.
K. Alomar, H. I. Aysel, and X. Cai, “CNNs, RNNs and Transformers in human action recognition: A survey and a hybrid model,” Artificial Intelligence Review, vol. 58, no. 12, pp. 1-44, 2025, doi: 10.1007/s10462-025-11388-3.
A. Jabbar, X. Li, and B. Omar, “A survey on generative adversarial networks: Variants, applications, and training,” ACM Computing Surveys (CSUR), vol. 54, no. 8, pp. 1-49, 2021, doi: 10.1145/3463475.
K. C. Alowonou and J. H. Han, “MSA-GCN: Exploiting Multi-Scale Temporal Dynamics With Adaptive Graph Convolution for Skeleton-Based Action Recognition,” IEEE Access, vol. 12, pp. 193552-193563, 2024, doi: 10.1109/ACCESS.2024.3520172.
T. Alam, F. Saidine, A. AL Faisal, A. Khan, and G. Hossain, “Smart-textile strain sensor for human joint monitoring,” Sensors and Actuators A: Physical, vol. 341, no. 1, pp. 113587-113599, 2022, doi: 10.2139/ssrn.4051542.
X. Yuan, A. Qi, H. Wu, J. Wang, Y. Guo, S. Li, et al., “Cross-modal feature alignment and fusion with contrastive learning in multimodal recommendation,” Knowledge-Based Systems, vol. 326, pp. 114020-114033, 2025, doi: 10.1016/j.knosys.2025.114020.
M. S. Junayed and M. B. Islam, “Consistent video inpainting using axial attention-based style transformer,” IEEE Transactions on Multimedia, vol. 25, no. 1, pp. 7494-7504, 2022, doi: 10.1109/TMM.2022.3222932.
M. Zeng and H. Zeng, “Research on Violin Audio Feature Recognition Based on Mel-frequency Cepstral Coefficient-based Feature Parameter Extraction,” Informatica, vol. 48, no. 19, pp. 1-12, 2024, doi: 10.31449/inf.v48i19.5966.
A. J. Marbun and F. R. Kodong, “Implementation of Mel-Frequency Cepstral Coefficient as Feature Extraction Method On Speech Audio Data,” Telematika: Jurnal Informatika dan Teknologi Informasi, vol. 21, no. 3, pp. 1-15, 2024, doi: 10.31315/telematika.v21i3.12339.
S. Deldari, H. Xue, A. Saeed, D. V. Smith, and F. D. Salim, “Cocoa: Cross modality contrastive learning for sensor data,” Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, vol. 6, no. 3, pp. 1-28, 2022, doi: 10.1145/3550316.
M. A. Manzoor, S. Albarri, Z. Xian, Z. Meng, P. Nakov, and S. Liang, “Multimodality representation learning: A survey on evolution, pretraining and its applications,” ACM Transactions on Multimedia Computing, Communications and Applications, vol. 20, no. 3, pp. 1-34, 2023, doi: 10.1145/3617833.
S. Rezo, M. C. Ferreira, J. J. Machado, and J. M. R. Tavares, “A multi-head attention-based transformer model for traffic flow forecasting with a comparative analysis to recurrent neural networks,” Expert Systems with Applications, vol. 202, no. 1, pp. 117275-117289, 2022, doi: 10.1016/j.eswa.2022.117275.