Construction of a Multimodal Learning-Based Emotion Computing Model for Dance Movements, Clothing, and Aesthetic Education
Main Article Content
Abstract
Modeling aesthetic perception from dynamic visual information requires simultaneous understanding of human motion and the physical behavior of textile materials, making multimodal feature fusion a critical challenge in intelligent perception systems. Similar information integration strategies have also attracted increasing attention in electromagnetic sensing and heterogeneous signal interpretation, where multiple data sources must be jointly analyzed to characterize complex targets. This study proposes a multimodal emotion computing framework (DC-AEM) that combines dance kinematics and garment dynamics for computational aesthetic assessment. The framework employs a GCN-LSTM network to encode spatio-temporal skeletal motion and constructs a dedicated costume representation by integrating color distribution, texture descriptors, shape characteristics, and optical-flow-based dynamic features through a convolutional architecture. A cross-modal attention mechanism is introduced to adaptively associate movement patterns with context-dependent fabric behavior, generating a unified representation for predicting gracefulness, power, and fluidity. Experimental evaluation on a ballet performance dataset demonstrates that the proposed model significantly outperforms movement-only and simple fusion baselines, achieving superior correlation and lower prediction error across all aesthetic dimensions. The results confirm that dynamic textile characteristics provide essential complementary information for affective perception and establish an effective multimodal analysis strategy that may offer methodological reference for intelligent visual sensing, electromagnetic perception systems, and data-driven costume design.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
C. Hou, “The Fusion of Modern Dance and Traditional Dance: Balancing Innovation and Heritage,” Journal of Education, Humanities, and Social Research, vol. 2, no. 1, pp. 137-144, 2025, doi: 10.71222/skerj284.
R. Huang, L. Zhang, and Y. Li, “Transforming dance education in China: Enhancing sustainable development and cultural preservation,” Research in Dance Education, pp. 1-29, 2025, doi: 10.1080/14647893.2025.2524151.
S. Banes, “Terpsichore in Sneakers: Post-Modern Dance,” Middletown, CT, USA: Wesleyan University Press; 1987.
A. Carter and J. O’Shea, “The Routledge Dance Studies Reader Second Edition,” London, UK: Routledge; 2010.
D. Barbieri, “Costume in Performance: Materiality, Culture, and the Body,” London, UK: Bloomsbury Publishing; 2017.
J. Imparato, “Relations between body and clothing in performance: Costume as an activator of bodily actions,” Studies in Costume & Performance, vol. 6, no. 2, pp. 171-184, 2021, doi: 10.1386/scp_00045_1.
B. Faber, “Color Psychology and Color Therapy,” New York, NY, USA: McGraw-Hill; 1950.
A. Kraut, “Choreographing Copyright: Race, Gender, and Intellectual Property Rights in American Dance,” New York, NY, USA: Oxford University Press; 2015.
J. Wang, Y. Chen, S. Hao, X. Peng, and L. Hu, “Deep learning for sensor-based activity recognition: A survey,” Pattern Recognition Letters, vol. 119, pp. 3-11, 2019, doi: 10.1016/j.patrec.2018.02.010.
R. W. Picard, “Affective Computing,” Cambridge, MA, USA: MIT Press; 2000.
I. Rallis, A. Voulodimos, N. Bakalos, E. Protopapadakis, N. Doulamis, and A. Doulamis, “Machine learning for intangible cultural heritage: A review of techniques on dance analysis,” Visual Computing for Cultural Heritage, pp. 103-119, 2020, doi: 10.1007/978-3-030-37191-3_6.
M. Joshi and S. Chakrabarty, “An extensive review of computational dance automation techniques and applications,” Proceedings of the Royal Society A, vol. 477, no. 2251, Art. no. 20210071, 2021, doi: 10.1098/rspa.2021.0071.
C. Zheng, W. Wu, C. Chen, T. Yang, S. Zhu, J. Shen, et al., “Deep learning-based human pose estimation: A survey,” ACM Computing Surveys, vol. 56, no. 1, pp. 1-37, 2023, doi: 10.1145/3603618.
L. V. Wang and Wu H-i, “Biomedical Optics: Principles and Imaging,” Hoboken, NJ, USA: John Wiley & Sons; 2007.
S. Yan, Y. Xiong, D. Lin, and editors, “Spatial temporal graph convolutional networks for skeleton-based action recognition,” Proceedings of the AAAI Conference on Artificial Intelligence; 2–7 February 2018; New Orleans, LA, USA. Palo Alto, CA, USA: AAAI Press; 2018, doi: 10.1609/aaai.v32i1.12328.
S. Afzal, H. A. Khan, M. J. Piran, and J. W. Lee, “A comprehensive survey on affective computing: Challenges, trends, applications, and future directions,” IEEE Access, vol. 12, pp. 96150-96168, 2024, doi: 10.1109/ACCESS.2024.3422480.
Y. Deng, C. C. Loy, and X. Tang, “Image aesthetic assessment: An experimental survey,” IEEE Signal Processing Magazine, vol. 34, no. 4, pp. 80-106, 2017, doi: 10.1109/MSP.2017.2696576.
R. Datta, D. Joshi, J. Li, J. Z. Wang, and editors, “Studying aesthetics in photographic images using a computational approach,” European Conference on Computer Vision. Berlin, Heidelberg: Springer; 2006, doi: 10.1007/11744078_23.
X. Lu, Z. Lin, H. Jin, J. Yang, J. Z. Wang, and editors, “Rapid: Rating pictorial aesthetics using deep learning,” Proceedings of the 22nd ACM International Conference on Multimedia; 3-7 November 2014; Orlando, FL, USA. New York, NY, USA: ACM; 2014, doi: 10.1145/2647868.2654927.
L. Mai, H. Jin, F. Liu, and editors, “Composition-preserving deep photo aesthetics assessment,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; 27-30 June 2016; Las Vegas, NV, USA. New York, NY, USA: IEEE; 2016, [Online]. Available: https://openaccess.thecvf.com/content_cvpr_2016/papers/Mai_Composition-Preserving_Deep_Photo_CVPR_2016_paper.pdf.
S. Fan, B. L. Koenig, Q. Zhao, and M. S. Kankanhalli, “A deeper look at human visual perception of images,” SN Computer Science, vol. 1, no. 1, pp. 58, 2020, doi: 10.1007/s42979-019-0061-5.
M. Hadi Kiapour, X. Han, S. Lazebnik, A. C. Berg, T. L. Berg, and editors, “Where to buy it: Matching street clothing photos in online shops,” Proceedings of the IEEE International Conference on Computer Vision; 7-13 December 2015; Santiago, Chile. New York, NY, USA: IEEE; 2015, doi: 10.1109/ICCV.2015.382.
X. Han, Z. Wu, Z. Wu, R. Yu, and L. Davis, “Viton: An image-based virtual try-on network,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; 18-23 June 2018; Salt Lake City, UT, USA. New York, NY, USA: IEEE; 2018, doi: 10.1109/CVPR.2018.00787.
Z. Liu, P. Luo, S. Qiu, X. Wang, X. Tang, and editors, “Deepfashion: Powering robust clothes recognition and retrieval with rich annotations,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; 27-30 June 2016; Las Vegas, NV, USA. New York, NY, USA: IEEE; 2016, doi: 10.1109/CVPR.2016.124.
R. M. Haralick, K. Shanmugam, and I. H. Dinstein, “Textural features for image classification,” IEEE Transactions on Systems, Man, and Cybernetics, vol. SMC-3, no. 6, pp. 610-621, 1973, doi: 10.1109/TSMC.1973.4309314.
K.-J. Choi and H.-S. Ko, “Stable but responsive cloth,” ACM SIGGRAPH 2005 Courses. New York, NY, USA: Association for Computing Machinery; 2005. p. 1-es, doi: 10.1145/1198555.1198571.
D. Baraff and A. Witkin, “Large steps in cloth simulation,” Seminal Graphics Papers: Pushing the Boundaries. New York, NY, USA: Association for Computing Machinery; 2023. Volume 2, p. 767-778, doi: 10.1145/3596711.3596792.
T. Baltrušaitis, C. Ahuja, and L.-P. Morency, “Multimodal machine learning: A survey and taxonomy,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 41, no. 2, pp. 423-443, 2018, doi: 10.1109/TPAMI.2018.2798607.
W. Guo, J. Wang, and S. Wang, “Deep multimodal representation learning: A survey,” IEEE Access, vol. 7, pp. 63373-63394, 2019, doi: 10.1109/ACCESS.2019.2916887.
K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhudinov, editors, et al., “Show, attend and tell: Neural image caption generation with visual attention,” Proceedings of the 32nd International Conference on Machine Learning, PMLR; 6-11 July 2015; Lille, France, 2015, doi: 10.48550/arXiv.1502.03044.
S. Chaudhari, V. Mithal, G. Polatkan, and R. Ramanath, “An attentive survey of attention models,” ACM Transactions on Intelligent Systems and Technology (TIST), vol. 12, no. 5, pp. 1-32, 2021, doi: 10.1145/3465055.