UGC-Driven Dynamic Evolution of Public Art Installations via Multimodal Transformer: A Tourist Content Perspective
Main Article Content
Abstract
Public art installations have long been limited by their inability to perceive and respond in real time to visitor-generated content in situ. To address this limitation, this study proposes a dynamically evolving public art installation system driven by tourist user-generated content and supported by a multimodal Transformer architecture. The system processes heterogeneous image, text, and audio inputs through a hierarchical cross-modal attention mechanism to extract collective affective semantics from visitor-generated signals. A differentiable parameter-mapping network then converts these semantic representations into real-time control actions for installation color, morphology, and rhythmic behavior. The framework also incorporates edge-side deployment, lowlatency communication, and continuous data-stream processing, which are essential for public environments involving dense wireless signal propagation and multimodal sensing. Three field tests were conducted in urban public plazas in China. The results show that tri-modal semantic understanding achieved a fusion accuracy of 91.2%, while end-to-end latency remained below 500 ms. Compared with static installations, the proposed system increased average visitor dwell time by approximately 100.2% and secondary user-generated content publication by 121.8%. During 30 days of continuous operation, system stability reached 99.2%. These results demonstrate the feasibility of integrating multimodal semantic understanding, real-time signal acquisition, and responsive visual control for human-machine co-creative public artworks, and provide an engineering pathway for interactive installations operating in complex wireless and acoustic environments.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
J. M. Nair, C. L. Poo, K. L. Ming, et al., “Gait-ViT: Gait recognition with vision transformer,” Sensors, vol. 22, no. 19, Art. no. 7362, 2022, doi: 10.3390/S22197362.
W. Ruhan, R. Philip, and C. Fan, “A hybrid quantum-classical neural network for learning transferable visual representation,” Quantum Sci. Technol., vol. 8, no. 4, 2023, doi: 10.1088/2058-9565/ACF1C7.
Y. Liu, B. Zhang, C. Wang, et al., “Vision-language representation learning with breadth and depth attention pre-training,” Knowl.-Based Syst., vol. 310, Art. no. 112941, 2025, doi: 10.1016/J.KNOSYS.2024.112941.
W. Huang, C. Li, H. Yang, et al., “Enhancing the vision-language foundation model with key semantic knowledge-emphasized report refinement,” Med. Image Anal., vol. 97, Art. no. 103299, 2024, doi: 10.1016/J.MEDIA.2024.103299.
D. Huiming, W. Sen, X. Zhifeng, et al., “A fine-grained vision and language representation framework with graph-based fashion semantic knowledge,” Comput. Graph., vol. 115, pp. 216–225, 2023, doi: 10.1016/J.CAG.2023.07.025.
L. Shiyi and W. Panpan, “Multi-dimensional fusion: transformer and GANs-based multimodal audiovisual perception robot for musical performance art,” Front. Neurorobot., vol. 17, Art. no. 1281944, 2023, doi: 10.3389/FNBOT.2023.1281944.
B. M. Ammar, A. Mendoza, N. Belkhir, et al., “Foundation models and transformers for anomaly detection: a survey,” Inf. Fusion, vol. 126, Art. no. 103517, 2026, doi: 10.1016/J.INFFUS.2025.103517.
C. Chen, Y. Wu, Q. Dai, et al., “A survey on graph neural networks and graph transformers in computer vision: a task-oriented perspective,” IEEE Trans. Pattern Anal. Mach. Intell., 2024, doi: 10.1109/TPAMI.2024.3445463.
A. Reza, K. Amirhossein, H. Moein, et al., “Advances in medical image analysis with vision transformers: a comprehensive review,” Med. Image Anal., vol. 91, Art. no. 103000, 2024, doi: 10.1016/J.MEDIA.2023.103000.
Y. Dazhi and S. Yunxue, “A data efficient transformer based on Swin Transformer,” Vis. Comput., vol. 40, no. 4, pp. 2589–2598, 2023, doi: 10.1007/S00371-023-02939-2.
G. J. Izabela, “Reading the unnamable, naming the visible, telling stories in a new way: picturebooks in teaching Polish as a foreign language,” Lang.: Codification, Competence, Commun., vol. 2, no. 11, pp. 23–40, 2024, doi: 10.2478/LCCC-2024-0008.
B. J. L. Nixon, “Do deep learning models accurately measure visual destination image? A comparison of a fine-tuned model to past work,” Inf. Technol. Tourism, vol. 26, no. 3, pp. 377–406, 2024, doi: 10.1007/S40558-024-00293-0.
T. Hu and J. Geng, “Research on the perception of the terrain image of the tourism destination based on multimodal user-generated content data,” PeerJ Comput. Sci., vol. 10, Art. no. e1801, 2024, doi: 10.7717/PEERJ-CS.1801.
F. Xiuqing, W. Fang, and W. Ying, “Inbound tourists’ perception of tourist destination image classified by UGC picture computer program,” J. Electr. Comput. Eng., vol. 2022, Art. no. 3100892, 2022, doi: 10.1155/2022/3100892.
M. K. Aboalganam, F. S. AlFraihat, and S. Tarabieh, “The impact of user-generated content on tourist visit intentions: the mediating role of destination imagery,” Administrative Sciences, vol. 15, no. 4, Art. no. 117, 2025, doi: 10.3390/ADMSCI15040117.
C. V. Fajardo, R. I. Rodríguez, and P. M. Cabrera, “From words to visuals: a transformer-based multimodal framework for emotion-driven tourism analytics,” Inf. Technol. Tourism, vol. 27, no. 4, pp. 1–41, 2025, doi: 10.1007/S40558-025-00334-2.
S. Zhang, Y. Li, X. Song, et al., “Multi-dimensional perceptual recognition of tourist destination using deep learning model and geographic information system,” PLoS ONE, vol. 20, no. 2, Art. no. e0318846, 2025, doi: 10.1371/JOURNAL.PONE.0318846.
A. Twil, O. Bencharef, and S. Kaloun, “Analyzing tourism reviews using an LDA topic-based sentiment analysis approach,” MethodsX, vol. 9, Art. no. 101894, 2022, doi: 10.1016/J.MEX.2022.101894.
S. S. A. Muazzam and O. YuYen, “TRP-BERT: discrimination of transient receptor potential (TRP) channels using contextual representations from deep bidirectional transformer based on BERT,” Comput. Biol. Med., vol. 137, Art. no. 104821, 2021, doi: 10.1016/J.COMPBIOMED.2021.104821.
B. Abayomi, N. SinChun, and L. ManFai, “A BERT framework to sentiment analysis of tweets,” Sensors, vol. 23, no. 1, Art. no. 506, 2023, doi: 10.3390/S23010506.
L. Wenfeng, Y. Jing, H. Zhanliang, et al., “An improved BERT and syntactic dependency representation model for sentiment analysis,” Comput. Intell. Neurosci., vol. 2022, Art. no. 5754151, 2022, doi: 10.1155/2022/5754151.
Y. Yang, J. Xu, L. Zhao, et al., “How users’ personality traits predict sentiment tendencies of user-generated content in social media: a mixed method of configuration analysis and machine learning,” J. Personality, vol. 93, no. 5, pp. 1175–1188, 2024, doi: 10.1111/JOPY.13000.
P. A. Kirilenko and S. Stepchenkova, “Facilitating topic modeling in tourism research: comprehensive comparison of new AI technologies,” Tourism Manage., vol. 106, Art. no. 105007, 2025, doi: 10.1016/J.TOURMAN.2024.105007.
S. Z. Rahmani, A. R. Hossein, and M. B. Sadegh, “Persian text sentiment analysis based on BERT and neural networks,” Iranian J. Sci. Technol., Trans. Electr. Eng., vol. 47, no. 4, pp. 1623–1634, 2023, doi: 10.1007/S40998-023-00626-5.
J. Li, C. Zhu, S. Zheng, et al., “ToPoFM: topology-guided pathology foundation model for high-resolution pathology image synthesis with cellular-level control,” IEEE Trans. Med. Imag., 2025, doi: 10.1109/TMI.2025.3548872.
N. I. Sari and W. Du, “Weighted similarity-confidence Laplacian synthesis for high-resolution art painting completion,” Appl. Sci., vol. 14, no. 6, Art. no. 2397, 2024, doi: 10.3390/APP14062397.
C. E. Salman, M. Mau, and S. Karnay, “Book review: Push the Button: Interactive Television and Collaborative Journalism in Japan, by Elizabeth Rodwell,” Television New Media, vol. 26, no. 8, pp. 935–937, 2025, doi: 10.1177/15274764251324933.
K. Yamada, “Push the Button: Interactive Television and Collaborative Journalism in Japan,” Japan Forum, vol. 37, no. 2, pp. 300–302, 2025, doi: 10.1080/09555803.2024.2408763.
P. Antonio, “Exploring the potential of generative AI (ChatGPT) for foreign language instruction: applications and challenges,” Hispania, vol. 106, no. 3, pp. 355–362, 2023.
X. Chen, Z. Ibrahim, and A. A. Aziz, “Predicting emotional responses in interactive art using random forests: a model grounded in enactive aesthetics,” Front. Psychol., vol. 16, Art. no. 1609103, 2025, doi: 10.3389/FPSYG.2025.1609103.
J. L. Hyun, J. M. O, and M. K. Jeong, “Characterizing smart environments as interactive and collective platforms: a review of the key behaviors of responsive architecture,” Sensors, vol. 21, no. 10, Art. no. 3417, 2021, doi: 10.3390/S21103417.
C. Marianna, “Interactive art as reflective experience: imagineers and ultra-technologists as interaction designers,” Vis. Resources, vol. 36, no. 4, pp. 382–396, 2020, doi: 10.1080/01973762.2022.2041218.
S. Pan and Q. Shi, “Exploring the evolution of museum knowledge organization systems,” Cataloging Classification Quart., vol. 63, no. 6–7, pp. 475–493, 2025, doi: 10.1080/01639374.2025.2544142.
M. L. Arenas, M. Gromaz, D. F. Fontela, et al., “WED-449 Exploring the evolution of circulating protein biomarkers in liver transplantation setting for MASH, ALD, and MetALD,” J. Hepatol., vol. 82, no. S1, p. S561, 2025, doi: 10.1016/S0168-8278(25)01522-3.
E. Lamboglia, G. Cambone, D. F. Frate, et al., “The evolution of Earth observation: exploring space sustainability through European case studies,” Int. J. Sustain. Eng., vol. 17, no. 1, pp. 632–641, 2024, doi: 10.1080/19397038.2024.2387427.