Multimodal Fusion Comprehension Model for English Listening and Literature Appreciation Based on the Perceiver IO Architecture
Main Article Content
Abstract
Traditional English multimodal listening comprehension models suffer from semantic fragmentation and weak task transferability because they lack a unified cross-modal attention mechanism, have difficulty processing high-dimensional images and long-duration speech signals simultaneously, and often rely on loosely coupled fusion strategies such as post -concatenation or simple weighting. To address these issues, this paper uses Wav2Vec 2.0 and BERT to achieve semantic alignment between speech and text, ensuring spatiotemporal consistency of linguistic information. A unified input representation structure is then constructed to improve representational consistency across modalities. A fusion encoding module centered on Perceiver IO is built, and a cross-attention mechanism is introduced to realize contextual coupling among modalities. Finally, a multi-objective prediction structure is implemented at the output layer to identify rhetorical elements, sentiment nodes, and key sentences in literary listening materials, including novels, poetry recitations, and technical literary expressions. The model achieves sentiment classification accuracy of 0.86 with F1 of 0.85, and F1 above 0.80 in cross-text rhetorical recognition. The speech-text alignment distance decreases to nearly 0.2, with BLEU ranging from 0.78 to 0.85, confirming effective multimodal fusion and comprehension.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
P. Zhang and Y. Peng, “Multimodal reading in reading-only versus reading-while-listening modes: Evidence from Chinese language learners,” Chinese as a Second Language Research, vol. 13, no. 2, pp. 215-236, 2024, doi: 10.1515/caslar-2024-2003.
H. Cartner and D. Cameron, “Investigating metacognitive strategy awareness for multimodal listening,” E-Learning and Digital Media, vol. 20, no. 5, pp. 424-441, 2023, doi: 10.1177/20427530221108014.
M. Shamsitdinova, “Developing Multimodal Listening Skills in English for Specific (or special) Purposes: A pedagogical framework,” The American Journal of Social Science and Education Innovations, vol. 6, no. 06, pp. 15-21, 2024, doi: 10.37547/tajssei/Volume06Issue06-03.
S. Kaowiwattanakul, “Multimodal literature and CEFR reading proficiency: Improving literary reading skills in EFL learners,” Indonesian Journal of Applied Linguistics, vol. 15, no. 1, pp. 20-34, 2025, doi: 10.17509/ijal.v15i1.75921.
R. Masinde, D. Barasa, and L. Mandillah, “Effectiveness of using multimodal approaches in teaching and learning listening and speaking skills,” Nairobi Journal of Humanities and Social Sciences, vol. 7, no. 2, pp. 91-112, 2023, doi: 10.58256/njhs.v7i2.1333.
B. E. Smith, N. Amgott, and I. Malova, ““It made me think in a different way”: Bilingual students’ perspectives on multimodal composing in the English language arts classroom,” Tesol Quarterly, vol. 56, no. 2, pp. 525-551, 2022, doi: 10.1002/tesq.3064.
N. Trisanti, D. Suherdi, and D. Sukyadi, “Multimodality reflected in EFL teaching materials: Indonesian EFL in-service teacher’s multimodality literacy perception,” Language Circle: Journal of Language and Literature, vol. 17, no. 1, pp. 129-139, 2022, doi: 10.15294/lc.v17i1.38776.
N. A. Jamil and A. A. Aziz, “The use of multimodal text in enhancing students’ reading habit,” Malaysian Journal of Social Sciences and Humanities (MJSSH), vol. 6, no. 9, pp. 487-492, 2021, doi: 10.47405/mjssh.v6i9.977.
E. Tour and M. Barnes, “Engaging English language learners in digital multimodal composing: Pre-service teachers’ perspectives and experiences,” Language and Education, vol. 36, no. 3, pp. 243-258, 2022, doi: 10.1080/09500782.2021.1912083.
K. V. Alvionita, L. Widyaningrum, and A. Prayogo, “EFL learners’ reflection on digitally mediated multimodal project-based learning: Multimodal enactment in a listening-speaking class,” Language Circle: Journal of Language and Literature, vol. 17, no. 1, pp. 87-97, 2022, doi: 10.15294/lc.v17i1.36466.
F. V. Lim and L. Unsworth, “Multimodal composing in the English classroom: Recontextualising the curriculum to learning,” English in Education, vol. 57, no. 2, pp. 102-119, 2023, doi: 10.1080/04250494.2023.2187696.
P. Aedo and C. Millafilo, “Increasing vocabulary acquisition and retention in EFL young learners through the use of multimodal texts (memes),” Colombian Applied Linguistics Journal, vol. 24, no. 2, pp. 251-269, 2022, doi: 10.14483/22487085.18312.
A. von Zansen, R. Hilden, and E. Laihanen, “The multimodal listening test in a high-stakes context: Gender-neutral or not?,” International Journal of Listening, vol. 36, no. 2, pp. 152-170, 2022, doi: 10.1080/10904018.2021.1993446.
E. Salamanti, D. Park, N. Ali, and S. Brown, “The efficacy of collaborative and multimodal learning strategies in enhancing English language proficiency among ESL/EFL Learners: A quantitative analysis,” Research Studies in English Language Teaching and Learning, vol. 1, no. 2, pp. 78-89, 2023, doi: 10.62583/rseltl.v1i2.11.
S. Hartle, R. Facchinetti, and V. Franceschi, “Teaching communication strategies for the workplace: A multimodal framework,” Multimodal Communication, vol. 11, no. 1, pp. 5-15, 2022, doi: 10.1515/mc-2021-0005.
C. Tyrer, “The voice, text, and the visual as semiotic companions: An analysis of the materiality and meaning potential of multimodal screen feedback,” Education and Information Technologies, vol. 26, no. 4, pp. 4241-4260, 2021, doi: 10.1007/s10639-021-10455-w.
S. Salsabila, N. Nurami, R. T. Oktaviana, S. Muljanto, and A. Kurnia, “Seeing How Multimodal Source Are Used in EFL Classroom,” English Education and Applied Linguistics Journal (EEAL Journal), vol. 6, no. 2, pp. 102-106, 2023, [Online]. Available: https://api.semanticscholar.org/CorpusID:269469218.
L. Jiang, M. M. Gu, and F. Fang, “Multimodal or multilingual? Native English teachers’ engagement with translanguaging in Hong Kong TESOL classrooms,” Applied Linguistics Review, vol. 15, no. 4, pp. 1299-1319, 2024, doi: 10.1515/applirev-2022-0062.
A. Corbitt, M. Wargo J, and C. O’Connor, “Encountering unnatural E-literature: tracing interpretation and relationality across multimodal response and digital annotation,” English in Education, vol. 56, no. 2, pp. 186-200, 2022, doi: 10.1080/04250494.2021.1933424.
C. Smith, “Deconstructing innercirclism: A critical exploration of multimodal discourse in an English as a foreign language textbook,” Discourse: Studies in the Cultural Politics of Education, vol. 44, no. 1, pp. 88-105, 2023, doi: 10.1080/01596306.2021.1963212.
K. A. Sherwani and M. K. Harchegani, “The Impact of Multimodal Discourse Analysis on the Improvement of Iraqi EFL Learners’ Reading Comprehension Skill,” Journal of Tikrit University for Humanities, vol. 29, no. 12, 2, pp. 1-19, 2022, doi: 10.25130/jtuh.29.12.2.2022.22.
C. M. Ponzio and M. R. Deroo, “Harnessing multimodality in language teacher education: Expanding English-dominant teachers’ translanguaging capacities through a multimodalities entextualization cycle,” International Journal of Bilingual Education and Bilingualism, vol. 26, no. 8, pp. 975-991, 2023, doi: 10.1080/13670050.2021.1933893.
A. N. Kluger, M. Lehmann, H. Aguinis, and G. Itzchakov, “A meta-analytic systematic review and theory of the effects of perceived listening on work outcomes,” Journal of Business and Psychology, vol. 39, no. 2, pp. 295-344, 2024, doi: 10.1007/s10869-023-09897-5.
C. Jia, K. F. Hew, and M. Li, “Towards a flipped SEF-ARCS decoding model to improve foreign language listening proficiency,” Computer Assisted Language Learning, vol. 38, no. 3, pp. 369-396, 2025, doi: 10.1080/09588221.2023.2191655.
Y. Zhao and V. Aryadoust, “An automatized semantic analysis of two large-scale listening tests: A corpus-based study,” Language Testing, vol. 42, no. 3, pp. 312-343, 2025, doi: 10.1177/02655322241288598.
M. Mpumuje, G. Bazimaziki, and J. D. L. P. Muragijimana, “Exploring the Role of Oral Literature in Enhancing Learners’ Language Proficiency: A Case of Three Selected Secondary Schools in Rwanda,” African Journal of Empirical Research, vol. 5, no. 2, pp. 752-763, 2024, doi: 10.51867/ajernet.5.2.65.