Improving Cross-Modal Matching of College English Listening Materials Using Contrastive Learning
Main Article Content
Abstract
Low audio-text matching accuracy in college English listening materials is caused by modality gaps, pronunciation variability, and scarce annotated data. To address this problem, this paper proposes a cross-modal alignment method based on contrastive learning. Wav2Vec 2.0 and BERT are used to extract speech and text features, respectively, followed by modality-specific projection networks that map heterogeneous representations into a shared semantic space. Under extreme data scarcity with a sample size of 5, semantically constructed positive and negative pairs and an in-batch hard negative mining strategy are introduced to improve discriminability. The InfoNCE loss optimizes cross-modal similarity, while a symmetric dual-tower encoder balances semantic consistency and feature independence. Experimental results show Top-1 accuracy of 94.7% ± 1.3% and Recall@5 of 97.6% ± 0.9%, outperforming AudioCLIP, CLAP, and Wav2CLIP baselines. A similarity gap of 0.58 at 10 dB confirms a clearly separated cross-modal decision boundary. The proposed approach mitigates modality disparity and small-sample limitations, providing a robust semantic alignment framework for intelligent listening instruction and low-resource multimodal signal matching.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
M. Gao, W. Song, C. Zhang, Q. Zhou, and P. Su, “Situational teaching based evaluation of college students’ English reading, listening, and speaking ability,” International Journal of Emerging Technologies in Learning (iJET), vol. 17, no. 8, pp. 140-154, 2022, doi: 10.3991/ijet.v17i08.30561.
H. Liu, R. Chen, S. Cao, and H. Lv, “Evaluation of college English teaching quality based on grey clustering analysis,” International Journal of Emerging Technologies in Learning (iJET), vol. 16, no. 2, pp. 173-187, 2021, doi: 10.3991/ijet.v16i02.19727.
S. C. Tsai, “Learning with mobile augmented reality-and automatic speech recognition-based materials for English listening and speaking skills: Effectiveness and perceptions of non-English major English as a foreign language students,” Journal of Educational Computing Research, vol. 61, no. 2, pp. 444-465, 2023, doi: 10.1177/07356331221111203.
N. Li, “A fuzzy evaluation model of college English teaching quality based on analytic hierarchy process,” International Journal of Emerging Technologies in Learning (iJET), vol. 16, no. 2, pp. 17-30, 2021, doi: 10.3991/ijet.v16i02.19731.
T. Hao, H. Sheng, Y. Ardasheva, and Z. Wang, “Effects of dual subtitles on Chinese students’ English listening comprehension and vocabulary learning,” The Asia-Pacific Education Researcher, vol. 31, no. 5, pp. 529-540, 2022, doi: 10.1007/s40299-021-00601-w.
H. H. H. Ali, “The importance of the four English language skills: Reading, writing, speaking, and listening in teaching Iraqi learners,” Humanities & Natural Sciences Journal, vol. 3, no. 2, pp. 154-165, 2022, doi: 10.53796/hnsj3210.
R. Al-Jarf, “Mobile Audiobooks, Listening Comprehension and EFL College Students,” Online Submission, vol. 9, no. 4, pp. 410-423, 2021, doi: 10.29121/granthaalayah.v9.i4.2021.3868.
J. Yuan, Y. Liu, X. Han, A. Li, and L. Zhao, “Educational metaverse: an exploration and practice of VR wisdom teaching model in Chinese Open University English course,” Interactive Technology and Smart Education, vol. 20, no. 3, pp. 403-421, 2023, doi: 10.1108/ITSE-10-2022-0140.
Y. Wang, “Artificial intelligence technologies in college English translation teaching,” Journal of psycholinguistic research, vol. 52, no. 5, pp. 1525-1544, 2023, doi: 10.1007/s10936-023-09960-5.
N. Noviana and L. Oktaviani, “The correlation between college student personality types and English proficiency ability at Universitas Teknokrat Indonesia,” Journal of English Language teaching and learning, vol. 3, no. 1, pp. 54-60, 2022, doi: 10.33365/jeltl.v3i1.1709.
Y. Zhang, “The research on critical thinking teaching strategies in college English classroom,” Creative Education, vol. 13, no. 4, pp. 1469-1485, 2022, doi: 10.4236/ce.2022.134090.
C. Y. Jao, H. C. Yeh, W. R. Huang, and N. S. Chen, “Using video dubbing to foster college students’ English-speaking ability,” Computer Assisted Language Learning, vol. 37, no. 4, pp. 585-607, 2024, doi: 10.1080/09588221.2022.2049824.
R. L. Fitriani, “The development of English speaking proficiency to increase students’ communication skill in a business and technology college,” KOMVERSAL, vol. 4, no. 2, pp. 90-112, 2022, doi: 10.38204/komversal.v4i2.1041.
P. Gou, “Teaching English using mobile applications to improve academic performance and language proficiency of college students,” Education and Information Technologies, vol. 28, no. 12, pp. 16935-16949, 2023, doi: 10.1007/s10639-023-11864-9.
R. Gao, “The vocabulary teaching mode based on the theory of constructivism,” Theory and Practice in Language Studies, vol. 11, no. 4, pp. 442-446, 2021, doi: 10.17507/tpls.1104.14.
C. T. Y. Yang, S. L. Lai, and H. H. J. Chen, “The impact of intelligent personal assistants on learners’ autonomous learning of second language listening and speaking,” Interactive Learning Environments, vol. 32, no. 5, pp. 2175-2195, 2024, doi: 10.1080/10494820.2022.2141266.
R. Al-Jarf, “Learning vocabulary in the app store by EFL college students,” Online Submission, vol. 5, no. 1, pp. 216-225, 2022, doi: 10.47191/ijsshr/v5-i1-30.
L. Diao and P. Hu, “Deep learning and multimodal target recognition of complex and ambiguous words in automated English learning system,” Journal of Intelligent & Fuzzy Systems, vol. 40, no. 4, pp. 7147-7158, 2021, doi: 10.3233/JIFS-189543.
D. Thi Ngu, D. T. Huong, D. T. N. Huy, P. T. Than, and E. S. Dongul, “Language teaching application to English students at master’s grade levels on history and macroeconomic-banking management courses in universities and colleges,” Journal of Language and Linguistic Studies, vol. 17, no. 3, pp. 1457-1468, 2021, doi: 10.52462/jlls.105.
H. C. Yeh, W. Y. Chang, H. Y. Chen, and L. Heng, “Effects of podcast-making on college students’ English speaking skills in higher education,” Educational technology research and development, vol. 69, no. 5, pp. 2845-2867, 2021, doi: 10.1007/s11423-021-10026-3.
S. Khurana, A. Laurent, and J. Glass, “Samu-xlsr: Semantically-aligned multimodal utterance-level cross-lingual speech representation,” IEEE Journal of Selected Topics in Signal Processing, vol. 16, no. 6, pp. 1493-1504, 2022, doi: 10.1109/JSTSP.2022.3192714.
S. Ge, J. Ren, Y. Shi, Y. Zhang, S. Yang, and J. Yang, “Audio-Text Multimodal Speech Recognition via Dual-Tower Architecture for Mandarin Air Traffic Control Communications,” CMC-COMPUTERS MATERIALS & CONTINUA, vol. 78, no. 3, pp. 3215-3245, 2024, doi: 10.32604/cmc.2023.046746.
A. Chakhtouna, S. Sekkate, and A. Adib, “Efficient bimodal emotion recognition system based on speech/text embeddings and ensemble learning fusion,” Annals of Telecommunications, vol. 80, no. 5, pp. 379-399, 2025, doi: 10.1007/s12243-025-01088-y.
S. Sekkate, M. Khalil, and A. Adib, “A statistical feature extraction for deep speech emotion recognition in a bilingual scenario,” Multimedia Tools and Applications, vol. 82, no. 8, pp. 11443-11460, 2023, doi: 10.1007/s11042-022-14051-z.
V. Poluboina, A. Pulikala, and A. N. Pitchaimuthu, “Deep Speech Denoising with Minimal Dependence on Clean Speech Data,” Circuits, Systems, and Signal Processing, vol. 43, no. 6, pp. 3909-3926, 2024, doi: 10.1007/s00034-024-02644-y.
S. J. Johnson, M. R. Murty, and I. Navakanth, “A detailed review on word embedding techniques with emphasis on word2vec,” Multimedia Tools and Applications, vol. 83, no. 13, pp. 37979-38007, 2024, doi: 10.1007/s11042-023-17007-z.
G. Di Gennaro, A. Buonanno, and F. A. N. Palmieri, “Considerations about learning Word2Vec,” The Journal of Supercomputing, vol. 77, no. 11, pp. 12320-12335, 2021, doi: 10.1007/s11227-021-03743-2.
S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, Z. Chen, et al., “Wavlm: Large-scale self-supervised pre-training for full stack speech processing,” IEEE Journal of Selected Topics in Signal Processing, vol. 16, no. 6, pp. 1505-1518, 2022, doi: 10.1109/JSTSP.2022.3188113.
R. Mao, Q. Liu, K. He, W. Li, and E. Cambria, “The biases of pre-trained language models: An empirical study on prompt-based sentiment analysis and emotion detection,” IEEE transactions on affective computing, vol. 14, no. 3, pp. 1743-1753, 2022, doi: 10.1109/TAFFC.2022.3204972.
L. Xu, H. Xie, Z. Li, F. L. Wang, W. Wang, and Q. Li, “Contrastive learning models for sentence representations,” ACM Transactions on Intelligent Systems and Technology, vol. 14, no. 4, pp. 1-34, 2023, doi: 10.1145/3593590.
J. Yu, X. Xia, T. Chen, L. Cui, N. Q. V. Hung, and H. Yin, “XSimGCL: Towards extremely simple graph contrastive learning for recommendation,” IEEE Transactions on Knowledge and Data Engineering, vol. 36, no. 2, pp. 913-926, 2023, doi: 10.1109/TKDE.2023.3288135.
T. Ao, Z. Zhang, and L. Liu, “Gesturediffuclip: Gesture diffusion model with clip latents,” ACM Transactions on Graphics (TOG), vol. 42, no. 4, pp. 1-18, 2023, doi: 10.1145/3592097.
X. Tan, J. Chen, H. Liu, J. Cong, C. Zhang, Y. Liu, et al., “Naturalspeech: End-to-end text-to-speech synthesis with human-level quality,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 6, pp. 4234-4245, 2024, doi: 10.1109/TPAMI.2024.3356232.
Y. Yasuda and T. Toda, “Investigation of Japanese PnG BERT language model in text-to-speech synthesis for pitch accent language,” IEEE Journal of Selected Topics in Signal Processing, vol. 16, no. 6, pp. 1319-1328, 2022, doi: 10.1109/JSTSP.2022.3190672.
Y. Fan, Z. Lin, J. Saito, W. Wang, and T. Komura, “Joint audio-text model for expressive speech-driven 3d facial animation,” Proceedings of the ACM on Computer Graphics and Interactive Techniques, vol. 5, no. 1, pp. 1-15, 2022, doi: 10.1145/3522615.