magazinelogo

Journal of Humanities, Arts and Social Science

ISSN Online: 2576-0548 ISSN Print: 2576-0556 CODEN: JHASAY
Frequency: monthly Email: jhass@hillpublisher.com
Total View: 6586154 Downloads: 1934746 Citations: 437 (From Dimensions)
ArticleOpen Access http://dx.doi.org/10.26855/jhass.2026.06.010

From Acoustic Diagnosis to Generative Interaction: An Integrated Model of Digital Tools in College English Speaking Instruction

Yijing Zhang

Department of Foreign Language Teaching and Research, Hebei Normal University, Shijiazhuang 050024, Hebei, China.

*Corresponding author: Yijing Zhang

This work was supported by the Teaching Reform Research Project of Hebei Normal University: Research on the Reform of English Speaking Instruction Based on Praat Visualized Speech Software (No. 2023XJJG082).
Published: June 30, 2026

Abstract

This article examines how digital speaking tools can be integrated into college English speaking instruction in Chinese higher education. It reviews three tool types: acoustic-analysis tools such as Praat, automatic speech recognition (ASR)-based pronunciation and speaking platforms, and large language model (LLM)-powered conversational agents. Drawing on research in technology-assisted language learning, pronunciation instruction, mobile-assisted practice, and artificial intelligence (AI)-supported oral interaction, the article compares their functions and limits in classroom teaching and learner autonomy. The analysis shows that these tools serve different stages of speaking development: acoustic-analysis tools support pronunciation diagnosis, ASR-based platforms and video-dubbing applications support imitation and feedback, and LLM-powered conversational agents support situated oral production. Based on a focus on form, situated cognition, and task-based language teaching, the article proposes a “diagnosis-imitation-situated production” model. The model empha-sizes teacher-designed tasks, process-based evidence, and a balance among pronunciation accuracy, oral fluency, and pragmatic appropriateness.

Keyword

College English speaking; digital speaking tools; situated learning

References

Amrate, M., & Tsai, P. (2025). Computer-assisted pronunciation training: A systematic review. ReCALL, 37(1), 22-42.

https://doi.org/10.1017/S0958344024000181

Brown, J. S., Collins, A., & Duguid, P. (1989). Situated cognition and the culture of learning. Educational Researcher, 18(1), 32-42. https://doi.org/10.3102/0013189X018001032

Chapelle, C. A. (2003). English language learning and technology: Lectures on applied linguistics in the age of information and communication technology. John Benjamins.
https://doi.org/10.1075/lllt.7

Dai, Y., & Wu, Z. (2023). Mobile-assisted pronunciation learning with feedback from peers and/or automatic speech recognition: A mixed-methods study. Computer Assisted Language Learning, 36(5-6), 861-884. 

https://doi.org/10.1080/09588221.2021.1952272

Derwing, T. M., & Munro, M. J. (2015). Pronunciation fundamentals: Evidence-based perspectives for L2 teaching and research. John Benjamins.

Du, J., & Daniel, B. K. (2024). Transforming language education: A systematic review of AI-powered chatbots for English as a foreign language speaking practice. Computers and Education: Artificial Intelligence, 6, Article 100230.

https://doi.org/10.1016/j.caeai.2024.100230

Ellis, R. (2003). Task-based language learning and teaching. Oxford University Press.

Godwin-Jones, R. (2023). Emerging spaces for language learning: AI bots, ambient intelligence, and the metaverse. Language Learning & Technology, 27(2), 6-27.
https://hdl.handle.net/10125/73501

Golonka, E. M., Bowles, A. R., Frank, V. M., Richardson, D. L., & Freynik, S. (2014). Technologies for foreign language learning: A review of technology types and their effectiveness. Computer Assisted Language Learning, 27(1), 70-105.

https://doi.org/10.1080/09588221.2012.700315

Hsu, L. (2016). An empirical examination of EFL learners’ perceptual learning styles and acceptance of ASR-based computer-assisted pronunciation training. Computer Assisted Language Learning, 29(5), 881-900. 

https://doi.org/10.1080/09588221.2015.1069747

Kessler, G. (2018). Technology and the future of language teaching. Foreign Language Annals, 51(1), 205-218.

https://doi.org/10.1111/flan.12318

Kukulska-Hulme, A., & Shield, L. (2008). An overview of mobile assisted language learning: From content delivery to supported collaboration and interaction. ReCALL, 20(3), 271-289.
https://doi.org/10.1017/S0958344008000335

Levis, J. M. (2007). Computer technology in teaching and researching pronunciation. Annual Review of Applied Linguistics, 27, 184-202. https://doi.org/10.1017/S0267190508070098

Levis, J. M., & Grant, L. (2003). Integrating pronunciation into ESL/EFL classrooms. TESOL Journal, 12(2), 13-19.

Long, M. H. (1991). Focus on form: A design feature in language teaching methodology. In K. de Bot, R. B. Ginsberg, & C. Kramsch (Eds.), Foreign language research in cross-cultural perspective (pp. 39-52). John Benjamins.

https://doi.org/10.1075/sibil.2.07lon

López-Molines, L. (2025). A systematic review of generative artificial intelligence-based tools to improve oral skills in English as a foreign language. The EuroCALL Review, 32(2), 183-195. 

https://doi.org/10.4995/eurocall.2025.23941

Lyster, R., Saito, K., & Sato, M. (2013). Oral corrective feedback in second language classrooms. Language Teaching, 46(1), 1-40. https://doi.org/10.1017/S0261444812000365

Nickolai, D., Schaefer, E., & Figueroa, P. (2024). Aggregating the evidence of automatic speech recognition research claims in CALL. System, 121, Article 103250.
https://doi.org/10.1016/j.system.2024.103250

Reinders, H., & White, C. (2016). 20 years of autonomy and technology: How far have we come and where to next? Language Learning & Technology, 20(2), 143-154.
https://hdl.handle.net/10125/44466

Wang, X., & Lee, S.-M. (2025). The impact of video dubbing app on Chinese college students' oral language skills across different proficiency levels. International Journal of Educational Research, 130, Article 102521.

https://doi.org/10.1016/j.ijer.2024.102521

Copyright

© 2026 by the author(s).
This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution-NonCommercial-NoDerivatives (CC BY-NC-ND) license, which permits non-commercial use, distribution, and reproduction in any medium, provided the original work is properly cited and is not modified or adapted.
https://creativecommons.org/licenses/by-nc-nd/4.0/

How to cite this paper

From Acoustic Diagnosis to Generative Interaction: An Integrated Model of Digital Tools in College English Speaking Instruction

How to cite this paper: Yijing Zhang. (2026) From Acoustic Diagnosis to Generative Interaction: An Integrated Model of Digital Tools in College English Speaking Instruction. Journal of Humanities, Arts and Social Science10(6), 672-680.

DOI: http://dx.doi.org/10.26855/jhass.2026.06.010