X. Shi, X. Wang, Z. Guo, Y. Wang, P. Zhang, X. Zhang, Z. Guo, H. Hao, Y. Xi, B. Yang, J. Xu, J. Zhou, and J. Lin, “Qwen3-ASR technical report,” Qwen Team, Tech. Rep., 2026.
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” Proc. ICML, 28492-28518 (2023).
J.-U. Bang, S. Yun, S.-H. Kim, M.-Y. Choi, M.-K. Lee, Y.-J. Kim, D.-H. Kim, J. Park, Y.-J. Lee, and S.-H. Kim, “KsponSpeech: Korean spontaneous speech corpus for automatic speech recognition,” Appl. Sci. 10, 6936 (2020).
10.3390/app10196936Zeroth-Korean: Korean Open-source Speech Corpus for Speech Recognition by Zeroth Project, https:// www.openslr.org/40/, (Last viewed September 7, 2026).
J.-W. Ha, K. Nam, J. Kang, S.-W. Lee, S. Yang, H. Jung, E. Kim, H. Kim, S. Kim, H. A. Kim, K. Doh, C. K. Lee, N. Sung, and S. Kim, “ClovaCall: Korean goal-oriented dialog speech corpus for automatic speech recognition of contact centers,” Proc. Interspeech , 409-413 (2020).
10.21437/Interspeech.2020-1136C. Tang, W. Yu, G. Sun, X. Chen, T. Tan, W. Li, L. Lu, Z. Ma, and C. Zhang, “SALMONN: Towards generic hearing abilities for large language models,” Proc. ICLR, 16607-16629 (2024).
Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-scale Audio-language Models, https://arxiv.org/abs/2311.07919, (Last viewed September 9, 2026).
Y. Chu, J. Xu, Q. Yang, H. Wei, X. Wei, Z. Guo, Y. Leng, Y. Lv, J. He, J. Lin, C. Zhou, and J. Zhou, “Qwen2-Audio technical report,” Qwen Team, Tech. Rep., 2024.
Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition, https://arxiv. org/abs/2407.04675, (Last viewed September 9, 2026).
J. Wu, Y. Gaur, Z. Chen, L. Zhou, Y. Zhu, T. Wang, J. Li, S. Liu, B. Ren, L. Liu, and Y. Wu, “On decoder- only architecture for speech-to-text and large language model integration,” Proc. ASRU, 1-8 (2023).
10.1109/ASRU57964.2023.10389705X. Geng, T. Xu, K. Wei, B. Mu, H. Xue, H. Wang, Y. Li, P. Guo, Y. Dai, L. Li, M. Shao, and L. Xie, “Unveiling the potential of LLM-based ASR on Chinese open-source datasets,” Proc. ISCSLP, 26-30 (2024).
10.1109/ISCSLP63861.2024.10800077A. Holtzman, J. Buys, L. Du, M. Forbes, and Y. Choi, “The curious case of neural text degeneration,” Proc. ICLR, 1-16 (2020).
CTRL: A Conditional Transformer Language Model for Controllable Generation, https://arxiv.org/abs/ 1909.05858, (Last viewed September 9, 2026).
and Qwen3-ASR 1.7B-hf Model Configurations, https://huggingface.co/collections/Qwen/ qwen3-asr, (Last viewed September 9, 2026).
- Publisher :The Acoustical Society of Korea
- Publisher(Ko) :한국음향학회
- Journal Title :The Journal of the Acoustical Society of Korea
- Journal Title(Ko) :한국음향학회지
- Volume : 45
- No :5
- Pages :582-590
- Received Date : 2026-07-15
- Revised Date : 2026-08-11
- Accepted Date : 2026-08-25
- DOI :https://doi.org/10.7776/ASK.2026.45.5.582



The Journal of the Acoustical Society of Korea









