All Issue

2026 Vol.45, Issue 5 Preview Page
30 September 2026. pp. 582-590
Abstract
References
1

X. Shi, X. Wang, Z. Guo, Y. Wang, P. Zhang, X. Zhang, Z. Guo, H. Hao, Y. Xi, B. Yang, J. Xu, J. Zhou, and J. Lin, “Qwen3-ASR technical report,” Qwen Team, Tech. Rep., 2026.

2

A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” Proc. ICML, 28492-28518 (2023).

3

J.-U. Bang, S. Yun, S.-H. Kim, M.-Y. Choi, M.-K. Lee, Y.-J. Kim, D.-H. Kim, J. Park, Y.-J. Lee, and S.-H. Kim, “KsponSpeech: Korean spontaneous speech corpus for automatic speech recognition,” Appl. Sci. 10, 6936 (2020).

10.3390/app10196936
4

Zeroth-Korean: Korean Open-source Speech Corpus for Speech Recognition by Zeroth Project, https:// www.openslr.org/40/, (Last viewed September 7, 2026).

5

J.-W. Ha, K. Nam, J. Kang, S.-W. Lee, S. Yang, H. Jung, E. Kim, H. Kim, S. Kim, H. A. Kim, K. Doh, C. K. Lee, N. Sung, and S. Kim, “ClovaCall: Korean goal-oriented dialog speech corpus for automatic speech recognition of contact centers,” Proc. Interspeech , 409-413 (2020).

10.21437/Interspeech.2020-1136
6

AI-Hub, https://www.aihub.or.kr/, (Last viewed September 9, 2026).

7

C. Tang, W. Yu, G. Sun, X. Chen, T. Tan, W. Li, L. Lu, Z. Ma, and C. Zhang, “SALMONN: Towards generic hearing abilities for large language models,” Proc. ICLR, 16607-16629 (2024).

8

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-scale Audio-language Models, https://arxiv.org/abs/2311.07919, (Last viewed September 9, 2026).

9

Y. Chu, J. Xu, Q. Yang, H. Wei, X. Wei, Z. Guo, Y. Leng, Y. Lv, J. He, J. Lin, C. Zhou, and J. Zhou, “Qwen2-Audio technical report,” Qwen Team, Tech. Rep., 2024.

10

Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition, https://arxiv. org/abs/2407.04675, (Last viewed September 9, 2026).

11

J. Wu, Y. Gaur, Z. Chen, L. Zhou, Y. Zhu, T. Wang, J. Li, S. Liu, B. Ren, L. Liu, and Y. Wu, “On decoder- only architecture for speech-to-text and large language model integration,” Proc. ASRU, 1-8 (2023).

10.1109/ASRU57964.2023.10389705
12

X. Geng, T. Xu, K. Wei, B. Mu, H. Xue, H. Wang, Y. Li, P. Guo, Y. Dai, L. Li, M. Shao, and L. Xie, “Unveiling the potential of LLM-based ASR on Chinese open-source datasets,” Proc. ISCSLP, 26-30 (2024).

10.1109/ISCSLP63861.2024.10800077
13

A. Holtzman, J. Buys, L. Du, M. Forbes, and Y. Choi, “The curious case of neural text degeneration,” Proc. ICLR, 1-16 (2020).

14

CTRL: A Conditional Transformer Language Model for Controllable Generation, https://arxiv.org/abs/ 1909.05858, (Last viewed September 9, 2026).

15

and Qwen3-ASR 1.7B-hf Model Configurations, https://huggingface.co/collections/Qwen/ qwen3-asr, (Last viewed September 9, 2026).

16

K. Manohar and L. G. Pillai, “What is lost in normalization? Exploring pitfalls in multilingual ASR model evaluations,” Proc. EMNLP, 10864-10869 (2024).

10.18653/v1/2024.emnlp-main.607
17

S. Karita, R. Sproat, and H. Ishikawa, “Lenient evaluation of Japanese speech recognition: Modeling naturally occurring spelling inconsistency,” Proc. CAWL, 61-70 (2023).

10.18653/v1/2023.cawl-1.8
Information
  • Publisher :The Acoustical Society of Korea
  • Publisher(Ko) :한국음향학회
  • Journal Title :The Journal of the Acoustical Society of Korea
  • Journal Title(Ko) :한국음향학회지
  • Volume : 45
  • No :5
  • Pages :582-590
  • Received Date : 2026-07-15
  • Revised Date : 2026-08-11
  • Accepted Date : 2026-08-25