Energetika, elektronika i telekomunikacije
Vol. 41 No. 08 (2026): Proceedings of the Faculty of Technical Sciences
Performance analysis of the Whisper model in translating Serbian speech into English
Abstract
This paper analyzes methods for automatic speech translation from Serbian to English, focusing on Whisper models and large language models (LLMs). It compares a direct end-to-end approach with a cascade system that separates speech recognition and translation. Evaluation was conducted on multiple test sets using BLEU, METEOR, and WER metrics. Results show that the cascade approach, combining Whisper transcription and LLM translation, outperforms direct Whisper translation. Fine-tuning Whisper for automatic speech recognition in Serbian significantly improves transcription and translation quality. LLMs demonstrate
robustness even with transcription errors, highlighting their potential in this domain. Future work includes further fine-tuning and integration of language models to enhance translation accuracy.
References
- [1] A. B´erard, O. Pietquin, C. Servan, and L. Besacier,
- “Listen and translate: A proof of concept for end-to
- end speech-to-text translation,” in NIPS Workshop on
- End-to-End Learning for Speech and Audio
- Processing, 2016.
- [2] R. J. Weiss, J. Chorowski, N. Jaitly, Y. Wu, and Z.
- Chen, “Sequence-to-sequence models can directly
- translate foreign speech,” in Interspeech. ISCA, 2017,
- pp. 2625–2629.
- [3] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L.
- Jones, A. N. Gomez, Kaiser, and I. Polosukhin,
- “Attention is all you need,” in Advances in Neural
- Information Processing Systems (NeurIPS), 2017, pp.
- 5998–6008.
- [4] A. Radford, J. W. Kim, T. Xu, G. Brockman, C.
- McLeavey, and I. Sutskever, “Robust speech
- recognition via large-scale weak supervision,” arXiv
- preprint arXiv:2212.04356, 2022.
- [5] Deng, L., & Yu, D. (2014). Deep learning: methods
- and applications. Foundations and trends® in signal
- processing, 7(3–4), 197-387.
- [6] I. Goodfellow, Y. Bengio, and A. Courville, Deep
- Learning. Cambridge, MA: MIT Press, 2016, ch. 9:
- “Convolutional Networks”.
- [7] Hendrycks, D. and Gimpel, K. Gaussian error linear
- units (gelus). arXiv preprint arXiv:1606.08415, 2016.
- [8] Press, O. and Wolf, L. Using the output embedding to
- improve language models. In Proceedings of the 15th
- Conference of the European Chapter of the Associa
- tion for Computational Linguistics: Volume 2, Short
- Papers, pp. 157–163, Valencia, Spain, April 2017. As
- sociation for Computational Linguistics. URL https:
- //aclanthology.org/E17-2025.
- [9] Child, R., Gray, S., Radford, A., and Sutskever, I. Gen
- erating long sequences with sparse transformers.
- arXiv preprint arXiv:1904.10509, 2019.
- [10] Sennrich, R., Haddow, B., and Birch, A. Neural
- machine translation of rare words with subword units.
- arXiv preprint arXiv:1508.07909, 2015.
- [11] Radford, A., Wu, J., Child, R., Luan, D., Amodei, D.,
- and Sutskever, I. Language models are unsupervised
- multitask learners. 2019.
- [12] K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu,
- “Bleu: a method for automatic evaluation of machine
- translation,” in Proceedings of the 40th Annual
- Meeting of the Association for Computational
- Linguistics, 2002, pp. 311–318.
- [13] S. Banerjee and A. Lavie, “Meteor: An automatic
- metric for mt evalua-tion with improved correlation
- with human judgments,” in Proceedings of the ACL
- Workshop on Intrinsic and Extrinsic Evaluation
- Measures for Machine Translation and/or
- Summarization, 2005, pp. 65–72.
- [14] A. Ali and S. Renalds, “Word Error Rate Estimation
- for Speech Recognition: e-WER,” in Proceedings of
- the 56th Annual Meeting of the Association for
- Computational Linguistics (Short Papers), pp. 20–24
- Melbourne, Australia, July 15 - 20, 2018.
- [15] V. Timmel, C. Paonessa, M. Vogel, D. Perruchoud
- and R. Kakooe, “Fine-tuning Whisper on
- Low-Resource Languages for Real-World
- Applications,” arXiv preprint arXiv:2412.15726,
- Apr. 2025.
- [16] R. Ma, M. Qian, Y. Fathullah, S. Tang, M. Gales and
- K. Knill, “Cross-Lingual Transfer Learning for Speech
- Translation,” arXiv preprint, arXiv:2407.01130,
- Feb. 2025.
- [17] P. Peng, B. Yan, S. Watanabe and D. Harwath,
- “Prompting the hidden talent of web-scale speech
- models for zero-shot task generalization,” arXiv
- preprint, arXiv:2305.11095v3, Aug. 2023.