Skip to main navigation menu Skip to main content Skip to site footer

Energetika, elektronika i telekomunikacije

Vol. 41 No. 08 (2026): Proceedings of the Faculty of Technical Sciences

Performance analysis of the Whisper model in translating Serbian speech into English

  • Milana Vitković
DOI:
https://doi.org/10.24867/
Submitted
September 7, 2026
Published
2026-09-09

Abstract

This paper analyzes methods for automatic speech translation from Serbian to English, focusing on Whisper models and large language models (LLMs). It compares a direct end-to-end approach with a cascade system that separates speech recognition and translation. Evaluation was conducted on multiple test sets using BLEU, METEOR, and WER metrics. Results show that the cascade approach, combining Whisper transcription and LLM translation, outperforms direct Whisper translation. Fine-tuning Whisper for automatic speech recognition in Serbian significantly improves transcription and translation quality. LLMs demonstrate
robustness even with transcription errors, highlighting their potential in this domain. Future work includes further fine-tuning and integration of language models to enhance translation accuracy. 

References

  1. [1] A. B´erard, O. Pietquin, C. Servan, and L. Besacier,
  2. “Listen and translate: A proof of concept for end-to
  3. end speech-to-text translation,” in NIPS Workshop on
  4. End-to-End Learning for Speech and Audio
  5. Processing, 2016.
  6. [2] R. J. Weiss, J. Chorowski, N. Jaitly, Y. Wu, and Z.
  7. Chen, “Sequence-to-sequence models can directly
  8. translate foreign speech,” in Interspeech. ISCA, 2017,
  9. pp. 2625–2629.
  10. [3] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L.
  11. Jones, A. N. Gomez, Kaiser, and I. Polosukhin,
  12. “Attention is all you need,” in Advances in Neural
  13. Information Processing Systems (NeurIPS), 2017, pp.
  14. 5998–6008.
  15. [4] A. Radford, J. W. Kim, T. Xu, G. Brockman, C.
  16. McLeavey, and I. Sutskever, “Robust speech
  17. recognition via large-scale weak supervision,” arXiv
  18. preprint arXiv:2212.04356, 2022.
  19. [5] Deng, L., & Yu, D. (2014). Deep learning: methods
  20. and applications. Foundations and trends® in signal
  21. processing, 7(3–4), 197-387.
  22. [6] I. Goodfellow, Y. Bengio, and A. Courville, Deep
  23. Learning. Cambridge, MA: MIT Press, 2016, ch. 9:
  24. “Convolutional Networks”.
  25. [7] Hendrycks, D. and Gimpel, K. Gaussian error linear
  26. units (gelus). arXiv preprint arXiv:1606.08415, 2016.
  27. [8] Press, O. and Wolf, L. Using the output embedding to
  28. improve language models. In Proceedings of the 15th
  29. Conference of the European Chapter of the Associa
  30. tion for Computational Linguistics: Volume 2, Short
  31. Papers, pp. 157–163, Valencia, Spain, April 2017. As
  32. sociation for Computational Linguistics. URL https:
  33. //aclanthology.org/E17-2025.
  34. [9] Child, R., Gray, S., Radford, A., and Sutskever, I. Gen
  35. erating long sequences with sparse transformers.
  36. arXiv preprint arXiv:1904.10509, 2019.
  37. [10] Sennrich, R., Haddow, B., and Birch, A. Neural
  38. machine translation of rare words with subword units.
  39. arXiv preprint arXiv:1508.07909, 2015.
  40. [11] Radford, A., Wu, J., Child, R., Luan, D., Amodei, D.,
  41. and Sutskever, I. Language models are unsupervised
  42. multitask learners. 2019.
  43. [12] K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu,
  44. “Bleu: a method for automatic evaluation of machine
  45. translation,” in Proceedings of the 40th Annual
  46. Meeting of the Association for Computational
  47. Linguistics, 2002, pp. 311–318.
  48. [13] S. Banerjee and A. Lavie, “Meteor: An automatic
  49. metric for mt evalua-tion with improved correlation
  50. with human judgments,” in Proceedings of the ACL
  51. Workshop on Intrinsic and Extrinsic Evaluation
  52. Measures for Machine Translation and/or
  53. Summarization, 2005, pp. 65–72.
  54. [14] A. Ali and S. Renalds, “Word Error Rate Estimation
  55. for Speech Recognition: e-WER,” in Proceedings of
  56. the 56th Annual Meeting of the Association for
  57. Computational Linguistics (Short Papers), pp. 20–24
  58. Melbourne, Australia, July 15 - 20, 2018.
  59. [15] V. Timmel, C. Paonessa, M. Vogel, D. Perruchoud
  60. and R. Kakooe, “Fine-tuning Whisper on
  61. Low-Resource Languages for Real-World
  62. Applications,” arXiv preprint arXiv:2412.15726,
  63. Apr. 2025.
  64. [16] R. Ma, M. Qian, Y. Fathullah, S. Tang, M. Gales and
  65. K. Knill, “Cross-Lingual Transfer Learning for Speech
  66. Translation,” arXiv preprint, arXiv:2407.01130,
  67. Feb. 2025.
  68. [17] P. Peng, B. Yan, S. Watanabe and D. Harwath,
  69. “Prompting the hidden talent of web-scale speech
  70. models for zero-shot task generalization,” arXiv
  71. preprint, arXiv:2305.11095v3, Aug. 2023.