Writing
A Quick Arabic TTS Comparison
In the name of Allah the most racious the most merciful بسم الله الرحمن الرحيم
While reviewing Microsoft’s AI models, the MAI Models, I tested MAI-Voice-2-Flash for text-to-speech. The result was not bad, especially considering that Arabic is not officially listed among the supported languages.
I remembered that I had tested some of the Gemini TTS models a few years ago, so I decided to do a quick comparison between these models, the well-known ElevenLabs platform, and Hakim.
I used a short sentence across all of them and decided to share the results in a brief article.
The text:
“يتنوعُ الأدبُ العربيُّ عبر عصوره الطويلةِ بين الشعر والنثر، ويشملُ نماذجَ خالدةً تعكسُ عبقريةَ اللسان العربي.”
| Provider | Model Name | Voice Name | Result |
|---|---|---|---|
| Microsoft | MAI-Voice-2-Flash | Klaus | Play |
| Microsoft | MAI-Voice-2-Flash | Bence | Play |
| ElevenLabs | Eleven-V3 | Chaouki | Play |
| ElevenLabs | Eleven-V3 | Hassan Mort | Play |
| Hakim | Hakim-Fast-V1 | Ali | Play |
| Hakim | Hakim-Fast-V1 | Amir | Play |
| Gemini-3.8-Flash-TTS | Algenib | Play | |
| Gemini-3.8-Flash-TTS | Neno | Play |
At some point, I started wondering what all of this means for people who work in voice recording, especially audiobook narrators.
I reassured myself with the thought that AI models may be able to reproduce the content, but not necessarily the feeling and emotion that a human voice can convey.