Writing

A Quick Arabic TTS Comparison


In the name of Allah the most racious the most merciful بسم الله الرحمن الرحيم

While reviewing Microsoft’s AI models, the MAI Models, I tested MAI-Voice-2-Flash for text-to-speech. The result was not bad, especially considering that Arabic is not officially listed among the supported languages.

I remembered that I had tested some of the Gemini TTS models a few years ago, so I decided to do a quick comparison between these models, the well-known ElevenLabs platform, and Hakim.

I used a short sentence across all of them and decided to share the results in a brief article.

The text:

“يتنوعُ الأدبُ العربيُّ عبر عصوره الطويلةِ بين الشعر والنثر، ويشملُ نماذجَ خالدةً تعكسُ عبقريةَ اللسان العربي.”

Provider Model Name Voice Name Result
Microsoft MAI-Voice-2-Flash Klaus Play
Microsoft MAI-Voice-2-Flash Bence Play
ElevenLabs Eleven-V3 Chaouki Play
ElevenLabs Eleven-V3 Hassan Mort Play
Hakim Hakim-Fast-V1 Ali Play
Hakim Hakim-Fast-V1 Amir Play
Google Gemini-3.8-Flash-TTS Algenib Play
Google Gemini-3.8-Flash-TTS Neno Play

At some point, I started wondering what all of this means for people who work in voice recording, especially audiobook narrators.

I reassured myself with the thought that AI models may be able to reproduce the content, but not necessarily the feeling and emotion that a human voice can convey.