Uzbek speech to text: why global engines fail and what actually works

Try dictating Uzbek to Siri or running an Uzbek voice message through a generic transcription service. The result is usually a word salad. Global speech engines are trained mostly on English and a handful of major languages; Uzbek, Kazakh and Kyrgyz are agglutinative, data-poor and heavily accented across regions, so error rates climb above 30% - every third word is wrong.

sezish is a speech-to-text service built specifically for Central Asian languages. Its model is trained for the region and holds around a 9% word error rate on Uzbek, while also recognizing Kazakh, Kyrgyz, Russian and English.

Three ways to use it

Telegram bot. Forward any voice message to @sezishbot and get text back. Free for 10 minutes a day. Add the bot to a group chat and it transcribes every voice message automatically. Each transcript comes with translation buttons for six languages, including Karakalpak - a language most translation services simply do not cover, despite its two million speakers.

Desktop app. The Mac app does system-wide dictation and meeting transcription. Everything runs locally: the model is a 225 MB download and works fully offline. Meeting recordings never leave your computer - no cloud, no third parties.

Mobile. The iPhone app is live on the App Store (free), with live conversation translation, dictation and meeting notes. The Android version is landing on Google Play now.

Privacy

Audio sent to the server travels over TLS, gets transcribed in memory and is not stored: the server returns text and forgets the audio. Nothing is used for model training. For fully offline work, use the desktop app. Details are in the privacy policy on sezi.sh.

You can test the recognition right on the sezi.sh home page - press the record button and say something in Uzbek.

Read more