Mazandarani (Latin script) Speech to Text: A Complete Guide
Mazandarani Speech to Text: Transcribing an Iranian Heritage Language
Mazandarani (مازندرانی), also known as Tabari, is a Northwestern Iranian language spoken primarily in the province of Mazandaran along the Caspian Sea. With an estimated 3–4 million native speakers, it is the most widely spoken minority language in Iran after Azeri and Kurdish. Yet unlike those languages, Mazandarani has received very little attention from speech technology developers. This is now changing with Speechyou, a dedicated AI-powered transcription service that converts Mazandarani audio and video into accurate Latin-script text with subtitles in SRT and VTT formats.
Where Is Mazandarani Spoken?
The language area stretches roughly from the Alborz mountain range to the Caspian coast, including major cities such as Sari, Babol, Amol, Gorgan (in adjacent Golestan province), and the capital city of Mazandaran, Sari. A diaspora also exists in Tehran and other Iranian cities. Mazandarani belongs to the Caspian branch of Iranian languages and is closely related to Gilaki. Its writing system has historically been the Perso-Arabic script (Nastaliq style), but in recent decades a Latin-based orthography has emerged, especially on digital platforms and among younger speakers — this is the script supported by Speechyou.
Why Accurate Transcription Matters
For Mazandarani speakers, the lack of speech-to-text tools means lost opportunities:
- Cultural preservation: Oral traditions, folk music, and elderly voices at risk of being lost.
- Media accessibility: Mazandarani-language YouTube channels, podcasts, and local TV programs cannot provide subtitles for non-speakers or the hearing impaired.
- Research: Linguists studying the language's unique phonology and grammar often transcribe hours of fieldwork manually.
- Education: Schools in Mazandaran that teach the language (often as a second elective) need ready-made subtitled materials.
Speechyou fills this gap by offering automatic, high-accuracy transcription tailored to Mazandarani's specific features.
Specific Transcription Challenges
Mazandarani presents several difficulties for automated speech recognition:
1. Phonological Uniqueness
Mazandarani retains consonants that have shifted in Persian, such as /d/ and /g/ in words like dend (teeth) vs. Persian dandân. It also has front rounded vowels /y/ and /ø/ (as in pür – son) and a uvular fricative /ʁ/. Generic ASR models often map these to Persian or other sounds, causing errors.
2. Dialect Variation
Three main dialect groups exist — Central (Sari), Western (Babol), and Eastern (Gorgan) — with noticeable phonetic differences. For instance, the vowel in “water” is /u/ in Sari, /o/ in Babol, and /uː/ in Gorgan. Speechyou's model is trained on data from all three regions.
3. Code-Switching
Urban speakers frequently alternate between Mazandarani and Persian within single sentences. A robust ASR must recognize both languages and correctly assign the output script. Speechyou uses a bilingual language identification module that tags each segment, ensuring Persian words are romanized consistently.
4. Limited Training Data
Mazandarani is a low-resource language, meaning few transcribed speech corpora are publicly available. Speechyou overcomes this through transfer learning — first training on Persian and other Iranian languages, then fine-tuning on a custom collected Mazandarani dataset of over 200 hours of read and spontaneous speech.
Use Cases in Action
- Podcast transcriptions: Convert episodes of popular Mazandarani podcast Mazandaran Nâ into text for show notes and SEO.
- Subtitle generation: A local filmmaker adding SRT subtitles to her Mazandarani-language documentary on Caspian ecology.
- Academic research: A PhD student transcribing 40 hours of interviews with Mazandarani weavers for a dissertation on textile heritage.
- Oral history: An NGO digitizing interviews with elderly speakers of the endangered Kaleh Dashi dialect.
- Accessibility: A university adding live captions for a Mazandarani poetry night attended by deaf community members.
How Speechyou Helps
Speechyou is the only commercial speech-to-text service that explicitly supports Mazandarani in Latin script. It delivers:
- Automatic transcription with timestamps (SRT/VTT)
- Support for multiple dialects
- Real-time and batch processing
- Unlimited transcription on the Solo plan
- Export to plain text, Word, or subtitle formats
Using Speechyou, you can transcribe a 2-hour Mazandarani interview in minutes — a task that would take a human transcriber 6–8 hours. The accuracy exceeds 95% for clear studio recordings and drops only moderately (to ~85%) for noisy field recordings, thanks to adaptive denoising.
Getting Started
To transcribe your Mazandarani content, simply upload your audio or video file to Speechyou, select "Mazandarani (Latin script)" as the language, and receive your transcript and subtitles in minutes. Whether you're a linguist, content creator, or community activist, accurate Mazandarani speech recognition is now just a click away.







