Northern Kurdish (Latin script) Speech to Text: A Complete Guide
Understanding Kurmancî (Northern Kurdish) Speech to Text
Kurmancî is the most widely spoken Kurdish dialect, used by millions in Turkey, Syria, Iraq, Iran, and the Kurdish diaspora. It is written in the Latin script in Turkey and Syria, while the Arabic script is used in Iraq and Iran. For transcription purposes, the Latin script version is common in digital media.
Why Accurate Kurdish Transcription Matters
Kurdish media is growing rapidly. From news channels like Rudaw to independent podcasts, there is a rising demand for subtitles and transcripts. Accurate speech-to-text helps:
- Make content accessible to deaf and hard-of-hearing viewers.
- Improve search engine visibility for Kurdish video content.
- Enable translation into other languages.
- Preserve oral history and linguistic data.
Transcription Challenges in Kurmancî
Kurmancî presents several challenges for automatic speech recognition:
Phonological Complexity
The language has a rich consonant inventory, including ejective stops (pʼ, tʼ, kʼ) and a three-way voicing distinction in some dialects. Vowel length is phonemic, meaning words like "gir" (heavy) and "gîr" (capture) are distinguished only by vowel duration.
Dialectal Variation
Major dialects include Botan, Hakkâri, Serhed, and Torî. Each has its own pronunciation patterns. For instance:
- Botan: Strong vowel harmony, influenced by Turkish.
- Hakkâri: Conservative, preserving older Kurdish sounds.
- Serhed: Many loanwords from Armenian and Turkish.
- Torî: Transitional between Kurmancî and Soranî.
Orthographic Consistency
While the Latin script is standard in Turkey, spelling conventions can vary. Some users write "ş" for /ʃ/, others use "sh". Speechyou's model is trained on multiple orthographic variants to handle this.
Practical Use Cases
- Podcasts and Radio: Transcribe Kurdish talk shows for show notes and accessibility.
- Film and Video: Generate SRT subtitles for Kurdish documentaries and films.
- Education: Create subtitles for Kurdish language courses.
- Research: Transcribe interviews for linguistic and sociological studies.
- Oral History: Convert spoken narratives from elders into searchable text.
How Speechyou Helps
Speechyou offers a dedicated Kurmancî speech-to-text engine that:
- Recognizes ejective consonants and vowel length accurately.
- Adapts to major dialects (Botan, Hakkâri, Serhed, Torî).
- Handles loanwords from Turkish, Arabic, and Persian.
- Exports SRT and VTT subtitle files.
- Works with uploaded audio/video or real-time recording.
With unlimited transcription on the Solo plan, Speechyou is a cost-effective solution for Kurdish content creators, educators, and researchers. Whether you need subtitles for a YouTube video or a transcript of an oral history interview, Speechyou delivers reliable results.
Getting Started
Simply upload your Kurdish audio or video file, select Kurmancî (Latin script) as the language, and let the AI process it. In minutes, you will have a transcript or subtitle file ready to download. No special training or technical skills required.







