Desano Speech to Text: A Complete Guide
Desano Speech-to-Text: A New Tool for an Endangered Amazonian Language
Desano is a Tucanoan language spoken by approximately 2,000 people in the Vaupés region of Colombia and the neighboring Amazonas state of Brazil. It is a language rich in tone, nasality, and vowel length, features that are critical for distinguishing meaning. For example, the word 'pĩ' (with a high tone and nasalization) means 'tobacco', while 'pi' (low tone, oral) means 'to pull'. Accurate speech-to-text for Desano must capture these nuances.
Why Accurate Transcription Matters
Desano is an endangered language, with most speakers over 40 years old. Younger generations are increasingly fluent in Spanish or Portuguese. Tools that transcribe Desano audio into text help preserve the language by creating written records of oral traditions, songs, and everyday speech. They also support bilingual education, where Desano is taught alongside the national language.
Challenges in Desano Speech Recognition
Building an ASR system for Desano comes with unique hurdles:
- Tonal distinctions: High and low tones change the meaning of words. Our model uses tonal embeddings to differentiate them.
- Nasalization: Vowels and some consonants are nasalized, requiring the model to detect nasal airflow features.
- Vowel length: Long and short vowels are phonemic. Time-based acoustic features are used to capture duration.
- Data scarcity: Only a few hours of transcribed Desano speech exist publicly. Speechyou employs transfer learning from related Tucanoan languages and synthetic data augmentation.
Despite these challenges, Speechyou achieves over 90% accuracy on clear recordings, and the system improves with user feedback.
Use Cases for Desano Transcription
- Oral history preservation: Elders can record stories and songs, which are automatically transcribed for archives.
- Subtitles for videos: Desano-language YouTube channels and community videos can add subtitles in Desano or Spanish/Portuguese.
- Language documentation: Linguists transcribe field recordings, including ritual language and conversations.
- Accessibility: Deaf or hard-of-hearing Desano speakers can access video content through captions.
- Podcast and radio: Indigenous media producers can publish searchable transcripts of their shows.
How Speechyou Helps
Speechyou is the first AI transcription service to support Desano specifically. Our model handles the tonal and nasal features of the language, and we export standard SRT and VTT subtitle files. The interface is simple: upload an audio or video file, select Desano, and receive a transcription with time codes. You can edit the text and download subtitles for any video platform.
For communities and researchers, Speechyou offers a free tier with unlimited transcription for solo users. This makes it accessible for non-profit language documentation projects. We also support custom fine-tuning for enterprise users who need higher accuracy for specific dialects or speakers.
Preserving Desano in the Digital Age
By providing reliable speech-to-text for Desano, Speechyou empowers speakers to document their language, educators to create teaching materials, and researchers to analyze the language in depth. It is a small but meaningful step toward ensuring that Desano continues to be spoken and written for generations to come.







