Speech technology for Ethiopian languages. Built from the ground up.
Featured
Research Note
Whisper-small fine-tuned on Amharic runs out of output room around the twenty-fourth second of speech because each Ge'ez character costs three BPE tokens. A controlled study isolates the cause and shows the fix is preprocessing.
May 9, 2026
Announcement
We are releasing 14.7 hours of read Tigrinya speech from 161 speakers. Speaker-disjoint evaluation splits, gender-balanced dev and test sets, CC-BY-4.0. Phonetico's first public speech dataset.
May 6, 2026
The Problem
Ethiopia has over 80 languages and 120 million speakers. Most have no voice assistant. No transcription. Nothing.
Phonetico builds for these languages.
What We Build
Data collection, quality assurance, model training. We do all of it.
Offline-first app, works on rural connections. Speakers record, consent is tracked per recording, and contributors get paid.
Every recording is checked automatically. Signal quality, content match, fraud detection. Bad data doesn't get through.
We train and ship speech models. Research is published openly. A portion of every dataset goes public.
Where It Goes
Speech-to-text is where it starts. Everything else follows.
Live for Amharic, Afaan Oromo, Tigrinya, Sidama, Wolayta, and Sebat Bet Gurage. Used daily by Ethiopians at home and in the diaspora.
Try it on TelegramComing next. Accessibility, voice alerts, information for people who can't read.
Talk to your phone in Amharic. Ask it something in Afaan Oromo. That's where this goes.
From one Ethiopian language to another, or to international languages. Building bridges.
How We Work
Contributors earn roughly 12x Ethiopia's national median wage. Better pay, better data.
Ethiopian data stays in Ethiopian hands. Ethiopian-led team. Jobs in Ethiopia.
Every speaker knows exactly how their voice will be used. Consent is per-recording, with a full audit trail.
What we build here works for underserved languages anywhere. Four language families. Five countries. One starting point.
Help us build speech technology for your language.