Addis AI Models
Open-sourced early models and Amharic-tuned Whisper models for on-device speech recognition
Library Rated 7 Oct 2026
- Built by
Biniyam &
Chapi Dev Talks &
Henok | Neural Nets
Float
335/500Where it landed on the scale
0
300
400
Summary
This is a steady, credible body of open Amharic AI work: models, datasets, a benchmark, a live demo and a Docker server anyone can run. The Addis AI product itself hasn't shipped yet, and nothing shows anyone paying, so viability stays at the plausible stage. The moat comes from Amharic-specific data and know-how rather than from the code.
Float · Real and alive, but nothing stops a copy yet.
Breakdown
Five measures, 100 points each. Open the receipts under any of them to see the posts and pages behind the points.
Viability
Is there a real problem, someone who would pay, and a market this builder can actually reach?
The gap is real, since Amharic has little open speech tooling, and an API or bot could plausibly be charged for. But everything shown is free under CC BY, and Addis AI is only described as coming, so there's no tested way to charge yet.
40/100
Rubric 21–40A real problem, but no plausible path to anyone paying, or a market well out of this builder's reach.
Receipts (3)
States Amharic has close to no open speech data, which defines the problem, and releases the data free.
“We're releasing the first batch of Dataset.ET open speech data for Amharic. 22.7 hours. 7,405 recordings. 320 speakers. Free, CC BY 4.0, on Hugging Face. Amharic has close to no open speech data. That's the gap we're trying to close, and this is the first step rather than the finished thing. Some honesty about what this is: it's a first batch. The audio isn't preprocessed yet. We're not going t”
Post · @chapidevtalks · 26 Aug 2026 · open on Telegram (opens in a new tab)
Hohe ASR runs on cheap CPUs and is offered free through a bot, but no pricing or paying users are shown.
“Hohe ASR is out. Open Amharic speech to text. Start with the speed, because it is the part people do not expect. On four ordinary CPU cores, a five second voice note is transcribed in 0.39 seconds. A thirty second clip takes 7.75 seconds. An hour of audio goes through in under ten minutes, on a machine that costs about nine cents an hour to rent. No GPU anywhere in that sentence. It is a CTC mod”
Post · @chapidevtalks · 21 Sept 2026 · open on Telegram (opens in a new tab)
Addis AI is mentioned as the product being built, with details promised later and no release yet.
“I've recently open-sourced some of the early models I've been experimenting with while building Addis AI - there are some models with exact training params and training details... along with our amharic tuned whisper models that can actually run on your phone (more on that in the coming weeks) But feel free on experimenting, seeing how they work and shoot me your feedback. Everything Is on huggin”
Post · @b1n1yamBuilds · 6 Dec 2025 · open on Telegram (opens in a new tab)
Moat
What stops someone copying it?
Amharic-specific depth would take real effort to copy: 880 hours of training audio with regional dialects, a contributor-reviewed speech corpus, a held-out benchmark and a gemination lexicon. Because the weights and data are open, the edge is the accumulated data and know-how rather than exclusivity.
50/100
Rubric 41–60Local depth that takes real effort to copy: payment integrations, Amharic, regulation, partnerships, operational know-how.
Receipts (4)
880 hours of training audio, including five dialects and phone-line audio, with honest error figures.
“Hohe ASR is out. Open Amharic speech to text. Start with the speed, because it is the part people do not expect. On four ordinary CPU cores, a five second voice note is transcribed in 0.39 seconds. A thirty second clip takes 7.75 seconds. An hour of audio goes through in under ten minutes, on a machine that costs about nine cents an hour to rent. No GPU anywhere in that sentence. It is a CTC mod”
Post · @chapidevtalks · 21 Sept 2026 · open on Telegram (opens in a new tab)
320 speakers contributed voices and reviewed each other's recordings, a data-collection effort that's hard to replicate.
“We're releasing the first batch of Dataset.ET open speech data for Amharic. 22.7 hours. 7,405 recordings. 320 speakers. Free, CC BY 4.0, on Hugging Face. Amharic has close to no open speech data. That's the gap we're trying to close, and this is the first step rather than the finished thing. Some honesty about what this is: it's a first batch. The audio isn't preprocessed yet. We're not going t”
Post · @chapidevtalks · 26 Aug 2026 · open on Telegram (opens in a new tab)
An 86,022-entry Amharic gemination lexicon is a niche linguistic resource.
“ገና can mean “still” or “Christmas.” Same spelling, different pronunciation. We’ve released the Amharic Gemination Lexicon at Dataset.ET to help capture these differences: which consonants are held longer when we speak. 86,022 word entries, with pronunciation patterns where available and unknown or ambiguous words flagged. Open under CC BY 4.0, with data, code and validation results. Built for ”
In Englishገና can mean "still" or "Christmas." Same spelling, different pronunciation. We've released the Amharic Gemination Lexicon at Dataset.ET to help capture these differences: which consonants are held longer when we speak. 86,022 word entries, with pronunciation patterns where available and unknown or ambiguous words flagged. Open under CC BY 4.0, with data, code and validation results. Built for p
Post · @chapidevtalks · 25 Sept 2026 · open on Telegram (opens in a new tab)
Built a benchmark on recordings the models couldn't have seen, plus a 538 MB Amharic language model.
“We tested every open Amharic speech recognition model. Nobody had measured them on data the models couldn't have already seen. So we did, on 1,548 recordings where we knew exactly what was said. Three things surprised us: • The smaller model beat the bigger one • A monolingual model beat the multilingual one • Two fair ways of counting mistakes disagreed about the winner Then we built a 538 MB”
Post · @chapidevtalks · 6 Sept 2026 · open on Telegram (opens in a new tab)
Momentum
Did the updates keep coming?
Worked on across 27 months with updates in 12 of them, and activity was still recent as of late September 2026. Output picked up a lot from December 2025, with a burst of releases in August and September 2026.
78/100
- Sustained shipping60/60
Worked on across 27 months.
- Consistency18/40
Updates in 12 of 27 months.
Receipts (2)
Earliest post about it.
“You know Llama right, so let me introduce you to LLaMAX specifically trained for translation and it includes Amharic and Afan Oromo from Ethiopian languages. It says they used millions of Amharic sentences and thousands of Afan Oromo sentences, I'll test the models and see how they do later. Paper Model on HF”
Post · @neural_netss · 12 Jul 2024 · open on Telegram (opens in a new tab)
Most recent post about it.
“ገና can mean “still” or “Christmas.” Same spelling, different pronunciation. We’ve released the Amharic Gemination Lexicon at Dataset.ET to help capture these differences: which consonants are held longer when we speak. 86,022 word entries, with pronunciation patterns where available and unknown or ambiguous words flagged. Open under CC BY 4.0, with data, code and validation results. Built for ”
In Englishገና can mean "still" or "Christmas." Same spelling, different pronunciation. We've released the Amharic Gemination Lexicon at Dataset.ET to help capture these differences: which consonants are held longer when we speak. 86,022 word entries, with pronunciation patterns where available and unknown or ambiguous words flagged. Open under CC BY 4.0, with data, code and validation results. Built for p
Post · @chapidevtalks · 25 Sept 2026 · open on Telegram (opens in a new tab)
Infrastructure
Does a real, working product exist?
Real, working artifacts exist: downloadable Amharic ASR and LLM models, a live demo, a benchmark and a Docker server with documented deployment. Addis AI as a finished product doesn't exist yet, and the early models are experimental.
85/100
- It's live20/20
https://huggingface.co/b1n1yam responded when checked.
- HTTPS5/5
Served over HTTPS.
- Real product25/35
Working models, a live demo space and a runnable ASR server exist, and the author documents the weaknesses. The Addis AI product itself isn't out and several items are experimental or early, so it's working but still thin as a product.
- Own home10/10
Lives on its own domain (huggingface.co).
- Maintained15/15
Last sign of shipping 2026-09-25.
- Operations10/15
There's real operational material: a deployable server, an AWS spot and load balancer setup, a Telegram bot, a GitHub repo, measured speed numbers and multiple hosting platforms. There's no status page or accounts system for the project itself.
Receipts (5)
https://huggingface.co/b1n1yam responded when checked.
Checked 7 Oct 2026 · huggingface.co/b1n1yam (opens in a new tab)
Hohe ASR ships with an open server that runs from three Docker commands and streams long audio.
“Hohe ASR now ships with the server we run ourselves, and it is open. Three commands and you have Amharic speech to text running on your own machine. No GPU, no account, no API key, nothing to sign. git clone https://github.com/snapwre/hohe-serve && cd hohe-serve docker build -t hohe-asr . && docker run --rm -p 8080:8080 hohe-asr curl -F audio=@clip.ogg http://localhost:8080/transcribe The ima”
Post · @chapidevtalks · 21 Sept 2026 · open on Telegram (opens in a new tab)
Public demo space and benchmark let anyone try and check the models.
“We tested every open Amharic speech recognition model. Nobody had measured them on data the models couldn't have already seen. So we did, on 1,548 recordings where we knew exactly what was said. Three things surprised us: • The smaller model beat the bigger one • A monolingual model beat the multilingual one • Two fair ways of counting mistakes disagreed about the winner Then we built a 538 MB”
Post · @chapidevtalks · 6 Sept 2026 · open on Telegram (opens in a new tab)
The profile lists 32 models, including Amharic ASR and LLM checkpoints, and several dataset collections.
Checked 7 Oct 2026 · huggingface.co/b1n1yam (opens in a new tab)
Repo documents AWS spot deployment behind a load balancer and publishes the speed measurements.
“Hohe ASR now ships with the server we run ourselves, and it is open. Three commands and you have Amharic speech to text running on your own machine. No GPU, no account, no API key, nothing to sign. git clone https://github.com/snapwre/hohe-serve && cd hohe-serve docker build -t hohe-asr . && docker run --rm -p 8080:8080 hohe-asr curl -F audio=@clip.ogg http://localhost:8080/transcribe The ima”
Post · @chapidevtalks · 21 Sept 2026 · open on Telegram (opens in a new tab)
Sustainability
Is it being set up to last?
Distribution is open weights and a Telegram bot, with no revenue or price points and an explicit call for funding. The Amharic data and benchmark work show seriousness, but there's no business model yet.
82/100
- Distribution25/30
A web app on its own domain.
- Payments15/15
Payments wired in: Stripe.
- Pricing10/10
A pricing page or stated prices on the site.
- Revenue0/5
No post mentions customers, sales or income; everything is released free under CC BY. The Stripe and pricing flags in the crawler checks come from Hugging Face's own site, not from this project.
- Privacy policy7/7
Has a privacy policy.
- Terms5/5
Has terms of service.
- Support8/8
A way to reach support.
- Commitment12/20
There's a small team (Addis AI with a co-builder, the Dataset.ET group), a promised next release with a proper data pipeline, a project bot and 320 volunteer contributors. There's no funding or accelerator, and Addis AI promises haven't been delivered yet.
Receipts (6)
A web app on its own domain.
Checked 7 Oct 2026 · huggingface.co/b1n1yam (opens in a new tab)
Payments wired in: Stripe.
Checked 7 Oct 2026 · huggingface.co/b1n1yam (opens in a new tab)
Has a privacy policy.
Checked 7 Oct 2026 · huggingface.co/privacy (opens in a new tab)
Offered free to try and open to download, with no paid tier mentioned.
“Hohe ASR is out. Open Amharic speech to text. Start with the speed, because it is the part people do not expect. On four ordinary CPU cores, a five second voice note is transcribed in 0.39 seconds. A thirty second clip takes 7.75 seconds. An hour of audio goes through in under ten minutes, on a machine that costs about nine cents an hour to rent. No GPU anywhere in that sentence. It is a CTC mod”
Post · @chapidevtalks · 21 Sept 2026 · open on Telegram (opens in a new tab)
Asks for cloud credits or funding, which suggests there's no revenue yet.
“New dataset drop: Afaan Oromoo speech data is now live on Hugging Face! https://huggingface.co/datasets/snapwre/afaan-oromoo-speech Still early-stage, but every contribution helps us grow coverage for Afaan Oromoo alongside the other Ethiopian languages we're building for at Dataset.ET. If you or your org can help with cloud credits or funding to scale this up, please reach out even small suppo”
Post · @chapidevtalks · 30 Aug 2026 · open on Telegram (opens in a new tab)
Dataset.ET team releases a first batch and commits to a preprocessed, validated next release.
“We're releasing the first batch of Dataset.ET open speech data for Amharic. 22.7 hours. 7,405 recordings. 320 speakers. Free, CC BY 4.0, on Hugging Face. Amharic has close to no open speech data. That's the gap we're trying to close, and this is the first step rather than the finished thing. Some honesty about what this is: it's a first batch. The audio isn't preprocessed yet. We're not going t”
Post · @chapidevtalks · 26 Aug 2026 · open on Telegram (opens in a new tab)