Dataset.ET
Open community platform for building Ethiopian language AI datasets
Web app Rated 7 Oct 2026
- Link
- dataset.et
- Built by
Chapi Dev Talks
Float
343/500Where it landed on the scale
0
300
400
Summary
Dataset.ET is a working, actively shipped data-collection pipeline: a live Telegram bot, a website with a resources directory, airtime payouts, open Amharic and Afaan Oromoo speech datasets, a benchmark, a lexicon and an open ASR model. It has no business model, since it is a public-good project funded by credits and volunteers, so viability is plausible but unproven. Its edge is the contributor pipeline and the data it accumulates, and the many released artifacts show real commitment.
Float · Real and alive, but nothing stops a copy yet.
Breakdown
Five measures, 100 points each. Open the receipts under any of them to see the posts and pages behind the points.
Viability
Is there a real problem, someone who would pay, and a market this builder can actually reach?
The problem is real: Amharic and other Ethiopian languages have almost no open speech data. There is a believable path to funding through credits, grants and the demand for ASR/TTS. But there is no pricing and no payment provider, and the project itself says it is non-commercial, so no revenue model has been shown to work.
50/100
Rubric 41–60Plausible users and a believable way to charge them. Not yet shown to work.
Receipts (4)
States Amharic has close to no open speech data and releases 22.7 hours from 320 speakers, which shows the gap is real.
“We're releasing the first batch of Dataset.ET open speech data for Amharic. 22.7 hours. 7,405 recordings. 320 speakers. Free, CC BY 4.0, on Hugging Face. Amharic has close to no open speech data. That's the gap we're trying to close, and this is the first step rather than the finished thing. Some honesty about what this is: it's a first batch. The audio isn't preprocessed yet. We're not going t”
Post · @chapidevtalks · 26 Aug 2026 · open on Telegram (opens in a new tab)
Reports a $10k AWS credit via NVIDIA Inception, which shows external support but not revenue.
“BIG UPDATE ON DATASET.ET WE GOT A 10k USD from AWS from NVIDEA INCETPTION. This is gonna help us a looot.”
Post · @chapidevtalks · 28 Aug 2026 · open on Telegram (opens in a new tab)
The terms say datasets may be used commercially, and contributors are paid in airtime, but there is no paid product.
Checked 7 Oct 2026 · dataset.et/terms (opens in a new tab)
Asks for cloud credits or funding to scale, so the project is not self-sustaining yet.
“New dataset drop: Afaan Oromoo speech data is now live on Hugging Face! https://huggingface.co/datasets/snapwre/afaan-oromoo-speech Still early-stage, but every contribution helps us grow coverage for Afaan Oromoo alongside the other Ethiopian languages we're building for at Dataset.ET. If you or your org can help with cloud credits or funding to scale this up, please reach out even small suppo”
Post · @chapidevtalks · 30 Aug 2026 · open on Telegram (opens in a new tab)
Moat
What stops someone copying it?
There is local depth that takes real effort to copy: a validated community of voice contributors, an Ethio Telecom airtime reward loop, and specialised Amharic data such as the gemination lexicon and a held-out benchmark. The data is openly licensed, so it does not stay proprietary, and the edge is mostly the ongoing contributor pipeline and credibility in the field.
44/100
Rubric 41–60Local depth that takes real effort to copy: payment integrations, Amharic, regulation, partnerships, operational know-how.
Receipts (4)
Airtime payouts through Ethio Telecom are a locally specific contributor incentive.
“🇪🇹 Thank you, all of you ❤️ To our community, Thank you, truly, for lending us your voice and helping build AI for Ethiopian languages. 🙏 Today we come with a small token of thanks — airtime rewards are now live! You can turn the contributions you've made into Ethio Telecom airtime. Open the bot(@dataset_et_bot) and tap /redeem. Let's be honest with you: we're building this without big fu”
Post · @chapidevtalks · 29 Jun 2026 · open on Telegram (opens in a new tab)
An 86,022-entry Amharic gemination lexicon is specialised linguistic work that would take effort to reproduce.
“ገና can mean “still” or “Christmas.” Same spelling, different pronunciation. We’ve released the Amharic Gemination Lexicon at Dataset.ET to help capture these differences: which consonants are held longer when we speak. 86,022 word entries, with pronunciation patterns where available and unknown or ambiguous words flagged. Open under CC BY 4.0, with data, code and validation results. Built for ”
In Englishገና can mean "still" or "Christmas." Same spelling, different pronunciation. We've released the Amharic Gemination Lexicon at Dataset.ET to help capture these differences: which consonants are held longer when we speak. 86,022 word entries, with pronunciation patterns where available and unknown or ambiguous words flagged. Open under CC BY 4.0, with data, code and validation results. Built for p
Post · @chapidevtalks · 25 Sept 2026 · open on Telegram (opens in a new tab)
A benchmark on held-out recordings plus a custom Amharic language model shows technical depth beyond a wrapper.
“We tested every open Amharic speech recognition model. Nobody had measured them on data the models couldn't have already seen. So we did, on 1,548 recordings where we knew exactly what was said. Three things surprised us: • The smaller model beat the bigger one • A monolingual model beat the multilingual one • Two fair ways of counting mistakes disagreed about the winner Then we built a 538 MB”
Post · @chapidevtalks · 6 Sept 2026 · open on Telegram (opens in a new tab)
Open CC BY 4.0 release means the data itself is not defensible, only the pipeline that produces it.
“We're releasing the first batch of Dataset.ET open speech data for Amharic. 22.7 hours. 7,405 recordings. 320 speakers. Free, CC BY 4.0, on Hugging Face. Amharic has close to no open speech data. That's the gap we're trying to close, and this is the first step rather than the finished thing. Some honesty about what this is: it's a first batch. The audio isn't preprocessed yet. We're not going t”
Post · @chapidevtalks · 26 Aug 2026 · open on Telegram (opens in a new tab)
Momentum
Did the updates keep coming?
Worked on for about seven months, with updates in five of them and a steady rise in output from the March launch to the late-September releases.
89/100
- Sustained shipping60/60
Worked on across 7 months.
- Consistency29/40
Updates in 5 of 7 months.
Receipts (2)
Earliest post about it.
“🚀 Introducing Dataset.ET The future of AI should speak Ethiopian languages. Today we are launching Dataset.ET — an open community initiative to build the largest dataset for Ethiopian languages. Why this matters: Most AI systems today barely understand Amharic, Afaan Oromo, Tigrinya, Somali, and many other Ethiopian languages. Without datasets, AI will ignore our languages. Dataset.ET is chan”
Post · @chapidevtalks · 9 Mar 2026 · open on Telegram (opens in a new tab)
Most recent post about it.
“ገና can mean “still” or “Christmas.” Same spelling, different pronunciation. We’ve released the Amharic Gemination Lexicon at Dataset.ET to help capture these differences: which consonants are held longer when we speak. 86,022 word entries, with pronunciation patterns where available and unknown or ambiguous words flagged. Open under CC BY 4.0, with data, code and validation results. Built for ”
In Englishገና can mean "still" or "Christmas." Same spelling, different pronunciation. We've released the Amharic Gemination Lexicon at Dataset.ET to help capture these differences: which consonants are held longer when we speak. 86,022 word entries, with pronunciation patterns where available and unknown or ambiguous words flagged. Open under CC BY 4.0, with data, code and validation results. Built for p
Post · @chapidevtalks · 25 Sept 2026 · open on Telegram (opens in a new tab)
Infrastructure
Does a real, working product exist?
A real, working product exists: a live contribution and validation bot, a site with a resources directory, published speech datasets, an open ASR model with a demo, and airtime redemption.
92/100
- It's live20/20
https://dataset.et responded when checked.
- HTTPS5/5
Served over HTTPS.
- Real product35/35
A full working product: a live Telegram bot for recording and validation, a website with a resources directory, and published datasets on Hugging Face. There is also a working ASR demo and model, and a redeem flow for airtime.
- Own home5/10
A Telegram bot, no home of its own.
- Maintained15/15
Last sign of shipping 2026-09-25.
- Operations12/15
There are real operations: a bot with contribution, validation and redemption flows, privacy and terms pages, open data releases with documentation, a model demo and a resources directory, and more than one platform (Telegram, web, Hugging Face). The crawler found no status page or API docs.
Receipts (5)
https://dataset.et responded when checked.
Checked 7 Oct 2026 · dataset.et (opens in a new tab)
Published a real Amharic speech dataset with 7,405 recordings from 320 speakers.
“We're releasing the first batch of Dataset.ET open speech data for Amharic. 22.7 hours. 7,405 recordings. 320 speakers. Free, CC BY 4.0, on Hugging Face. Amharic has close to no open speech data. That's the gap we're trying to close, and this is the first step rather than the finished thing. Some honesty about what this is: it's a first batch. The audio isn't preprocessed yet. We're not going t”
Post · @chapidevtalks · 26 Aug 2026 · open on Telegram (opens in a new tab)
A public ASR demo and benchmark, all usable.
“We tested every open Amharic speech recognition model. Nobody had measured them on data the models couldn't have already seen. So we did, on 1,548 recordings where we knew exactly what was said. Three things surprised us: • The smaller model beat the bigger one • A monolingual model beat the multilingual one • Two fair ways of counting mistakes disagreed about the winner Then we built a 538 MB”
Post · @chapidevtalks · 6 Sept 2026 · open on Telegram (opens in a new tab)
Hohe ASR is live at dataset.et/ai and via a bot.
“The Fact that hohe ASR is running on 4VCPU is insane honestly and so small parameter 0.6B and it performs very well. anytime try it on https://dataset.et/ai or @dataset_ai_bot AND OPEN SOURCE 🤯”
Post · @chapidevtalks · 21 Sept 2026 · open on Telegram (opens in a new tab)
The Telegram bot is live.
Checked 7 Oct 2026 · t.me/dataset_et_bot (opens in a new tab)
Sustainability
Is it being set up to last?
Distribution through a Telegram bot works for collecting data, and contributor payouts and open releases show commitment. But there is no payment provider or pricing, funding comes from credits, and the project itself says it is non-commercial, so sustainability is unproven.
68/100
- Distribution20/30
A Telegram bot or mini app.
- Payments0/15
No payment provider found.
- Pricing10/10
Prices stated in the creator's posts.
- Revenue1/5
There is no customer revenue. The only money mentioned is a $10k cloud credit and a request for funding, and the project pays contributors rather than being paid.
- Privacy policy7/7
Has a privacy policy.
- Terms5/5
Has terms of service.
- Support8/8
A way to reach support.
- Commitment17/20
Seven months of steady shipping, a dedicated community channel and bot, a team referenced in posts, a kept path from collection to a first dataset, a benchmark, a lexicon and an ASR model, plus cloud credits through NVIDIA Inception. It is still small and recruiting partnerships.
Receipts (6)
A Telegram bot or mini app.
Checked 7 Oct 2026 · t.me/dataset_et_bot (opens in a new tab)
Has a privacy policy.
Checked 7 Oct 2026 · dataset.et/privacy (opens in a new tab)
A $10k credit from AWS/NVIDIA Inception is support, not paying customers.
“BIG UPDATE ON DATASET.ET WE GOT A 10k USD from AWS from NVIDEA INCETPTION. This is gonna help us a looot.”
Post · @chapidevtalks · 28 Aug 2026 · open on Telegram (opens in a new tab)
Says it is building without big funding.
“🇪🇹 Thank you, all of you ❤️ To our community, Thank you, truly, for lending us your voice and helping build AI for Ethiopian languages. 🙏 Today we come with a small token of thanks — airtime rewards are now live! You can turn the contributions you've made into Ethio Telecom airtime. Open the bot(@dataset_et_bot) and tap /redeem. Let's be honest with you: we're building this without big fu”
Post · @chapidevtalks · 29 Jun 2026 · open on Telegram (opens in a new tab)
Launched with a dedicated bot and community channel.
“🚀 Introducing Dataset.ET The future of AI should speak Ethiopian languages. Today we are launching Dataset.ET — an open community initiative to build the largest dataset for Ethiopian languages. Why this matters: Most AI systems today barely understand Amharic, Afaan Oromo, Tigrinya, Somali, and many other Ethiopian languages. Without datasets, AI will ignore our languages. Dataset.ET is chan”
Post · @chapidevtalks · 9 Mar 2026 · open on Telegram (opens in a new tab)
Promised preprocessing and a real pipeline in the next release, a roadmap stated openly.
“We're releasing the first batch of Dataset.ET open speech data for Amharic. 22.7 hours. 7,405 recordings. 320 speakers. Free, CC BY 4.0, on Hugging Face. Amharic has close to no open speech data. That's the gap we're trying to close, and this is the first step rather than the finished thing. Some honesty about what this is: it's a first batch. The audio isn't preprocessed yet. We're not going t”
Post · @chapidevtalks · 26 Aug 2026 · open on Telegram (opens in a new tab)