Data
Open data in, open data out. The database behind this site is free to download and reuse.
Krumeto's dictionary database is an adaptation of Wiktionary, Tatoeba, Open English WordNet, the CEFR-J and Octanove profiles, wordfreq and CMUdict. Because Wiktionary is CC BY-SA, our compiled database is published under the same license: Creative Commons Attribution-ShareAlike 4.0.
Download
The dump is produced by make dump in the repository and published with each data release. If no link appears here yet, the first public release has not been uploaded; write to us and we will send it.
What is inside
| Entries (lemma × part of speech) | 792,736 |
|---|---|
| Senses | 1,071,598 |
| Translations | 997,347 |
| Inflected forms | 1,038,342 |
| Sentences (all languages) | 5,577,551 |
| Sentence–lemma positions | 15,510,714 |
| WordNet synsets | 120,630 |
| Built | 2026-09-28 |
Per-word JSON
Every word is also available as JSON at /api/word/{word}, for example /api/word/decide. It is rate-limited and carries the license notice in the payload.
How to credit
“Data from Krumeto (CC BY-SA 4.0), compiled from Wiktionary, Tatoeba, Open English WordNet, CEFR-J/Octanove, wordfreq and CMUdict.”
Please keep the per-sentence contributor names from Tatoeba when you show sentences.