Data

Open data in, open data out. The database behind this site is free to download and reuse.

Krumeto's dictionary database is an adaptation of Wiktionary, Tatoeba, Open English WordNet, the CEFR-J and Octanove profiles, wordfreq and CMUdict. Because Wiktionary is CC BY-SA, our compiled database is published under the same license: Creative Commons Attribution-ShareAlike 4.0.

Download

The dump is produced by make dump in the repository and published with each data release. If no link appears here yet, the first public release has not been uploaded; write to us and we will send it.

What is inside

Entries (lemma × part of speech)792,736
Senses1,071,598
Translations997,347
Inflected forms1,038,342
Sentences (all languages)5,577,551
Sentence–lemma positions15,510,714
WordNet synsets120,630
Built2026-09-28

Per-word JSON

Every word is also available as JSON at /api/word/{word}, for example /api/word/decide. It is rate-limited and carries the license notice in the payload.

How to credit

“Data from Krumeto (CC BY-SA 4.0), compiled from Wiktionary, Tatoeba, Open English WordNet, CEFR-J/Octanove, wordfreq and CMUdict.”

Please keep the per-sentence contributor names from Tatoeba when you show sentences.