Sources

Krumeto is built only from openly licensed data. Every dataset, its license, the date we downloaded it, and the citation its authors ask for.

Wiktionary (English edition)

Senses, parts of speech, inflected forms, IPA, audio file references, per-sense translations, synonyms, antonyms, derived and related terms, etymology, hyphenation. Obtained through the kaikki.org machine-readable export (wiktextract). Downloaded —.

© Wiktionary contributors, CC BY-SA 4.0. Every word page links to its Wiktionary entry. Our compiled dataset is shared under the same license (see below).

Tatoeba

Example sentences and their translations into Ukrainian, Russian, German, Spanish, French, Italian, Polish, Turkish and Portuguese, with the contributor of every sentence. Downloaded — from tatoeba.org.

Sentence text is licensed CC BY 2.0 FR (some sentences CC0). Each sentence shows its contributor and links to its Tatoeba page. Tatoeba audio is not used.

Open English WordNet 2024

Synonym sets, definitions, antonyms, derivational links and example phrases. github.com/globalwordnet/english-wordnet, downloaded —. CC BY 4.0.

Citation: John P. McCrae, Alexandre Rademaker, Francis Bond, Ewa Rudnicka and Christiane Fellbaum (2019). English WordNet 2019 – An Open-Source WordNet for English. Proceedings of the 10th Global WordNet Conference.

CEFR-J Vocabulary Profile 1.5 and Octanove Vocabulary Profile C1/C2

Word levels A1–B2 come from the CEFR-J Vocabulary Profile (© Tono Lab, Tokyo University of Foreign Studies), free for research and commercial use with citation. C1/C2 levels come from the Octanove Vocabulary Profile, CC BY-SA 4.0. Both from github.com/openlanguageprofiles/olp-en-cefrj, downloaded —.

Citation: Tono, Yukio (ed.). CEFR-J Wordlist Version 1.5. Tokyo University of Foreign Studies. The ordering of the grammar library follows the CEFR-J Grammar Profile; all lesson text is our own.

wordfreq

Word frequencies (Zipf scale) from Robyn Speer's wordfreq (data frozen at 2021). Code MIT; data derived from openly licensed corpora as described in its repository.

CMUdict

Fallback US pronunciations from the CMU Pronouncing Dictionary, BSD-2-Clause, downloaded —. ARPAbet is converted to IPA by us.

Wikimedia Commons audio

Recorded pronunciations are played directly from Wikimedia Commons (not copied to our servers). Each recording shows its author and license; recordings whose license could not be resolved are not shown.

Fonts and software

Instrument Serif (Rodrigo Fuenzalida, Jordan Egstad), Newsreader (Production Type), Inter (Rasmus Andersson) and Noto Sans (Google) under the SIL Open Font License 1.1, self-hosted. The site runs on Next.js, React, Tailwind CSS, better-sqlite3 and SQLite, all under permissive open-source licenses.

Our dataset

The database compiled from the sources above (built 2026-09-28) is published under CC BY-SA 4.0 on the Data page.

Krumeto is not affiliated with the Wikimedia Foundation or Tatoeba.