← wordmathSources
English vocabulary
The complete large English vocabulary from wordfreq by Robyn Speer. Frequency estimates reflect language use through approximately 2021. The list includes names, inflected words, uncommon spellings and corpus artefacts; inclusion is not an endorsement or a dictionary definition.
wordfreq combines Wikipedia, SUBTLEX, OPUS OpenSubtitles, Google Books Ngrams, news, web and other corpora. We acknowledge the SUBTLEX authors Marc Brysbaert, Boris New, Emmanuel Keuleers and their collaborators, and OpenSubtitles. SUBTLEX is freely available data.
wordfreq software is Apache-licensed; its included data may be redistributed under CC BY-SA 4.0, with additional source attribution described in the NOTICE and licensing notes. We preserve the original source notices in the project’s build records. Text is normalized with Unicode NFKC, lowercasing and whitespace normalization for embedding and matching.
Wikipedia titles
Collections labelled “Wikipedia Vital Articles” include the Level 5 article list, selected and maintained by Wikipedia contributors. Article titles are resolved to their current canonical names, normalized, and deduplicated against the English wordlist. Source revisions and redirect aliases are saved with the build. Wikipedia content is available under CC BY-SA 4.0.
Embeddings and calculation
OpenAI text-embedding-3-large, 1,024 dimensions. Each input vector is normalized to length one. The calculator computes A − B + C and ranks indexed candidates using cosine similarity, excluding the input terms and duplicate spellings. New input embeddings and repeated results are cached on the server. No language model writes or chooses the answers.
Similarity is not confidence, factual correctness, or proof of an analogy. Results can reflect ambiguity and stereotypes in language. All displayed matches come from the stated answer collection.