Why raw type-token ratio punishes long documents
The basic measure of vocabulary variety is type-token ratio: unique words divided by total words. It has a serious flaw that this tool works around rather than hides — TTR falls mechanically as text gets longer, regardless of the writer.
The reason is unavoidable. English forces repetition of function words: "the", "of", "and" recur every few words in any text. In a 100-word sample those repeats are a small share of the total; across 5,000 words they dominate, so TTR drops even if the writer's vocabulary is expanding. Comparing the TTR of a paragraph with the TTR of a chapter tells you which is longer, not which is more varied.
That is why this tool also reports MATTR — a moving-average type-token ratio computed over a sliding 50-word window and averaged. Because every window is the same length, the length bias cancels, and two documents of very different sizes become genuinely comparable. When TTR and MATTR disagree, MATTR is the one to trust.
Hapax legomena, and what they tell you
The tool also lists hapax legomena — words that appear exactly once in the text. In a natural document these are a large fraction of the vocabulary, and their character is informative in a way the ratios are not.
Skim the list. If the one-off words are precise nouns and specific verbs, the writing is doing real work. If they are near-synonyms of words used elsewhere — "utilise" alongside "use", "commence" alongside "start" — the diversity is coming from elegant variation rather than from saying more things, which usually reads worse rather than better.