Skip to content
Redmoon Calculators
Text analysis

Lexical Density Calculator

Estimate lexical density: the ratio of content words to total words. Higher density = denser, more academic text.

Built and maintained by Paul Clark, Redmoon Software

When to use this

Use lexical density to compare register or formality across texts. High density (>55%) suggests dense, information-rich writing; low density (<40%) suggests conversational or narrative prose. Especially useful for academic writing coaches and editorial benchmarking.

How it compares

Density differs from lexical diversity (TTR): density asks "how much information per word?", diversity asks "how varied is your vocabulary?". Report both for a complete vocabulary picture.

0 chars

How it works

Lexical density is the proportion of content words (nouns, verbs, adjectives, adverbs) to total words.

Higher density is typical of formal, academic, or technical text. Lower density indicates more conversational or narrative prose.

This tool approximates content vs. function words using an English stopword list — sufficient for most diagnostic use.

Formula

FAQs

What is lexical density?

The proportion of content words (nouns, verbs, adjectives, adverbs) to total words. Function words (the, of, and) are excluded.

What is a typical range?

Speech: 35-45%. Conversational writing: 40-50%. News articles: 50-55%. Academic writing: 55-65%.

What counts as a content word?

Content words are nouns, main verbs, adjectives, and adverbs that carry meaning, while function words are articles, prepositions, conjunctions, pronouns, and auxiliary verbs. Lexical density is the share of content words in the total.

Does high lexical density mean better writing?

No. High density signals information-packed, formal text but can be harder to read, while lower density with more function words often flows more naturally in speech and conversational writing. Aim for a level that suits your audience, not the highest number.

Worked example

Input

The careful analysis of dense academic prose reveals striking lexical patterns.

Output

Lexical density: 64% — Academic / formal.

Of 11 words, 7 are content words (nouns, verbs, adjectives, adverbs). 7/11 = 64% — typical of academic writing.

Common pitfalls

  • Stopword lists are language-specific; using an English stopword list on French text will inflate density.
  • POS tagging by stopword approximation is imperfect; some content words may be misclassified.
  • Density alone is not a quality signal — dense writing can be unreadable; light writing can be brilliant.
  • Very short samples (under 50 words) produce unstable density figures.

Content words against function words

Lexical density is the share of words carrying meaning rather than grammar. Every word is checked against a stopword list — the articles, prepositions, pronouns, conjunctions and auxiliaries that hold sentences together — and anything not on that list counts as a content word. Density is content words divided by total words, as a percentage.

The number is a register measure more than a quality measure. Speech runs low, typically 40% or below, because conversation is full of function words and repetition. Written prose sits higher. Dense technical and academic writing, and anything compressed like a headline or an abstract, runs higher still.

Neither end is good or bad on its own. High density means information packed tightly, which is efficient for a motivated expert reader and exhausting for anyone else. Low density means the text is doing more work to guide the reader through, which is what you want in explanatory writing and what looks like padding in a specification.

What the stopword list decides for you

The whole measurement rests on one editorial judgement: which words count as grammar. This tool uses a general English stopword list, and that list is a convention rather than a fact — different lists disagree about borderline cases such as "very", "just", "many" and modal verbs like "should".

The practical effect is that absolute density figures are not comparable across tools that use different lists. Comparisons within this tool are consistent, because the same list is applied to both texts, which is why the useful workflow is comparing two drafts rather than benchmarking one against a published number.

Read more about this

Related tools

Send feedback

We read every message. Tell us what could be better or what you love.