What Readability Scores Actually Measure (and the Three Things They Cannot See)
Published 6/30/2026 · 11 min read · Text & language tools
Daniel Okonkwo — Front-end developer and tech writer at Allin
Web performance · File formats
Checked against 4 sources
Flesch Reading Ease and Flesch-Kincaid Grade Level take exactly two inputs: average words per sentence and average syllables per word. FRE = 206.835 - 1.015 x (words/sentences) - 84.6 x (syllables/words). FKGL = 0.39 x (words/sentences) + 11.8 x (syllables/words) - 15.59. Score a 32-word, 3-sentence, 47-syllable passage and you get FRE 71.8 and grade 5.9. Because those two counts are the whole model, the score is blind to three things that decide whether a reader understands you. It cannot see vocabulary difficulty: the sentence about a bank marking a swap to fair worth at each close is twelve one-syllable words, so it scores 110.1 and grade 0.9, the same as a sentence about a boy walking a dog. It cannot see whether the sentences are in a sensible order, because shuffling them changes no count. And it cannot see whether the text is true. Worse, the model is trivially gamed: split one 20-word sentence at its comma and the grade level drops from 11.68 to 7.78, nearly four school years, without a single word changing. Treat the score as a length gauge, not a comprehension measure, and never as a target.
Flesch Reading Ease and Flesch-Kincaid Grade Level count syllables and sentence length. Nothing else. Both formulas in full, one passage scored end to end, and the comma trick that buys 3.9 grade levels without changing a word.
Both formulas, in full, with nothing hidden
Rudolf Flesch published Reading Ease in 1948; a US Navy team led by J. Peter Kincaid re-fitted the same two inputs to a school-grade scale in 1975. Both use only two ratios. ASL is the average sentence length in words: total words divided by total sentences. ASW is the average word length in syllables: total syllables divided by total words. Then: Reading Ease = 206.835 - 1.015 x ASL - 84.6 x ASW, and Grade Level = 0.39 x ASL + 11.8 x ASW - 15.59.
Notice what is absent. There is no word list, no frequency table, no grammar, no discourse model, no fact check. A word is hard if it has many syllables and easy if it does not. That is the entire theory of difficulty inside both formulas, and everything that follows in this article is a consequence of it.
One passage, scored end to end
Take three sentences of ordinary workplace prose: Our new policy takes effect in March. Each employee must complete the online training before the deadline. Managers will receive a report that lists the names of staff who have not finished. That is 32 words in 3 sentences, and 47 syllables counted by hand. So ASL = 32 / 3 = 10.6667 and ASW = 47 / 32 = 1.46875.
Reading Ease: 206.835 - 1.015 x 10.6667 - 84.6 x 1.46875 = 206.835 - 10.8267 - 124.2563 = 71.75, which rounds to 71.8. Grade Level: 0.39 x 10.6667 + 11.8 x 1.46875 - 15.59 = 4.16 + 17.3313 - 15.59 = 5.90. Two numbers, one text. They do not contradict each other; they are two linear rescalings of the same pair of ratios, which is why they always move together.
A score of 71.8 lands in the band Flesch labelled fairly easy, and a grade of 5.9 says a competent eleven-year-old could decode it. Whether an eleven-year-old would know what a deadline for mandatory compliance training is, the formula has no opinion about.
Blind spot one: vocabulary difficulty
Here are two English sentences of twelve words each, every word one syllable, one sentence apiece. The boy fed the dog and then took it for a walk. The bank must mark the swap to fair worth at each close. Both have ASL 12.00 and ASW 1.000, so both score exactly 206.835 - 12.18 - 84.6 = 110.06 on Reading Ease and 4.68 + 11.8 - 15.59 = 0.89 on Grade Level.
One is a story about a dog. The other asks the reader to know what a swap is, what fair value measurement means, and why it has to be redone at every market close. The formulas cannot tell them apart, because syllable count is their only proxy for word difficulty and both sentences have the same one. Any technical field that has short jargon terms - bond, yield, swap, kernel, heap, lien, tort, writ, node - is systematically flattered by these scores.
Note also that 110.06 is above 100. Reading Ease has no ceiling and no floor: it is an unbounded linear function, and a legal sentence of 60 words averaging two syllables each will produce a negative number without anything having gone wrong.
Blind spots two and three: order, and truth
Take the 32-word passage and reverse the order of its three sentences. Word count unchanged, sentence count unchanged, syllable count unchanged, so Reading Ease is still 71.8 and the grade is still 5.90. The text now announces the consequence before the rule and mentions a deadline before saying what it is a deadline for, which is a real comprehension cost that the score is structurally incapable of registering.
The same holds for factual accuracy, and it is worth stating plainly because readability tools are sold into content workflows where the two get conflated. Replace March with a month that is wrong and the score does not move. Replace the whole passage with a confident, well-punctuated, entirely false description of a policy that does not exist and the score will be excellent. Readability is a property of the surface; correctness is a property of the claims. No formula that only counts syllables can be a proxy for the second.
The comma trick: 3.90 grade levels, zero words changed
This is the manipulation everyone does, usually without knowing it is a manipulation. Take a 20-word sentence with 33 syllables: The board approved the revised budget on Tuesday, the finance committee had explained why the original estimate was too low. As one sentence, ASL is 20.00 and ASW is 1.65, so Reading Ease is 206.835 - 20.30 - 139.59 = 46.95 and the grade is 7.80 + 19.47 - 15.59 = 11.68.
Now replace the comma with a full stop. Nothing else changes: same 20 words, same 33 syllables, same order, same meaning. But ASL halves to 10.00, so Reading Ease becomes 206.835 - 10.15 - 139.59 = 57.10 and the grade becomes 3.90 + 19.47 - 15.59 = 7.78. The grade level fell by 3.90 - close to four school years - and the Reading Ease rose by 10.15 points, purely because ASL is multiplied by -1.015 in one formula and by 0.39 in the other and you halved it.
Sometimes splitting genuinely helps. Sometimes it removes the connective that told the reader why the second clause follows the first, and the text gets harder while the score gets better. The formula cannot distinguish the two cases, so if you are editing to a score target you will make both changes indiscriminately and only one of them will be an improvement.
These formulas were calibrated on English, and only English
The constants 206.835, 1.015, 84.6, 0.39, 11.8 and 15.59 are not universal truths. They were fitted to English reading-comprehension data - Flesch against adult reading tests, Kincaid's team against US Navy enlisted personnel. Running them on French, German, Spanish, Portuguese or Italian produces a number, because arithmetic always produces a number, but the number has no established relationship to how hard the text is for a reader of that language.
The failure is not subtle. German compounds a noun phrase into one long word, which inflates syllables per word without making the text harder for a German reader - Geschwindigkeitsbegrenzung is six syllables and perfectly ordinary. Italian and Spanish have systematically higher syllable counts per word than English for reasons of phonology, not difficulty. French elides and liaises in ways that make syllable counting itself contested. Applying the English coefficients to any of them mainly measures the language, not the text.
Use a locally calibrated index instead. For French, Kandel and Moles re-estimated Flesch's coefficients on French texts in the 1950s; use their version, not Flesch's raw constants. For German, the Wiener Sachtextformel was developed and validated on German school and factual texts. Italian has the Gulpease index, built on letters rather than syllables precisely because syllabification is unreliable to automate. Spanish has the Fernandez Huerta adaptation and the later INFLESZ scale. And one index is genuinely language-agnostic in construction: Bjornsson's LIX is simply words per sentence plus the percentage of words longer than six characters. On the 32-word passage above that is 32/3 + 7 x 100/32 = 10.67 + 21.88 = 32.54, since seven of its words exceed six characters. LIX counts characters, not syllables, so it transfers across languages far better than Flesch does - though its bands were still set on Swedish material and should be read as rough.
How to use the number without being used by it
Treat the score as a smoke alarm, not a thermostat. A grade level that jumps from 9 to 16 between two drafts is worth investigating; a grade level of 8.4 versus 8.9 is noise. Read the two underlying ratios rather than the composite: if ASL is 28, you have a sentence-length problem, and if ASW is 1.9, you have a word-length problem, and the composite score tells you neither.
And never set a score as an acceptance criterion for a writer. The moment a number becomes a target, the cheapest way to hit it is the comma trick, and you will get text that is chopped into fragments, stripped of the connectives that carry the argument, and measurably easier by a formula that cannot read.
| Text | Words per sentence | Syllables per word | Reading Ease | Grade level |
|---|---|---|---|---|
| Policy passage: 32 words, 3 sentences, 47 syllables | 10.67 | 1.469 | 71.8 | 5.90 |
| Budget sentence as one sentence: 20 words, 33 syllables | 20.00 | 1.650 | 46.95 | 11.68 |
| Exactly the same 20 words, split at the comma into two sentences | 10.00 | 1.650 | 57.10 | 7.78 |
| The boy fed the dog and then took it for a walk. | 12.00 | 1.000 | 110.06 | 0.89 |
| The bank must mark the swap to fair worth at each close. | 12.00 | 1.000 | 110.06 | 0.89 |
Frequently asked questions
- What counts as a good Flesch Reading Ease score?
- Flesch's own bands put 90-100 at very easy, 60-70 at standard and 0-30 at very difficult, so general-audience prose is usually aimed at 60-70. But good depends entirely on the audience: a maintenance manual for trained technicians at 45 may be exactly right, and a consent form at 75 may still be unreadable because of the concepts, not the syllables. There is no threshold that is correct independent of who is reading.
- Why do two tools give different scores for the same text?
- The formulas are fixed, but the counting is not. Syllable detection is a heuristic - no tool has a full pronunciation dictionary, so words like fire, business and every are counted differently by different implementations. Sentence detection breaks on abbreviations, decimals and ellipses. And tools disagree on whether headings, list items, code and captions are sentences at all. Two implementations can easily differ by a full grade level on the same text; the direction of change between drafts is reliable, the absolute value is not.
- Is readability a Google ranking factor?
- Google has never published a readability score as a ranking signal, and its guidance on helpful content talks about expertise, originality and satisfying the reader, not about syllable counts. What is true is that unreadable pages tend to lose readers, and reader behaviour is downstream of everything. Optimising a Flesch number is not SEO; writing something a reader finishes is.
- Can I run Flesch-Kincaid on French, German or Italian text?
- You can compute it, but the result is not valid. The coefficients were estimated on English comprehension data and there is no published mapping from an English-fitted Flesch-Kincaid grade to a French, German or Italian reading level. Use Kandel-Moles for French, the Wiener Sachtextformel for German, Gulpease for Italian, Fernandez Huerta or INFLESZ for Spanish. If you need one number across all your languages, LIX is the least indefensible choice because it counts characters rather than syllables - just do not compare a LIX value to a Flesch value.
- Does splitting long sentences really make writing clearer?
- Often, yes - a 40-word sentence with three subordinate clauses usually does become clearer as two or three sentences. But the formula rewards the split whether or not it helped, and it rewards it just as much when the split deletes a because or a although that was carrying the logic. The reliable test is not the score: reread the split version and check that a reader can still tell why the second sentence follows the first. If they cannot, you traded comprehension for a lower number.
Articles you may find interesting
All guides →Related tools
Sources
- American Psychological Association — Flesch, R. (1948). A new readability yardstick. Journal of Applied Psychology, 32(3), 221-233.
- Defense Technical Information Center — Kincaid, Fishburne, Rogers & Chissom (1975). Derivation of New Readability Formulas for Navy Enlisted Personnel. Research Branch Report 8-75.
- Microsoft Support — Get your document's readability statistics (Flesch Reading Ease and Flesch-Kincaid Grade Level in Word)
- Google Search Central — Creating helpful, reliable, people-first content
Spotted a mistake in this article?