Skip to content
Allin

What Readability Scores Actually Measure (and the Three Things They Cannot See)

Published 6/30/2026 · 11 min read · Text & language tools

Daniel Okonkwo

Daniel OkonkwoFront-end developer and tech writer at Allin

Web performance · File formats

Checked against 4 sources

View profile
In short

Flesch Reading Ease and Flesch-Kincaid Grade Level take exactly two inputs: average words per sentence and average syllables per word. FRE = 206.835 - 1.015 x (words/sentences) - 84.6 x (syllables/words). FKGL = 0.39 x (words/sentences) + 11.8 x (syllables/words) - 15.59. Score a 32-word, 3-sentence, 47-syllable passage and you get FRE 71.8 and grade 5.9. Because those two counts are the whole model, the score is blind to three things that decide whether a reader understands you. It cannot see vocabulary difficulty: the sentence about a bank marking a swap to fair worth at each close is twelve one-syllable words, so it scores 110.1 and grade 0.9, the same as a sentence about a boy walking a dog. It cannot see whether the sentences are in a sensible order, because shuffling them changes no count. And it cannot see whether the text is true. Worse, the model is trivially gamed: split one 20-word sentence at its comma and the grade level drops from 11.68 to 7.78, nearly four school years, without a single word changing. Treat the score as a length gauge, not a comprehension measure, and never as a target.

Flesch Reading Ease and Flesch-Kincaid Grade Level count syllables and sentence length. Nothing else. Both formulas in full, one passage scored end to end, and the comma trick that buys 3.9 grade levels without changing a word.

Both formulas, in full, with nothing hidden

Rudolf Flesch published Reading Ease in 1948; a US Navy team led by J. Peter Kincaid re-fitted the same two inputs to a school-grade scale in 1975. Both use only two ratios. ASL is the average sentence length in words: total words divided by total sentences. ASW is the average word length in syllables: total syllables divided by total words. Then: Reading Ease = 206.835 - 1.015 x ASL - 84.6 x ASW, and Grade Level = 0.39 x ASL + 11.8 x ASW - 15.59.

Notice what is absent. There is no word list, no frequency table, no grammar, no discourse model, no fact check. A word is hard if it has many syllables and easy if it does not. That is the entire theory of difficulty inside both formulas, and everything that follows in this article is a consequence of it.

One passage, scored end to end

Take three sentences of ordinary workplace prose: Our new policy takes effect in March. Each employee must complete the online training before the deadline. Managers will receive a report that lists the names of staff who have not finished. That is 32 words in 3 sentences, and 47 syllables counted by hand. So ASL = 32 / 3 = 10.6667 and ASW = 47 / 32 = 1.46875.

Reading Ease: 206.835 - 1.015 x 10.6667 - 84.6 x 1.46875 = 206.835 - 10.8267 - 124.2563 = 71.75, which rounds to 71.8. Grade Level: 0.39 x 10.6667 + 11.8 x 1.46875 - 15.59 = 4.16 + 17.3313 - 15.59 = 5.90. Two numbers, one text. They do not contradict each other; they are two linear rescalings of the same pair of ratios, which is why they always move together.

A score of 71.8 lands in the band Flesch labelled fairly easy, and a grade of 5.9 says a competent eleven-year-old could decode it. Whether an eleven-year-old would know what a deadline for mandatory compliance training is, the formula has no opinion about.

Blind spot one: vocabulary difficulty

Here are two English sentences of twelve words each, every word one syllable, one sentence apiece. The boy fed the dog and then took it for a walk. The bank must mark the swap to fair worth at each close. Both have ASL 12.00 and ASW 1.000, so both score exactly 206.835 - 12.18 - 84.6 = 110.06 on Reading Ease and 4.68 + 11.8 - 15.59 = 0.89 on Grade Level.

One is a story about a dog. The other asks the reader to know what a swap is, what fair value measurement means, and why it has to be redone at every market close. The formulas cannot tell them apart, because syllable count is their only proxy for word difficulty and both sentences have the same one. Any technical field that has short jargon terms - bond, yield, swap, kernel, heap, lien, tort, writ, node - is systematically flattered by these scores.

Note also that 110.06 is above 100. Reading Ease has no ceiling and no floor: it is an unbounded linear function, and a legal sentence of 60 words averaging two syllables each will produce a negative number without anything having gone wrong.

Blind spots two and three: order, and truth

Take the 32-word passage and reverse the order of its three sentences. Word count unchanged, sentence count unchanged, syllable count unchanged, so Reading Ease is still 71.8 and the grade is still 5.90. The text now announces the consequence before the rule and mentions a deadline before saying what it is a deadline for, which is a real comprehension cost that the score is structurally incapable of registering.

The same holds for factual accuracy, and it is worth stating plainly because readability tools are sold into content workflows where the two get conflated. Replace March with a month that is wrong and the score does not move. Replace the whole passage with a confident, well-punctuated, entirely false description of a policy that does not exist and the score will be excellent. Readability is a property of the surface; correctness is a property of the claims. No formula that only counts syllables can be a proxy for the second.

The comma trick: 3.90 grade levels, zero words changed

This is the manipulation everyone does, usually without knowing it is a manipulation. Take a 20-word sentence with 33 syllables: The board approved the revised budget on Tuesday, the finance committee had explained why the original estimate was too low. As one sentence, ASL is 20.00 and ASW is 1.65, so Reading Ease is 206.835 - 20.30 - 139.59 = 46.95 and the grade is 7.80 + 19.47 - 15.59 = 11.68.

Now replace the comma with a full stop. Nothing else changes: same 20 words, same 33 syllables, same order, same meaning. But ASL halves to 10.00, so Reading Ease becomes 206.835 - 10.15 - 139.59 = 57.10 and the grade becomes 3.90 + 19.47 - 15.59 = 7.78. The grade level fell by 3.90 - close to four school years - and the Reading Ease rose by 10.15 points, purely because ASL is multiplied by -1.015 in one formula and by 0.39 in the other and you halved it.

Sometimes splitting genuinely helps. Sometimes it removes the connective that told the reader why the second clause follows the first, and the text gets harder while the score gets better. The formula cannot distinguish the two cases, so if you are editing to a score target you will make both changes indiscriminately and only one of them will be an improvement.

These formulas were calibrated on English, and only English

The constants 206.835, 1.015, 84.6, 0.39, 11.8 and 15.59 are not universal truths. They were fitted to English reading-comprehension data - Flesch against adult reading tests, Kincaid's team against US Navy enlisted personnel. Running them on French, German, Spanish, Portuguese or Italian produces a number, because arithmetic always produces a number, but the number has no established relationship to how hard the text is for a reader of that language.

The failure is not subtle. German compounds a noun phrase into one long word, which inflates syllables per word without making the text harder for a German reader - Geschwindigkeitsbegrenzung is six syllables and perfectly ordinary. Italian and Spanish have systematically higher syllable counts per word than English for reasons of phonology, not difficulty. French elides and liaises in ways that make syllable counting itself contested. Applying the English coefficients to any of them mainly measures the language, not the text.

Use a locally calibrated index instead. For French, Kandel and Moles re-estimated Flesch's coefficients on French texts in the 1950s; use their version, not Flesch's raw constants. For German, the Wiener Sachtextformel was developed and validated on German school and factual texts. Italian has the Gulpease index, built on letters rather than syllables precisely because syllabification is unreliable to automate. Spanish has the Fernandez Huerta adaptation and the later INFLESZ scale. And one index is genuinely language-agnostic in construction: Bjornsson's LIX is simply words per sentence plus the percentage of words longer than six characters. On the 32-word passage above that is 32/3 + 7 x 100/32 = 10.67 + 21.88 = 32.54, since seven of its words exceed six characters. LIX counts characters, not syllables, so it transfers across languages far better than Flesch does - though its bands were still set on Swedish material and should be read as rough.

How to use the number without being used by it

Treat the score as a smoke alarm, not a thermostat. A grade level that jumps from 9 to 16 between two drafts is worth investigating; a grade level of 8.4 versus 8.9 is noise. Read the two underlying ratios rather than the composite: if ASL is 28, you have a sentence-length problem, and if ASW is 1.9, you have a word-length problem, and the composite score tells you neither.

And never set a score as an acceptance criterion for a writer. The moment a number becomes a target, the cheapest way to hit it is the comma trick, and you will get text that is chopped into fragments, stripped of the connectives that carry the argument, and measurably easier by a formula that cannot read.

Words per sentence
Five texts, both formulas, all arithmetic shown. Rows 2 and 3 are the same 20 words with different punctuation; rows 4 and 5 have identical counts and wildly different difficulty.
TextWords per sentenceSyllables per wordReading EaseGrade level
Policy passage: 32 words, 3 sentences, 47 syllables10.671.46971.85.90
Budget sentence as one sentence: 20 words, 33 syllables20.001.65046.9511.68
Exactly the same 20 words, split at the comma into two sentences10.001.65057.107.78
The boy fed the dog and then took it for a walk.12.001.000110.060.89
The bank must mark the swap to fair worth at each close.12.001.000110.060.89
Readability scoreMeasure how easy your text is to read (Flesch score).Try the tool

Frequently asked questions

What counts as a good Flesch Reading Ease score?
Flesch's own bands put 90-100 at very easy, 60-70 at standard and 0-30 at very difficult, so general-audience prose is usually aimed at 60-70. But good depends entirely on the audience: a maintenance manual for trained technicians at 45 may be exactly right, and a consent form at 75 may still be unreadable because of the concepts, not the syllables. There is no threshold that is correct independent of who is reading.
Why do two tools give different scores for the same text?
The formulas are fixed, but the counting is not. Syllable detection is a heuristic - no tool has a full pronunciation dictionary, so words like fire, business and every are counted differently by different implementations. Sentence detection breaks on abbreviations, decimals and ellipses. And tools disagree on whether headings, list items, code and captions are sentences at all. Two implementations can easily differ by a full grade level on the same text; the direction of change between drafts is reliable, the absolute value is not.
Is readability a Google ranking factor?
Google has never published a readability score as a ranking signal, and its guidance on helpful content talks about expertise, originality and satisfying the reader, not about syllable counts. What is true is that unreadable pages tend to lose readers, and reader behaviour is downstream of everything. Optimising a Flesch number is not SEO; writing something a reader finishes is.
Can I run Flesch-Kincaid on French, German or Italian text?
You can compute it, but the result is not valid. The coefficients were estimated on English comprehension data and there is no published mapping from an English-fitted Flesch-Kincaid grade to a French, German or Italian reading level. Use Kandel-Moles for French, the Wiener Sachtextformel for German, Gulpease for Italian, Fernandez Huerta or INFLESZ for Spanish. If you need one number across all your languages, LIX is the least indefensible choice because it counts characters rather than syllables - just do not compare a LIX value to a Flesch value.
Does splitting long sentences really make writing clearer?
Often, yes - a 40-word sentence with three subordinate clauses usually does become clearer as two or three sentences. But the formula rewards the split whether or not it helped, and it rewards it just as much when the split deletes a because or a although that was carrying the logic. The reliable test is not the score: reread the split version and check that a reader can still tell why the second sentence follows the first. If they cannot, you traded comprehension for a lower number.

Articles you may find interesting

All guides
ExplainerWord Frequency and Zipf's Law: We Counted Six Books in Six Languages and Fitted the SlopeThe nth most common word appears about 1/n as often as the first. We counted six public-domain books, printed rank × frequency, fitted log frequency against log rank, and got slopes between -1.02 and -1.08 in all six languages — plus the two places the law breaks.ExplainerWhere a Line May Break: The Unicode Algorithm Behind Every Wrapped Paragraph"Break at spaces" fails in most of the world's writing systems. UAX #14 gives every character a line-break class; we looked ours up in Unicode 17.0.0 and ran a conforming implementation over no-break spaces, soft hyphens, zero-width spaces, URLs, Japanese and Thai.GuideCharacter Limits That Actually Bite: Code Units, Code Points and GraphemesA character is three different things at once. One emoji with a skin tone is 1 grapheme, 2 code points and 4 UTF-16 units. Every count in this guide was measured in Node, plus why an SMS drops from 160 to 70 and why VARCHAR(255) is not 255 of anything in particular.ExplainerKeyword Density Is a Dead Metric, and What Replaced ItDensity counted occurrences because retrieval once counted occurrences. TF-IDF, then BM25 with its saturation curve, then embeddings replaced it. Here is the same 800-word page scored three ways, and why the three disagree.ExplainerCounting Characters Against a Limit Someone Else SetOne emoji is 1 character, or 7, or 11, or 25, depending on who is counting. Which one your form, your database and your SMS gateway mean — and a one-paste test that tells you which you are facing.ExplainerCounting Words Is Ambiguous, and Every Tool Answers DifferentlyA word count is a definition, not a measurement. We counted the same paragraph four ways and got 25, 28, 33 and 38 — then counted 50,000 characters of ordinary prose and got agreement to within 4.5%. The gap is entirely driven by compounds, figures and URLs.

Related tools

Sources

Spotted a mistake in this article?