Skip to content
Allin

A Thousand French Words Covers 80% — and Tells You Nothing

Published 3/10/2026 · 17 min read · Everyday calculators

Lena Hoffmann

Lena HoffmannScience & education writer at Allin

Mathematics · Physics

Checked against 6 sources

View profile
In short

On 312,994,233 words of French subtitles, the 100 most frequent forms cover 57.68% of running text, the top 1,000 cover 80.77% and the top 2,000 cover 85.90%. Those numbers are real and they are also the most over-sold statistic in language learning, for four reasons. First, the denominator is a 50,000-form slice, not the language: measured against the whole corpus — 834,768 distinct forms, 318,214,450 words — the top 1,000 cover 79.45%, and reaching 95% takes 13,909 forms rather than the 9,556 the truncated slice suggests. Second, this is film and television dialogue: je ranks 2nd, vous 8th, tu 9th, and first- and second-person pronouns alone are 9.09% of every word spoken, against 2.10% for il, elle, ils and elles combined. Third, forms are not dictionary entries — the top 1,000 forms are 912 lemmas, a collapse of 8.8%, which is smaller than most people expect. Fourth, and this is the point: the 80% you would know is almost entirely grammar. The first 100 forms alone supply 71.4% of that 80%, and the highest-ranked ordinary noun in them is chose, "thing", at rank 92. Take the sentence "Il a dit qu'on ne pouvait pas savoir si le prévenu avait été mis en examen pour recel aggravé." Sixteen of its twenty tokens are in the top 1,000 — exactly 80% coverage — and the four that are not are prévenu, examen, recel and aggravé. A reader with a perfect command of the top 1,000 reads "He said that one could not know whether the ___ had been placed under ___ for ___ ___" and has learned nothing at all.

We measured the coverage curve of French on 313 million words of subtitles: the top 1,000 forms account for 80.77% of running text. Then we took a normal French sentence in which 80% of the tokens are inside that top 1,000, and showed that a reader who knows every one of them learns nothing from it.

What the 80 percent number actually measures

The measurement behind the claim is simple and it is honest. Take a large body of French, count every word occurrence in it, sort the distinct forms by how often they appear, and then ask: if I walked through the text knowing only the N most frequent forms, what fraction of the words I met would I recognise? That fraction is called coverage. Our figures come from 312,994,233 words of French subtitles, in which 50,000 distinct forms were counted. The 10 most frequent forms alone cover 18.54% of everything said. The first 100 cover 57.68%. The first 1,000 cover 80.77%, and the first 2,000 cover 85.90%.

It is worth seeing the head of that list before drawing any conclusion from it. In order, the twenty most frequent forms of French in this corpus are: de, je, est, pas, le, que, la, vous, tu, un, c', à, et, il, a, l', ne, les, j', en. There is not one noun among them, not one adjective, and not one verb that names an action rather than an auxiliary. They are the connective tissue of the language — prepositions, pronouns, articles, negation, the copula. That is exactly why they are frequent: French needs them in almost every sentence, whatever the sentence is about.

The curve those forms trace is very steep at the start and then almost flat, which is the general shape of word frequency in every language — the site's article on Zipf's law fits the slope on six books in six languages and will not be re-derived here. What matters for a learner is the practical consequence of that shape. Getting from nothing to 80% coverage takes about a thousand forms. Getting from 80% to 90% takes about four thousand more. Getting from 90% to 95% takes nine thousand more on top of that. Every additional point of coverage is bought at a steeply rising price, and the article's four caveats are all about why those last points matter far more than the first eighty.

Caveat one: the denominator is a 50,000-form slice, not French

Every figure in the first column of the table is a percentage of a list that stops at 50,000 forms. Real French does not stop there. The same publisher also releases the untruncated list for this corpus, and it contains 834,768 distinct forms across 318,214,450 words — of which the 50,000-form slice accounts for 98.360%. So the slice is not a small sample, but it is not the whole thing either, and every coverage figure computed inside it is inflated by a factor of about 1.017. Against the real denominator, the top 1,000 forms cover 79.45%, not 80.77%, and the top 2,000 cover 84.49%, not 85.90%. Both columns are in the table so you can see the size of the correction rather than take it on trust.

One and a half points sounds like a rounding error. It is not, because of where on the curve the error lands. Expressed the other way round — how many forms do you need in order to reach a given coverage — the truncated slice says 9,556 forms reach 95% and 20,984 reach 98%. The whole corpus says 13,909 forms reach 95% and 40,652 reach 98%. Where the curve is flat, a small error in coverage becomes a very large error in vocabulary size: at the 98% mark, the truncated denominator understates the requirement by a factor of almost two. Anyone quoting a vocabulary target from a truncated frequency list is quoting a number that is wrong in the direction that flatters them.

The tail is also not empty of anything you would want. Of the 834,768 forms in the full list, 389,191 — nearly half the vocabulary of the corpus — occur exactly once in 318 million words. Those are proper names, misspellings, foreign words and genuine rarities, and they are individually negligible and collectively 0.12% of the text. But the band between rank 20,000 and rank 50,000 is not made of noise: it contains ordinary, boring, entirely learnable French that a subtitle writer used a few hundred times and that a coverage table rounds to nothing.

Caveat two: this is film dialogue, and the ranking shows it

The corpus is OpenSubtitles: what people say to each other on screen. The register is visible in the ranking without any statistical apparatus. je is the second most frequent form in the language according to this list; vous is eighth and tu is ninth. Counting the token mass, first- and second-person personal pronouns — je, j', me, m', moi, nous, tu, t', te, toi, vous — are 9.09% of every word in the corpus, while the third-person subject pronouns il, elle, ils and elles together are 2.10%. In narrative prose or journalism that ratio is inverted. Twenty-three of the hundred most frequent forms are first- or second-person markers, and the list also contains non at rank 42, oui at 54, quoi at 61, pourquoi at 75, bon at 93 and merci at 95.

There is a second, sharper fingerprint. Among the top 2,000 forms, 67 are hyphenated clusters of a verb and its pronoun: est-ce, avez-vous, a-t-il, es-tu, dis-moi, laissez-moi, allez-y, vas-y, excusez-moi. Those are things people say out loud to a person standing in front of them. A French newspaper contains almost none of them. So the honest description of the first column of the table is not "coverage of French" but "coverage of spoken French as written down by subtitlers", and the coverage curve for a novel or a broadsheet would be flatter and lower at every rank, because written registers spread their tokens over a wider vocabulary. We cannot say by how much: that would take a second corpus, and this article measures only the one it has.

Caveat three: forms are not dictionary entries — but the gap is only nine percent

A frequency list counts surface forms. est, a, été, sont and ont are five entries in it and two verbs in a dictionary. So "1,000 words" is not what the list gives you — it gives you 1,000 forms, which is fewer words. We measured the collapse against this site's own French lexicon, 46,159 entries imported from Wiktionnaire, by resolving each form to its lemma where the entry declares itself an inflection of another. The top 100 forms are 91 distinct lemmas, of which 21 are inflected forms. The top 500 are 453 lemmas, with 120 inflections. The top 1,000 are 912 lemmas, with 239 inflections. The top 2,000 are 1,816 lemmas, with 480 inflections. The collapse is 9.0%, 9.4%, 8.8% and 9.2% respectively.

The striking thing is how small that is. The correction people expect here is dramatic — halve the number, they say, because French verbs conjugate. It is nine percent, and it is stable across the whole range. The reason is that the head of the list is mostly invariable: prepositions, articles, pronouns and adverbs have no paradigm to collapse, and the verbs that do appear near the top are so few that even all their forms together barely dent the count. A separate question is whether the forms are in a dictionary at all: 94.0% of the top 2,000 have an entry in the site's lexicon. The remaining 6% is dominated by strings no dictionary has a headword for — the 14 elided clitics c', j', l', d', qu', n', s', t', m' and their kin, the 67 hyphenated verb-plus-pronoun clusters, and two abbreviations. Elisions alone are 14 forms, so they cannot be the whole of a gap of roughly 120; the honest statement is that most of the gap is structural, and the rest is the ordinary incompleteness of any dictionary.

Caveat four: the 80 percent you know carries almost none of the meaning

The first three caveats shave the number. This one dismantles it. Coverage weights every token equally: le counts as much as prévenu. But the two do not carry the same amount of information, and the ranking is built so that the tokens carrying least information sit at the very top. Of the 80.77% that the first thousand forms cover, the first hundred forms supply 71.4% — the whole 900-form stretch from rank 101 to rank 1,000 adds only 23.09 points to the total. And what is in that first hundred? Prepositions, articles, pronouns, negation, the auxiliaries, a handful of high-frequency verbs like faire, dire and savoir, and the conversational fillers oui, non, bon and merci. The highest-ranked ordinary noun in the entire top hundred is chose — "thing" — at rank 92. A vocabulary that names nothing cannot tell you what a sentence is about.

Here is the demonstration, on one sentence of perfectly ordinary French: « Il a dit qu'on ne pouvait pas savoir si le prévenu avait été mis en examen pour recel aggravé. » It has twenty tokens. Sixteen of them are inside the top 1,000 — il at rank 14, a at 15, dit at 71, qu' at 26, on at 21, ne at 17, pouvait at 697, pas at 4, savoir at 200, si at 38, le at 5, avait at 125, été at 99, mis at 397, en at 20, pour at 27. That is exactly 80% token coverage, the headline figure, achieved on a real sentence. The four tokens outside are prévenu (rank 2,710), examen (2,175), recel (26,425) and aggravé (20,608).

Now read what the reader who has mastered the top 1,000 perfectly, and nothing beyond it, actually sees: « Il a dit qu'on ne pouvait pas savoir si le ___ avait été mis en ___ pour ___ ___. » He knows there is a man, that someone said something about him, that it was not knowable, and that the tense is past. He does not know who the man is, what happened to him, or what he is alleged to have done — which is the entire content of the sentence. Worse, one of the four unknowns is a trap for the half-knowing: examen looks like a school examination, and mis en examen is a fixed legal expression meaning placed under formal investigation. Understanding the parts here produces a confident wrong reading rather than a blank. That is the gap between 80% coverage and comprehension, and it does not close gradually — it closes when the content words arrive, which is thousands of ranks further down the list.

The number worth aiming at, and what it costs in study hours

The vocabulary-size literature does not use 80% as a threshold for anything, because nobody reads at 80%. Nation's 2006 study put the coverage needed for unassisted comprehension at 98%, and estimated the vocabulary required to reach it in English at 8,000 to 9,000 word families for written text and 6,000 to 7,000 for spoken. Laufer and Ravenhorst-Kalovski found the same optimal threshold of 98%, at around 8,000 families, with a minimal threshold of 95% at 4,000 to 5,000 families below which comprehension degrades sharply. Cobb's work on the same question is where to look if you want the argument about whether reading alone can get you there. Their numbers are English word families and ours are French surface forms, so the two vocabulary sizes are not comparable — but the coverage thresholds are, and they say the same thing: 80% is not a milestone, it is the starting line.

Applied to this corpus, those thresholds have prices. Reaching 95% of spoken French as subtitled takes about 13,909 forms. Reaching 98% takes about 40,652. That is not a reason to despair, it is a reason to stop treating the first thousand as an achievement and start treating them as free. They are free: they arrive on their own within weeks, because you cannot read a page without meeting de, je, est and pas several times. What costs time is everything after them, and what makes it cost less is choosing the register you want. If you intend to read newspapers, the frequency list of subtitles will not hand you the vocabulary of newspapers, however deep you go into it.

For a time budget rather than a word budget, the calculator on this page is the honest instrument. It scales the Foreign Service Institute's difficulty categories, which put French in the easiest group for an English speaker, and asks for your actual pace. At its default assumptions — French, target level B2, first foreign language, one hour a day five days a week — it returns 600 study hours, which at that pace is 120 weeks, or a little under twenty-eight months. Change the pace to two hours a day six days a week and the same 600 hours land in 50 weeks. The vocabulary curve tells you what you are buying with those hours; the calculator tells you how many of them the plan you actually keep to would take.

Share of the 50,000-form slice
Cumulative share of running French text covered by the most frequent N forms — measured twice, against the 50,000-form slice and against the whole 2018 subtitle corpus
Most frequent N formsShare of the 50,000-form sliceShare of the whole corpus
1018.54%18.24%
5046.65%45.88%
10057.68%56.74%
25068.42%67.30%
50075.05%73.81%
1,00080.77%79.45%
2,00085.90%84.49%
3,00088.57%87.12%
5,00091.64%90.14%
10,00095.21%93.64%
20,00097.85%96.25%
50,000100.00% (by construction)98.36%

Worked with our own calculator

Language learning hours to fluency calculator

Given

Target language
Spanish (cat. I)
Target level (CEFR)
A1
Study hours per day
0.5
Study days per week
3
Prior experience
First foreign language

Result

Total study hours
100
Weeks at your pace
66.667
Months at your pace
15.343

These figures are produced by the calculator below, not typed in by hand — they are recomputed whenever the tool changes.

Run it on your own figures

On this site

Frequently asked questions

So is "1,000 words is 80% of French" simply false?
It is arithmetically defensible and rhetorically dishonest, which is a worse combination than being wrong. The arithmetic: on this corpus the top 1,000 forms do account for 80.77% of running tokens inside the 50,000-form slice, and 79.45% of the whole corpus. The dishonesty is in what the sentence invites you to conclude. "80% of French" sounds like "you will understand four sentences in five", and it means "four words in five will be familiar", which are wildly different claims because the fifth word is where the meaning lives. A more accurate version of the same true fact would be: the first thousand forms of French are the grammar, they arrive quickly, and once you have them the language has not started yet.
How many French words do I actually need, then?
Aim at a coverage figure, not a word count, and then convert. The research consensus is that unassisted comprehension needs about 98% coverage, and that 95% is the floor below which reading becomes decoding. On this corpus of spoken French, 95% coverage takes about 13,909 forms and 98% takes about 40,652. Those are surface forms in one register, so they are an upper bound on dictionary entries and a lower bound on what a newspaper would demand — nine percent of them collapse onto shared lemmas, and a written corpus would spread the same coverage over more vocabulary. Any answer in the low thousands is answering a different question: how many words to follow a simple conversation about familiar things, which is real and useful and is not the same as understanding French.
Does it really matter that the list counts forms rather than dictionary words?
Less than people assume, at least at the top of the list. We measured it: the top 1,000 forms resolve to 912 distinct lemmas, and the figure is 91 for the top 100, 453 for the top 500 and 1,816 for the top 2,000 — a collapse of between 8.8% and 9.4% throughout. The intuition that French conjugation must halve the count is wrong for this part of the list, because the head is dominated by invariable words. Where it matters is practical rather than statistical: the list will hand you avait and été as separate items, and if you learn them as separate items you have learned two facts instead of one system. The fix costs nothing — look the form up, and the dictionary entry will tell you it is a form of avoir or être rather than a word in its own right.
Would the numbers be different for a newspaper or a novel?
Yes, and in the direction that makes the 80% claim worse rather than better. Written registers use a wider vocabulary to say the same amount, so their frequency curves are flatter: the top thousand forms would cover less, and every threshold would sit further down the list. The evidence for this inside our own data is the register fingerprint — first- and second-person pronouns at 9.09% of all tokens, 67 spoken inversion forms like avez-vous and dis-moi inside the top 2,000, merci at rank 95. None of those belong to newspaper French. What we cannot do is tell you the size of the difference, because that would take a second corpus of written French measured the same way, and this article measures the one corpus it has rather than estimating the other.
Do these numbers apply to English, German or Italian too?
The shape does; the numbers do not. Every language produces the same steep-then-flat frequency curve, which is Zipf's law, and the site's article on it fits that slope across six languages. But the specific figures here — 80.77% at rank 1,000, 13,909 forms for 95%, the 9% form-to-lemma collapse — were measured on French and belong to French. A morphologically richer language spreads its tokens over more forms and would collapse harder onto lemmas; a more analytic one would do the opposite. We have not measured those, so we are not going to quote them. Everything on this page is about how much French you need, which is what you want to know whichever language you are reading this page in.

Articles you may find interesting

All guides
How-toHow Long Does It Take to Read a Book?Estimate reading time from word count and reading speed. Learn the simple formula, the average of about 250 words per minute, and a worked example for a full novel.ExplainerCounting Characters Against a Limit Someone Else SetOne emoji is 1 character, or 7, or 11, or 25, depending on who is counting. Which one your form, your database and your SMS gateway mean — and a one-paste test that tells you which you are facing.ExplainerWhat Readability Scores Actually Measure (and the Three Things They Cannot See)Flesch Reading Ease and Flesch-Kincaid Grade Level count syllables and sentence length. Nothing else. Both formulas in full, one passage scored end to end, and the comma trick that buys 3.9 grade levels without changing a word.How-toHow Long Will It Take to Download a File? Time, Bits vs Bytes, and OverheadEstimate download time from file size and connection speed. Learn the size ÷ speed formula, the crucial bits-versus-bytes conversion (divide by 8), and why real downloads run slower than the math predicts.ExplainerHow Fast Do You Read? Average Reading Speed and How to Measure ItMost adults read prose at about 200-250 words per minute. Learn what a normal reading speed is, how to measure yours in a minute, and why comprehension matters more than raw speed.ExplainerWhat Is a URL Slug? Why Lowercase, Hyphenated, and Good for SEOA URL slug is the readable, hyphenated part of a web address that names a page. Learn why slugs are lowercase, how they help SEO, and how to generate one from a title.

Related tools

This article measures one corpus. The frequency figures come from film and television subtitles collected in 2018, and a different corpus — a newspaper, a novel, a textbook — would move every number in the table. Coverage is a property of a text, not of a person: it counts how many running words you would recognise, not how much you would understand, and the article exists to keep those two apart. Nothing here is a promise about how fast anyone learns, and the study-hour figures are averages from institutional teaching, not a schedule for you.

Sources

Spotted a mistake in this article?