Why Your Subject Line Arrives as =?UTF-8?B? and a String of Letters
Published 9/8/2026 · 3 min read · File tools
Daniel Okonkwo — Front-end developer and tech writer at OneKitly
Web performance · File formats
Checked against 3 sources
A header may contain only plain ASCII, which leaves no room for é, ü, ñ or an em dash. The way round it is defined in RFC 2047: the text is wrapped as =?charset?encoding?data?=, where the encoding is B for base64 or Q for a quoted-printable variant in which a space becomes an underscore. So Réunion — ordre du jour travels either as =?UTF-8?B?UsOpdW5pb24g4oCUIG9yZHJlIGR1IGpvdXI=?= or as =?UTF-8?Q?R=C3=A9union_=E2=80=94_ordre_du_jour?=, and both decode to exactly the same words. Older messages use =?ISO-8859-1?Q?R=E9union?= and decode just as cleanly. The body is a separate problem with the same shape: it is usually quoted-printable, where =C3=A9 is é and a line ending in a bare equals sign continues on the next one. Convert the message to plain text and both layers come off at once, which is what makes a saved message searchable and quotable rather than something you squint at.
Mail headers are allowed to carry ASCII only, so an accented subject is encoded before it travels. Two encodings are in use, both start with =?, and decoding them is the whole difference between a readable message and a wall of symbols.
Why two encodings, and which you get
Q leaves ordinary letters alone and escapes only what it must, so a subject that is mostly English with one accent stays almost readable in the raw file. B encodes everything, which is shorter when most of the characters are non-ASCII — a subject in Greek or Japanese is far more compact as base64 than as a string of escapes. Mail programs choose per header, which is why one message can carry a Q-encoded subject and a B-encoded sender name, and why both have to be handled.
What plain text is good for
A message reduced to text is searchable by every tool on your machine, quotable without dragging a layout along, and small enough to keep by the thousand. It is also the form that survives: an HTML message from 2011 renders differently in every reader that has existed since, while its text alternative reads today exactly as it read then. Most messages carry both, and the text one is the one nobody looks at until they need it.
| In the raw file | Decoded |
|---|---|
| =?UTF-8?B?UsOpdW5pb24g4oCUIG9yZHJl…?= | Réunion — ordre du jour |
| =?UTF-8?Q?R=C3=A9union_=E2=80=94_ordre…?= | Réunion — ordre du jour |
| =?ISO-8859-1?Q?R=E9union?= | Réunion |
| Réunion (not encoded at all) | Réunion — passed through unchanged |
Frequently asked questions
- My subject shows as Réunion instead of Réunion. What is that?
- The opposite problem: the bytes were decoded, but with the wrong charset. é is what the two bytes of a UTF-8 é look like when read as Latin-1. It usually comes from a program that ignored the charset the header declared. Opening the original .eml and letting the tool read the declaration gives the right characters back — the file was never damaged, only misread.
- Does the text version lose anything?
- The layout, the images and the links' destinations — a link becomes its words, not its address, unless the sender wrote the address out. What it keeps is everything that was said. If you need the addresses too, take the HTML conversion or read the message in the viewer, where the links are still links.
- Can I convert a whole mailbox at once?
- Split the .mbox into individual .eml files first, then convert the ones you want. That two-step route is deliberate: a mailbox is usually far larger than the handful of messages you actually need as text, and splitting first lets you pick.
Articles you may find interesting
All guides →Related tools
Sources
Spotted a mistake in this article?