Repository navigation
Many MUA don't recognize charset "eucgb2312_cn" in email header #88726
Description
Activity
Email module is used for email message decode and encode, if the header content is gb2312 encoded for example "中文", by design we would finally have a rfc-2047 encoded header as below:
=?eucgb2312_cn?b?1tDOxA==?=the test script is as below:
from email import header, charset h = header.make_header([(str("中文").encode("gb2312"), charset.Charset("gb2312"))]) print(h.encode())My question is why don't we use "gb2312" as the charset in rfc-2047 encoded string, considering the "eucgb2312_cn" is only python awareness.
Thanks
- addedtype-bugAn unexpected behavior, bug, or errorAn unexpected behavior, bug, or error
on Jul 4, 2021 - added3.7 (EOL)end of lifeend of life3.8 (EOL)end of lifeend of life3.9 (EOL)end of lifeend of life
on Jul 6, 2021 I can't tell tell for sure if this behavior is intentional or not from a quick glance at the code (though like you I wouldn't think it would be).
That's part of the legacy api, at this point. The new api will just use utf8:
from email.message import EmailMessage m = EmailMessage() m['Subject'] = '中文' print(bytes(m))
results in
b'Subject: =?utf-8?b?5Lit5paH?=\n\n'
The fix, assuming it is correct, would be to add the line:
'eucgb2312_cn': 'gb2312',to the CODEC_MAP in email/charset.py, and then specify the internal codec name in your Charset call. I'm not sure that's right, though...once upon I time I think I understood the logic behind the charset module, but I no longer remember the details.
I'd recommend just using the new API and not the legacy API.
Anything before 3.9 only gets security patches.
- changed the title
[-]Unrecognized charset "eucgb2312_cn" in email header for many MUA[/-][+]Many MUA don't recognize charset "eucgb2312_cn" in email header[/+]on Jul 9, 2021 - changed the title
[-]Unrecognized charset "eucgb2312_cn" in email header for many MUA[/-][+]Many MUA don't recognize charset "eucgb2312_cn" in email header[/+]on Jul 9, 2021 CODEC_MAP maps gb2312 to eucgb2312_cn and big5 to big5_tw. Python initially did not support Asian codecs by default. They were implemented as third-party codecs, and CODEC_MAP mapped MIME names to names of these codecs. Most of it's content was gone in 4a44293 (see bpo-39645/gh-39645). There was a distinction between the charset name (included in the message) and the codec name, used to convert between Unicode strings and binary data. This distinction mostly gone in the current code, which uses
output_charsetandoutput_codecinterchangeably.The bug is a case when
output_codec(orinput_codec) was used for encoding and was included in the message. This, the private codec alias was leaked. Since Python supports gb2312 and big5, these mappings are not needed.- addedstdlibStandard Library Python modules in the Lib/ directoryStandard Library Python modules in the Lib/ directory3.13only security fixesonly security fixes3.14bugs and security fixesbugs and security fixes3.15bugs and security fixesbugs and security fixes3.16new features, bugs and security fixesnew features, bugs and security fixesand removed3.9 (EOL)end of lifeend of life
on May 17, 2026 - added 4 commits that reference this issue
on May 26, 2026
Note: these values reflect the state of the issue at the time it was migrated and might not reflect the current state.
Show more details
GitHub fields:
bugs.python.org fields:
Linked PRs