Skip to content

Many MUA don't recognize charset "eucgb2312_cn" in email header #88726

Description

@tommylikehu
BPO 44560
Nosy @warsaw, @terryjreedy, @bitdancer, @corona10

Note: these values reflect the state of the issue at the time it was migrated and might not reflect the current state.

Show more details

GitHub fields:

assignee = None
closed_at = None
created_at = <Date 2021-07-04.08:32:11.133>
labels = ['type-bug', 'expert-email', '3.9']
title = 'Many MUA don\'t recognize charset "eucgb2312_cn" in email header'
updated_at = <Date 2021-07-09.19:33:45.681>
user = 'https://bugs.python.org/tommylikehu'

bugs.python.org fields:

activity = <Date 2021-07-09.19:33:45.681>
actor = 'terry.reedy'
assignee = 'none'
closed = False
closed_date = None
closer = None
components = ['email']
creation = <Date 2021-07-04.08:32:11.133>
creator = 'tommylikehu'
dependencies = []
files = []
hgrepos = []
issue_num = 44560
keywords = []
message_count = 3.0
messages = ['396939', '397045', '397210']
nosy_count = 5.0
nosy_names = ['barry', 'terry.reedy', 'r.david.murray', 'corona10', 'tommylikehu']
pr_nums = []
priority = 'normal'
resolution = None
stage = None
status = 'open'
superseder = None
type = 'behavior'
url = 'https://bugs.python.org/issue44560'
versions = ['Python 3.9']

Linked PRs

Activity

  1. tommylikehu commented on Jul 4, 2021

    tommylikehumannequin
    MannequinAuthor

    Email module is used for email message decode and encode, if the header content is gb2312 encoded for example "中文", by design we would finally have a rfc-2047 encoded header as below:

    =?eucgb2312_cn?b?1tDOxA==?=
    

    the test script is as below:

    from email import header, charset
    
    h = header.make_header([(str("中文").encode("gb2312"),
                             charset.Charset("gb2312"))])
    print(h.encode())
    

    My question is why don't we use "gb2312" as the charset in rfc-2047 encoded string, considering the "eucgb2312_cn" is only python awareness.

    Thanks

  2. bitdancer commented on Jul 6, 2021

    @bitdancer
    Member

    I can't tell tell for sure if this behavior is intentional or not from a quick glance at the code (though like you I wouldn't think it would be).

    That's part of the legacy api, at this point. The new api will just use utf8:

    from email.message import EmailMessage
    
    m = EmailMessage()
    m['Subject'] = '中文'
    
    print(bytes(m))

    results in

    b'Subject: =?utf-8?b?5Lit5paH?=\n\n'

    The fix, assuming it is correct, would be to add the line:

    'eucgb2312_cn': 'gb2312',
    

    to the CODEC_MAP in email/charset.py, and then specify the internal codec name in your Charset call. I'm not sure that's right, though...once upon I time I think I understood the logic behind the charset module, but I no longer remember the details.

    I'd recommend just using the new API and not the legacy API.

  3. terryjreedy commented on Jul 9, 2021

    @terryjreedy
    Member

    Anything before 3.9 only gets security patches.

  4. changed the title [-]Unrecognized charset "eucgb2312_cn" in email header for many MUA[/-] [+]Many MUA don't recognize charset "eucgb2312_cn" in email header[/+] on Jul 9, 2021
  5. changed the title [-]Unrecognized charset "eucgb2312_cn" in email header for many MUA[/-] [+]Many MUA don't recognize charset "eucgb2312_cn" in email header[/+] on Jul 9, 2021
  6. transferred this issue fromon Apr 10, 2022
  7. serhiy-storchaka commented on May 17, 2026

    @serhiy-storchaka
    Member

    CODEC_MAP maps gb2312 to eucgb2312_cn and big5 to big5_tw. Python initially did not support Asian codecs by default. They were implemented as third-party codecs, and CODEC_MAP mapped MIME names to names of these codecs. Most of it's content was gone in 4a44293 (see bpo-39645/gh-39645). There was a distinction between the charset name (included in the message) and the codec name, used to convert between Unicode strings and binary data. This distinction mostly gone in the current code, which uses output_charset and output_codec interchangeably.

    The bug is a case when output_codec (or input_codec) was used for encoding and was included in the message. This, the private codec alias was leaked. Since Python supports gb2312 and big5, these mappings are not needed.

  8. added
    stdlibStandard Library Python modules in the Lib/ directory
    3.13only security fixes
    3.14bugs and security fixes
    3.15bugs and security fixes
    3.16new features, bugs and security fixes
    and removed on May 17, 2026
  9. added 4 commits that reference this issue on May 26, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    3.13only security fixes3.14bugs and security fixes3.15bugs and security fixes3.16new features, bugs and security fixesstdlibStandard Library Python modules in the Lib/ directorytopic-emailtype-bugAn unexpected behavior, bug, or error

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions