Skip to content

Double free / use-after-free in list.append() when the list grows under MemoryError (_CALL_LIST_APPEND) #151818

Description

@devdanzin

Crash report

What happened?

When list.append(x) has to grow the list's backing array and that allocation fails (i.e.
under MemoryError), the appended item x is decref'd twice. If x is referenced
elsewhere this is a use-after-free: the interpreter aborts with _Py_NegativeRefcount on a
debug build, or segfaults on a release build, instead of raising a recoverable
MemoryError.

This is reachable under genuine memory pressure (a real RLIMIT_AS reproducer with no test
API is included below), so a program that correctly catches MemoryError can still be left
with a corrupted interpreter.

Reproducer

Deterministic, pure Python, using _testcapi.set_nomemory to fail the grow allocation at a
controlled point:

import _testcapi

class C:
    __slots__ = ("ref",)
    def __init__(self, ref):
        self.ref = ref

def fill():
    items = [C(str(i) + "_unique") for i in range(200)]
    out = []
    for it in items:
        out.append(it.ref)          # CALL_LIST_APPEND; it.ref is also held by the C instance

fill()                              # warm up: specialize out.append(...) to CALL_LIST_APPEND
for start in range(1500):
    _testcapi.set_nomemory(start, start + 1)   # fail one allocation, then resume
    try:
        try:
            fill()
        finally:
            _testcapi.remove_mem_hooks()
    except BaseException:
        pass

On a --with-pydebug build this aborts:

./Include/refcount.h:520: _Py_NegativeRefcount: Assertion failed: object has negative ref count
Fatal Python error: _PyObject_AssertFailed

On a release build it segfaults.

Without any test API (real MemoryError)

The same double-free fires under a genuine allocation failure. With an RLIMIT_AS cap so the
list's grow allocation returns NULL naturally (run on a non-ASan build):

import resource

pool = [object() for _ in range(8_000_000)]      # uniquely-referenced items, built before the cap
warm = []
for i in range(3000):
    warm.append(pool[i])                          # specialize CALL_LIST_APPEND
del warm

cur = int(open("/proc/self/statm").read().split()[0]) * 4096   # current virtual size
resource.setrlimit(resource.RLIMIT_AS, (cur + 24 * 1024 * 1024,) * 2)

out = []
for x in pool:
    out.append(x)                                 # real list_resize failure -> double-free -> SIGSEGV
$ ./python natural.py
Segmentation fault            # faulthandler pins the crash to the `out.append(x)` line

Under the same cap, when the failing allocation is not a list-append grow (e.g. appending
large bytes), Python raises a clean, catchable MemoryError and does not crash — so the
segfault is specific to the buggy append path.

Root cause

In the specialized append bytecode _CALL_LIST_APPEND (Python/bytecodes.c):

op(_CALL_LIST_APPEND, (callable, self, arg -- none, c, s)) {
    ...
    int err = _PyList_AppendTakeRef((PyListObject *)self_o, PyStackRef_AsPyObjectSteal(arg));
    UNLOCK_OBJECT(self_o);
    if (err) {
        ERROR_NO_POP();
    }
    ...
}

arg is stolen via PyStackRef_AsPyObjectSteal and handed to _PyList_AppendTakeRef,
which consumes the reference on every path — including decref'ing the item when the grow
fails (_PyList_AppendTakeRefListResize → if (list_resize(...) < 0) { Py_DECREF(newitem); return -1; }, Objects/listobject.c).

But on that failure the uop takes ERROR_NO_POP(), which jumps to exception handling
without removing arg from the value stack. Since arg's reference was already
consumed, the stale arg stackref is now dangling. The eval loop's exception_unwind then
pops the frame's value stack and PyStackRef_XCLOSEs every slot, closing the stale arg
slot — a second decref of the item.

(Confirmed with ASan on a --with-pymalloc build: the item is freed by PyStackRef_XCLOSE
← _PyEval_EvalFrameDefault (the exception_unwind handler); both the item's allocation and
the second, use-after-free decref are visible in the report.)

Suggested fix

_CALL_LIST_APPEND must account for the consumed arg on the error path — once
_PyList_AppendTakeRef has taken the reference, the arg stackref is dead and must not be
left on the value stack for exception_unwind to close. The sibling ops already show the two
correct idioms:

  • the comprehension element-adds (LIST_APPEND, SET_ADD, MAP_ADD) call the same kind of
    steal/*TakeRef helper but use ERROR_IF(...), so the codegen drops the consumed input on
    the error path. Concretely, [x for x in ...] (LIST_APPEND, the same
    _PyList_AppendTakeRef helper) does not crash where lst.append(x)
    (_CALL_LIST_APPEND) does;
  • the consuming call ops (_DO_CALL_FUNCTION_EX, _PY_FRAME_EX) call INPUTS_DEAD(); SYNC_SP(); before ERROR_NO_POP().

I audited the other specialized ops: _CALL_LIST_APPEND is the only one that steals a
stack input and then takes a bare ERROR_NO_POP() without either form of accounting, which
is why it is the lone op affected.

Environment

  • CPython main (3.16.0a0); reproduced on --with-pydebug builds (abort) and release builds
    (segfault), both free-threaded and default GIL.

This report and the reduced reproducers were drafted with the assistance of Claude Code; I
have reviewed and reproduced them.

CPython versions tested on:

CPython main branch, 3.16

Operating systems tested on:

Linux

Output from running 'python -VV' on the command line:

Python 3.16.0a0 (heads/main:aec0aed1978, Jun 20 2026, 17:44:00) [Clang 22.1.2 (1ubuntu1)]

Linked PRs

Activity

  1. added
    interpreter-core(Objects, Python, Grammar, and Parser dirs)
    type-crashA hard crash of the interpreter, possibly with a core dump
    on Jun 20, 2026
  2. devdanzin commented on Jun 20, 2026

    @devdanzin
    MemberAuthor

    Related to #151119 / #151538, but I believe this is a distinct bug.

    Here's Claude analysis:

    #151119 / #151538 fix the missing eval-stack sync before _PyList_AppendTakeRef — the _Py_Dealloc stackpointer != NULL assert, which fires when the appended item has no other reference and is deallocated immediately.

    This issue is the double-free / use-after-free: _CALL_LIST_APPEND steals arg, _PyList_AppendTakeRef decrefs it when list_resize fails, but the op then takes ERROR_NO_POP() and leaves the consumed arg on the value stack, so exception_unwind closes it a second time. It surfaces when the item has another live reference (so that first decref doesn't free it).

    #151538 marks the append ops HAS_ESCAPES, which also adds the SP-sync to _CALL_LIST_APPEND — but its _CALL_LIST_APPEND hunk only adds the sync; the error path still JUMP_TO_ERROR()s with arg on the stack, so the double-free survives that PR. Notably _CALL_LIST_APPEND is the only append op using ERROR_NO_POP; LIST_APPEND / SET_ADD / MAP_ADD use the codegen-accounted ERROR_IF, which drops the consumed input. So the fix here is to account for the stolen arg on the error path (and could be folded into #151538).

  3. added 3 commits that reference this issue on Jun 21, 2026
  4. vstinner commented on Jun 24, 2026

    @vstinner
    Member

    Oh, I created a duplicate issue: #152090.

  5. vstinner commented on Jun 24, 2026

    @vstinner
    Member

    According to git bisect, the regression was introduced by commit 72f5665:

    commit 72f56654d06a6d23c91e892c05f9e4d70009315b
    Author: Mark Shannon <mark@hotpy.org>
    Date:   Wed Feb 12 17:44:59 2025 +0000
    
        GH-128682: Account for escapes in `DECREF_INPUTS` (GH-129953)
        
        * Handle escapes in DECREF_INPUTS
        
        * Mark a few more functions as escaping
        
        * Replace DECREF_INPUTS with PyStackRef_CLOSE where possible
    
  6. vstinner commented on Jun 24, 2026

    @vstinner
    Member

    A simpler reproducer:

    import _testcapi
    import sys
    import os
    import struct
    
    sizeof_empty_list = sys.getsizeof([])
    sizeof_pointer = struct.calcsize('P')
    
    def capacity(lst):
        size = (sys.getsizeof(lst) - sizeof_empty_list)
        if size % sizeof_pointer:
            raise ValueError("wrong list size")
        capacity = size // sizeof_pointer
        return capacity
    
    def func(offset):
        L=[1,2,3]
        memory_error = False
        small = 1024 * 1024
        for i in range(1000):
            # run a few iterations (25) to specialize the bytecode
            # to CALL_LIST_APPEND
            if i > 25 and capacity(L) == len(L):
                # the list is full: L.append() will have to resize the list
                os.uname()  # used as gdb traceback
                _testcapi.set_nomemory(offset, 0)
            try:
                try:
                    L.append(b'x' * small)
                finally:
                    _testcapi.remove_mem_hooks()
            except MemoryError:
                memory_error = True
                break
    
    for i in range(0, 10):
        func(i)
    print("ok")
  7. added 2 commits that reference this issue on Jun 24, 2026
  8. vstinner commented on Jun 24, 2026

    @vstinner
    Member

    #152114 and #152117 look like workarounds for a bug in Tools/cases_generator/. But well, it's better than no fix :-)

    I wrote #152117 to add a non-regression test.

  9. added 2 commits that reference this issue on Jun 24, 2026
  10. devdanzin commented on Jul 3, 2026

    @devdanzin
    MemberAuthor

    FYI — blast radius / presentations index. fusil's OOM-injection fuzzing keeps surfacing this as superficially-unrelated crashes, but every one rr-reduces to this same _CALL_LIST_APPEND double-free (_PyList_AppendTakeRefListResize, listobject.c:531). Recording them here so the next person who hits one of these signatures finds this issue instead of filing a new one.

    One producer, many victim/detector sites (depending on what the double-freed item was and what reused its freed slot) — on a debug build:

    Presentation (assert / crash) victim
    _Py_NegativeRefcount / <object is freed> any item with a 2nd owner
    DirEntry_dealloc (posixmodule.c) segv / UAF an os.DirEntry.path string
    pycore_stackref.h:726 PyStackRef_XCLOSE negref the leftover stolen stackref itself
    traceback.c:313 _PyTraceBack_FromFrame assert an object held via a traceback frame
    pycore_object.h Py_DECREF_MORTAL — !_Py_IsStaticImmortal(op) freed block read back as static-immortal
    dictobject.c:961 new_dict — mp == NULL || Py_IS_TYPE(mp, &PyDict_Type) a dict that landed on the dict freelist
    tupleobject.c:48 tuple_alloc — PyTuple_Check(op) a tuple on the tuple freelist
    gc.c:380 validate_list (at shutdown) a co_consts constant whose freed slot a weakref's PyGC_Head reused

    And it's reachable from many ordinary stdlib operations under real memory pressure — os.walk / os.scandir, pkgutil.get_importer / the import machinery (importlib's path.append), sys.path mutation + __import__ — not only a literal list.append. (Several of these arrived as separate reports that rr later folded into this one.)

    Since they all funnel through the single _CALL_LIST_APPEND steal-then-ERROR_NO_POP, the fix here should cover every presentation above; flagging mainly because the freelist / GC-head faces don't look like a list.append bug at a glance.

    Full triage, per-face rr ledgers, and reproducers for each presentation live in a sibling catalog of fusil's OOM findings — this bug is reports/OOM-0036-list-append-oom-double-free/ (report.md + meta.json; per-face backtraces backtrace_immortal.txt / backtrace_dict_freelist.txt / backtrace_validate_list.txt; reproducers repro.py, repro_natural.py (real RLIMIT_AS, no test API), repro_get_importer.py).


    Investigation and comment draft by Claude Code (Opus 4.8)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    interpreter-core(Objects, Python, Grammar, and Parser dirs)type-crashA hard crash of the interpreter, possibly with a core dump

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions