Collectible thread-static handle reuse (CoreCLR)
A thread that exits can free an object it never touched, in a collectible context it never
touched. When a thread has used a [ThreadStatic] of a type in collectible context A, context A
is unloaded while the thread lives on, and the thread-static index A's type held is handed to a type
in a newer collectible context B, then the thread's exit calls FreeHandle on B's
LoaderAllocator with a handle that was minted by A. That releases whatever object happens to sit
at that slot of B's handle table — a GC-statics box, a RuntimeType, anything. The next GC collects
it, and every later read of that static goes through a pointer to memory that now holds something
else.
This is a candidate mechanism for sighting #18, the dangling static
base of EqualityComparer<IScheduledObserver<IEnumerable<LineOfBusiness>>>. It is established in three
ways. It is a runtime defect, not this repository's. It reproduces deterministically in the complete
program below, on .NET 10.0.11 and 10.0.12. And it produces exactly the state #18's dump shows: a
box absent from m_slots while the rest of the handle table looks normal. One thing is not
established: that it is what happened in #18. The dump cannot identify the exiting thread or the freed
handle (see Evidence from the #18 dump below), and the production dumps of #4654 have not been read.
Investigated on Systemorph/MeshWeaver#4654.
The defect, in the runtime source
All references are dotnet/runtime release/10.0. main has the same code as of 2026-09-26.
- A thread's first touch of a collectible thread static allocates the thread's data block and
pins it with a handle in the type's
LoaderAllocator. The thread records that handle, an index into the allocator's managedm_slotsarray, in its ownpLoaderHandles[tlsIndex](threadstatics.cpp,GetThreadLocalStaticBase:*pLoaderHandle = pMT->GetLoaderAllocator()->AllocateHandle(gc.tlsEntry)). - When that allocator is destroyed,
FreeTLSIndicesForLoaderAllocatormarks its TLS indices cleared in the global map (TLSIndexToMethodTableMap::Clear). It does not touch any thread'spLoaderHandles, so the stale handle stays in every live thread that used the type.pLoaderHandlesappears in exactly two files,threads.handthreadstatics.cpp, and no other code scrubs it. - The next collectible type that needs a thread-static index reuses the cleared one
(
GetTLSIndexForThreadStatic→FindClearedIndex), and that type lives in a different allocator. - When the thread exits,
FreeLoaderAllocatorHandlesForTLSDatawalks the current map and, for every index where the thread'spLoaderHandlesentry is non-null, callsentry.pMT->GetLoaderAllocator()->FreeHandle(handle). That is the new type's allocator, called with the old allocator's handle.FreeHandlewritesnullinto that slot of the new allocator'sm_slotsand pushes the index onto the free stack. - The victim is whatever lived at that slot. For a GC-statics box of a collectible type
(
LoaderAllocator::AllocateGCHandlesBytesForStaticVariables), them_slotsentry is the only strong root.m_pGCStaticsis tracked by a weak interior handle, which is nulled when the box dies and leavesm_pGCStaticsuntouched. The box is collected, andm_pGCStaticskeeps its old address. - The next
AllocateHandlepops the freed index and fills the slot again, so the handle table shows no gap afterwards. Sighting #18 looked like that:m_slotshad no null entries inside the used range and held no box for V.
A thread that touches the new type first overwrites its stale entry (the old handle is leaked, which is harmless). The corruption needs the thread to exit without touching the new type.
The repro — deterministic, with a control
The full source is below: a standalone program, outside the repository's build and outside its
concurrency rules. It deliberately corrupts its own heap, so it must never run inside a test host.
There are two projects, lib and host. The library is generated, because it needs 400 static
classes, and it is loaded twice, into two collectible contexts.
lib/lib.csproj
<Project Sdk="Microsoft.NET.Sdk">
<PropertyGroup><TargetFramework>net10.0</TargetFramework><Nullable>disable</Nullable></PropertyGroup>
</Project>
lib/Lib.cs, generated by this script (python3 gen.py > lib/Lib.cs):
N = 400
print("namespace Lib;")
print("public sealed class Marker { public long Tag = 0x5EED; }")
print("public static class Tls { [System.ThreadStatic] static object t; "
"public static void Touch() => t = new object[] { new Marker() }; }")
for i in range(N):
print(f"public static class H{i} {{ public static object V = new Marker(); }}")
touch = " ".join(f"_ = H{i}.V;" for i in range(N))
check = " ".join(f"if (!(H{i}.V is Marker m{i} && m{i}.Tag == 0x5EED)) bad++;" for i in range(N))
print("public static class Probe { public static void TouchAll() { " + touch + " } "
"public static int CountBad() { int bad = 0; " + check + " return bad; } }")
host/host.csproj
<Project Sdk="Microsoft.NET.Sdk">
<PropertyGroup><OutputType>Exe</OutputType><TargetFramework>net10.0</TargetFramework><Nullable>disable</Nullable>
<ImplicitUsings>enable</ImplicitUsings><ServerGarbageCollection>false</ServerGarbageCollection></PropertyGroup>
</Project>
host/Program.cs
using System.Reflection;
using System.Runtime.CompilerServices;
using System.Runtime.Loader;
var libPath = args[0];
var mode = args.Length > 1 ? args[1] : "repro"; // repro | control | baseline
var touched = new ManualResetEventSlim();
var release = new ManualResetEventSlim();
Thread holder = null;
if (mode == "baseline") // no collectible code at all: the emulator control
{
for (int i = 0; i < 5; i++) { GC.Collect(); GC.WaitForPendingFinalizers(); }
var j = new object[200_000]; for (int i = 0; i < j.Length; i++) j[i] = new byte[48];
GC.Collect();
Console.WriteLine("baseline done");
return;
}
var old = StartOld();
for (int i = 0; old.IsAlive && i < 200; i++) { GC.Collect(); GC.WaitForPendingFinalizers(); }
Console.WriteLine($"old context collected: {!old.IsAlive}");
var b = Load(libPath, out var ctx2);
Call(b, "Lib.Probe", "TouchAll"); // 400 GC-statics boxes -> slots in ctx2's LoaderAllocator
Call(b, "Lib.Tls", "Touch"); // ctx2's Tls reuses the cleared TLS index
var countBad = b.GetType("Lib.Probe")!.GetMethod("CountBad")!;
Console.WriteLine($"bad before holder exit: {countBad.Invoke(null, null)}");
if (mode == "repro") { release.Set(); holder.Join(); Console.WriteLine("holder thread exited"); }
else Console.WriteLine("control: holder thread kept alive");
for (int i = 0; i < 5; i++) { GC.Collect(); GC.WaitForPendingFinalizers(); }
var junk = new object[200_000]; for (int i = 0; i < junk.Length; i++) junk[i] = new byte[48];
GC.Collect();
Console.WriteLine($"bad after holder exit + GC: {countBad.Invoke(null, null)}");
for (int i = 0; i < 400; i++)
{
var v = b.GetType("Lib.H" + i)!.GetField("V")!.GetValue(null);
if (v?.GetType().Name != "Marker") Console.WriteLine($" Lib.H{i}.V now reads: {(v is null ? "null" : v.GetType().FullName)}");
}
GC.KeepAlive(junk);
GC.KeepAlive(ctx2);
[MethodImpl(MethodImplOptions.NoInlining)]
WeakReference StartOld()
{
var cell = new Assembly[] { Load(libPath, out var ctx) };
var t = new Thread(() =>
{
TouchAndForget(cell); // this thread now holds a LOADERHANDLE in ctx1's allocator
touched.Set();
release.Wait(); // stays alive across the unload + reload
}) { IsBackground = true, Name = "tls-holder" };
t.Start();
touched.Wait();
holder = t;
var w = new WeakReference(ctx);
ctx.Unload();
return w;
}
[MethodImpl(MethodImplOptions.NoInlining)]
static void TouchAndForget(Assembly[] cell) { Call(cell[0], "Lib.Tls", "Touch"); cell[0] = null; }
static Assembly Load(string path, out AssemblyLoadContext ctx)
{
ctx = new AssemblyLoadContext("c" + Guid.NewGuid(), isCollectible: true);
return ctx.LoadFromStream(new MemoryStream(File.ReadAllBytes(path)));
}
static void Call(Assembly a, string type, string method) => a.GetType(type)!.GetMethod(method)!.Invoke(null, null);
Run it
dotnet build -c Release lib/lib.csproj -o out/lib
dotnet build -c Release host/host.csproj -o out/host
dotnet out/host/host.dll out/lib/lib.dll control # expect: bad after … : 0
dotnet out/host/host.dll out/lib/lib.dll # expect: bad after … : 1 (or AccessViolationException, exit 134)
# linux-x64 on an Apple-silicon host (colima): the out/ folder must live under $HOME to be mountable
docker run --rm --platform linux/amd64 -e DOTNET_EnableWriteXorExecute=0 -v "$PWD/out:/r" \
mcr.microsoft.com/dotnet/sdk:10.0 dotnet /r/host/host.dll /r/lib/lib.dll
What the host does:
- Loads the library into collectible context A. It starts a background thread that calls
Tls.Touch(), drops its reference to A, and parks. - Unloads A and runs
GC.Collectuntil A'sWeakReferenceis dead. The parked thread keeps its stale handle. - Loads the library into context B, calls
Probe.TouchAll()(400 GC-statics boxes, one slot each in B'sm_slots), then callsTls.Touch()on the main thread. B'sTlsreuses A's cleared index. - Repro: releases the parked thread and joins it (thread exit). Control: keeps it parked.
- Runs five
GC.Collects, allocates 200,000 small arrays, collects again, and counts statics that no longer read aMarker.
| runtime | repro (thread exits) | control (thread kept alive) |
|---|---|---|
10.0.11, macOS arm64 |
5 of 5 runs: 1 static corrupted | 5 of 5 runs: 0 |
10.0.12, linux-x64 (container, DOTNET_EnableWriteXorExecute=0, see below) |
3 of 3 runs: H0.V now reads null |
3 of 3 runs: 0 |
The failure can take either form. In one arm64 run the dangling base pointed at a reused object, and
calling GetType() on it died with System.AccessViolationException → Fatal error. → exit 134.
That is sighting #18's exit shape. On x64 the reclaimed memory read as zero.
🚨 Under x64 emulation (colima, --platform linux/amd64), WX on made even the control
segfault. A non-collectible GC baseline passed in the same container. With
DOTNET_EnableWriteXorExecute=0, the control passes every run and the repro corrupts every run. So
the emulated WX crash is an artefact of the emulator, not a finding. Do not read it as evidence
either way.
Evidence from the #18 dump that the preconditions hold in our process
Read with ClrMD over MeshWeaver.Futu-1773.dmp (the probe walks the three collectible allocators'
m_slots):
- The allocator whose box went missing (
LoaderAllocator 0x7fd74200c3c0, thev7-…NodeType compile) holds a collectible thread-static data block.m_slots[214]is anObject[1]containing aSharedArrayPoolThreadLocalArray[], which is the[ThreadStatic] t_tlsBucketsofSharedArrayPool<LineOfBusiness>. Next to it,[213]holds the pool's GC-statics box and[215]holds its lambda cache. AnyArrayPool<T>.Shareduse over a NodeType-compiledTcreates one, including the BCL's own pooled builders (LINQToArrayand similar). Our code never has to write[ThreadStatic]itself. - The process had three collectible NodeType contexts resident (
v7,v14and a third). Every recompile retires one and mints another, so TLS indices are cleared and reused all the time. - V's box is missing from
m_slots, andm_slotshas no null entry inside its used range. Step 6 predicts exactly that. The earlier reading of that fact as "never registered" is not the only explanation, because a freed-then-reused slot looks the same.
Not established from the dump: which thread exited, which index was reused, and which handle was
freed. The per-thread pLoaderHandles arrays of threads that have already exited no longer exist, and
the dump cannot show history. The claim is that this mechanism produces exactly #18's state, and all
its preconditions are present in #18's process. It does not prove that this mechanism is what
happened there.
Why this codebase meets it often: thread exits
The trigger is a thread exit. This process has two steady sources of them:
- ThreadPool workers retire after about 20 s idle. Those are the threads hub turns and
deserialisation run on, so they touch NodeType-typed
ArrayPool<T>constantly. - Until #4654, every blocking lane of the pools
IoPoolRegistryregisters started a fresh thread per burst, and that thread exited when its queue drained (it now keeps them parked — see Remedies). This covered the bounded production pools: themw-cpu-lanecompile lane and themw-io-laneIO lanes. It never coveredIoPool.Unbounded, whoseInvokeBlockingruns on the ThreadPool throughTask.Runand is covered by the ThreadPool bullet above.LimitedConcurrencyLevelTaskSchedulerthen rannew Thread(_ => DrainQueue())fromNotifyThreadPoolOfPendingWork, andDrainQueuereturned when_taskswas empty. The CPU lane worked this way since the dedicated-thread compile lane, and every IO lane since #5678 (blocking leaves off the ThreadPool, merged 2026-09-25 06:43Z; the images running on the control and public instances on 2026-09-26, core4c8530d7dd, contained it). Each blocking-leaf burst was therefore one thread exit, and one more chance to free a handle belonging to a context unloaded since that thread first ran.
Combined with the recompile cadence measured on the control instance (one new compiled assembly every ~8 s for hours, #4654), the three conditions the defect needs (retire a context, reuse its index, exit a thread that held the old handle) are routine, not rare.
What this does and does not explain
| sighting #18 (CI, exit 134, dangling collectible static base) | Explained by mechanism. It reproduces the exact state. The dump cannot show which thread exited |
sightings #1–#17 (CI, exit 139, a MethodTable word reading exactly zero inside libcoreclr) |
Consistent, not shown. A static that dangles into reclaimed memory reads null (the x64 repro) or a wrong object. Once code copies that pointer into a live object, the GC marks through a pointer whose "MethodTable" is zeroed free space, which is that family's fingerprint. Nobody has traced a single #1–#17 dump back to a freed handle |
#4654 and the ci.9218 portal deaths (exit 139 at addr (nil) and at 0x1880000005; exit 134 ×3) |
Not read. The production dumps are on the memex-data PVC behind break-glass access. Their fields are compatible with this mechanism and with others |
Remedies
The root fix is upstream. It has two possible shapes. FreeTLSIndicesForLoaderAllocator could
null every live thread's pLoaderHandles[index] for the indices it clears, under the same
g_TLSCrst that FreeLoaderAllocatorHandlesForTLSData already takes. Or a thread's handle could
carry the identity of the allocator that minted it, with FreeLoaderAllocatorHandlesForTLSData
skipping on a mismatch. The repro above is the report. Filing it on dotnet/runtime is a public act
on the maintainer's behalf, and it has not been done.
In this repository, the defect fires only on a thread exit, so two changes remove the trigger for the threads this process controls until a runtime fix ships. Both are stopgaps against a runtime defect, named as such, and both are applied (#4654):
- ThreadPool workers no longer retire. The portal chart sets
DOTNET_ThreadPool_ThreadsToKeepAlive=-1(deploy/helm/templates/memex-portal/deployment.yaml). Measured on .NET 10 withDOTNET_ThreadPool_ThreadTimeoutMs=500: a 40-item burst leaves 18 workers, and all of them retire within 3 s; with the keep-alive set, all 18 remain. The cost is that the pool's peak thread count stays allocated, parked. It takes effect on the next roll of a deployment that renders this chart. IIoPoolblocking lanes keep their threads.LimitedConcurrencyLevelTaskSchedulerstarts a lane thread on demand (never more than the cap), PARKS it when the queue is empty, wakes it for the next leaf, and lets it exit only when the owningIoPool's disposal completes (Complete()). Before, every burst at an idle lane started a thread and every drain ended one. Pinned byBlockingLaneThreadsAreKeptTest: 20 sequential bursts run on ONE thread (the exit-per-drain shape, restored as a negative control, uses 20), a warm lane still runs its full cap at once, and disposal releases every kept thread.
Neither is the fix. Threads still exit at mesh teardown (every test host does that), and anything outside these two pools that starts and ends threads is untouched.
Production sighting: memex-cloud, 2026-10-07 09:53:29Z
Read from the createdump output in Loki through governed Logs actions
(Ops/Actions/logs-memexcloud-20261007-q5kxg-* on the control instance). The dump itself was not
read.
- Pod
memex-portal-deployment-748f7f4577-q5kxg:Application startedat 08:59:29Z, so 54 minutes of uptime, not the ~21 h of the original #4654 crashes. Heap 2.69 GiB,serverGC=True, 22 ThreadPool threads andpoolPending=0at the last heartbeat (09:52:51Z). Nothing atinfo/warnin the last 90 s beyond heartbeats and two GitHub-webhook warnings. Crashing thread 45de signal 11;NT_SIGINFO … signo 11 code 0001 … addr 0x8:SEGV_MAPERRat offset 8 from a null base. NoUnwind: exception typeline, so this is not a managed exception routed throughcreatedump.- Thread
45dehad no managed frames. ItsUnwind: thread 45deline is followed directly by the next thread's, where every application thread printsUnwind: managed frames. Its id is among the newest of the 127 threads in the dump (ids run up to47cb; the long-lived runtime threads are0001–0036), so it was created shortly before the crash. A recently created thread with no managed frame is a thread starting or ending, which is the window this defect's free runs in. The logs cannot say which, and pids are shared with child processes (git), so the id gap is not a thread count. - The
RIPprinted for45de(00007f8e66665913, beside the other threads' libc wait addresses) is the signal handler's, not the fault's, and its stack was "found in other mapping" of 3 pages, the alternate signal stack. The fault registers exist only in the dump. - 🚨 The dump is gone.
DOTNET_DbgMiniDumpNamewrites to thememex-dumpsemptyDir, which dies with the pod, and the 13:09Z roll replaced the pod. With rolls this frequent, a production dump survives for hours at most; reading one needs a copy taken before the next roll. (Since then dumps go to the/dataclaim and outlive the pod, from each instance's next Reconcile: Debugging Native Crashes → "Production: where a dump lands".)
Not established: that this crash is this mechanism. It is compatible with it (a young thread with no managed frame, a null-based read) and with others.
The same day's restarts on the pearl instance are a different defect: all three of its dumps in
the 48 h to 16:08Z (2026-10-06 01:07Z and 12:42Z, 2026-10-07 12:49Z) are signo 6 with
Unwind: exception type System.OutOfMemoryException — a managed out-of-memory abort, not a native
fault (Ops/Actions/logs-pearl-20261007-crash-4654).
Related
- Debugging Native Crashes: sighting #18, and the fact-6 recipe that found the dangling base
- Controlled IO Pooling: the
IIoPoollanes, whose threads are now kept between bursts - Node Type Compilation: where collectible contexts are minted and retired