A supervisor that re-settles the thread it gave up on
Measured 2026-09-21 on LinkedIn/_Thread/bug-report-from-carson-via-claude-over-m-1475
(the public instance), filed as #2229, first half fixed in #2236.
What it looked like
One thread, supervisorRetries: 2, one pending user message from 2026-07-28 never ingested
(ingestedMessageIds: []). Its supervisorNote had reached 27,927 characters holding 45 nested
copies of the supervisor's own diagnosis — innermost the original Parked … never ingested, then
44 (Stale) layers wrapped around it, one every ~900 s from 08:20Z to 19:50Z. Summary was
"Error: " + supervisorNote, exactly, and the whole string was copied into the thread's
Admin/Threads row (lastFailure 27,993 chars).
Growth rate: +607 bytes per sweep, ≈58 KB/day, unbounded, on a field every reader of that row loads.
Four defects, and how they compose
A — three submissions in 563 ms for one thread
ThreadSupervisor is an IHostedService: one per process, and the mesh ran five replicas. The
terminal branch decided whether to file from thread.SupervisorFeedbackPath read out of its own
sweep snapshot, with no claim and no version pin. Three replicas reached that branch in one 60 s
tick and each filed — two became GitHub issues in a satellite repository, the third was orphaned.
SupervisorFeedbackPath is single-valued and last-writer-wins, so the node remembers one of them.
The tell that it was three snapshots rather than one event: the third submission quotes node version 214980 where the other two quote 215216.
B — the diagnosis read a field the settle writes
Settle writes the diagnosis into both Summary (as Error: {diagnosis}) and SupervisorNote,
in one with. Diagnosis embedded Summary as last error:. So the composition
Diagnosis ∘ Settle's-write grew its own input on every application — neither line wrong when
read alone, which is why no test saw it.
The guard that should have stopped the second application was a case list:
if (thread.SupervisorFeedbackPath is not null && verdict.Kind == VerdictKind.Parked)
return …; // inside a block whose `case Parked:` and `case Stale:` share one body
Stale fell through. The kind was never what decided it; whether it had been filed was.
C — the settle re-arms the round it just gave up on
Settle writes ResetExecution(), which sets Status = Idle and deliberately preserves
PendingUserMessages — "the inbox queue — the next round drains anything still pending". That
write reaches a cold address, activating the hub, whose submission watcher sees Idle + pending
and claims → StartingExecution without stamping a timestamp. The round does not advance; 900 s
later the sweep reads that as Stale; defect B re-settles; round again. The supervisor is the
loop's clock — supervisorLastActionAt moved 19:35:47.84 → 19:50:52.10, i.e. 904 s.
This is why "SETTLED" is not a terminal state today, and it is a design question rather than a bug to patch: the remedy changes what happens to a real user's unanswered message.
D — the staleness gauge is blind to the writes actually happening
Classify measures quiet from node.LastModified. On this node ~250–640 writes/s did not advance
lastModified: it stayed at 19:50:52.997 across reads at 19:47, 19:49:56, 19:50:10 and 20:04:19 while
version climbed by ~300,000. So the gauge read "unchanged for 15 minutes" about a node being
written several hundred times a second, and classified it Stale on that basis.
🚨 The write storm is NOT the supervisor. 182,624 writes landed in a window sitting entirely
between two settles. The emitter is unestablished: the code names a candidate shape (the submission
claim oscillating StartingExecution → rollback → Idle → re-claim, and RollClaimToIdle) but that is
a hypothesis, and the one instrument that would close it — a Logs instance action — failed to
answer. Do not read a still-climbing version as the B fix having failed.
What #2236 changed, and the two traps in doing it
The invariant, not the case list
internal static bool AlreadyFiled(MeshThread thread) => thread.SupervisorFeedbackPath is not null;
and the branch choice is now one pure function, Choose(thread, verdict, cfg), with the filed arm
ahead of Wake, Relaunch and Settle.
🚨 The ORDER is load-bearing, and putting the guard in the terminal branch is not enough. cfg is
the live ThreadSupervisorStatus node, so an operator can raise MaxRetries at runtime — and a
thread already given up on and filed at SupervisorRetries == 2 then matched the Relaunch arm the
moment the cap went to 3, and was recycled for a thread nobody was waiting on. Report stays ahead
of the filed arm on purpose: a filed thread the pool has picked up is live activity, and suppressing
that row would hide it. So the arm covers the branches that act, which is what the invariant is
about — it is not the case-list-over-kinds that caused the defect.
🚨 The provenance test is the CO-WRITTEN PAIR, not "has the supervisor written here"
The first attempt asked SupervisorNote is null, and it was a regression:
SupervisorNoteis not settle-specific — the wake and the relaunch write it too;- a round's own terminal write sets
Summarywhile clearingExecutionStatus.
So on round fails → supervisor wakes it → next sweep Parked/Stale, that form dropped a genuine
provider error and emitted (the thread carries no error text — nothing ran far enough to write one).
A confident false statement, which is worse than the verbose true one it replaced — and invisible
to a control that leaves SupervisorNote null, which the first control did.
Because Settle writes the two fields in one with:
var supervisorsOwn = thread.SupervisorNote is not null
&& thread.Summary == $"Error: {thread.SupervisorNote}";
var roundError = supervisorsOwn ? null : thread.Summary;
That is "is this Summary still the one Settle wrote?". Any later writer — a round's terminal
write, a wake, a relaunch — breaks the pair, and a broken pair means Summary is the round's and is
read. It inspects no text for supervisor-ish wording; it asks whether two fields are still the pair
one write produced.
Why Summary keeps being written
Thread.Summary is a contract, not a scratch field: a delegating parent reads it as its tool-call
result, and every terminal write is required to carry one (a failed round's starting Error:). A
settle that wrote none would hand a delegating parent an empty result for a thread that genuinely
failed. The fix is to stop reading it back, not to stop writing it.
Two things this cost, stated because they are the reusable lessons
It did not only make the note grow — it SUPPRESSED. A non-empty Summary short-circuits ahead of
ExecutionStatus, so for any thread the supervisor had ever settled, every later genuine failure was
reported with the supervisor's stale sentence instead of the real error. That was not in the original
report; the control found it. Verbose-but-true and silently-wrong are different severities, and this
was both.
The identity is not the fingerprint. summary == "Error: " + supervisorNote holds after one
correct settle, so asserting its absence would have forced the contract break above. The defect's
fingerprint is growth under iteration — Diagnosis(SettleProjection(t)) ⊃ Diagnosis(t) — which is
why the live node carried 45 nestings and not one. The same identity is the right discriminator for
whose Summary this is; it is simply not evidence of the loop.
How to test a fixed point
The defect lives in the composition, so Settle's inline write is lifted to a pure
SettleProjection and Diagnosis made internal — the seam Classify already has, for the same
reason. ThreadSupervisorSettleLoopTest then iterates the pair rather than inspecting either line.
🚨 A test of the helper is not a test of its use. The first version asserted AlreadyFiled
directly, which passes even if nothing consults it and cannot see where in the branch order it is
consulted. Testing Choose — the single decider Act switches on — closes both.
Controls that make the suite non-vacuous, verified in both directions against builds that differ only in the fix under test:
| reverted | cases that fail |
|---|---|
| B (the whole defect) | growth at sweep 2, containment, five-re-entry (visibly nested five deep), suppression, and AlreadyFiled false for a filed Stale thread — 5 of 14 |
| only the two review fixes | RaisingMaxRetries_DoesNotRelaunchAThreadAlreadyFiled ("Expected None … found Relaunch") and the three provenance cases — 4 of 14 |
The cases that pass on both sides are the controls: a round's own error still reaching the diagnosis, an unfiled thread still actionable, and the other three arms still firing.
C, fixed: the settle is TERMINAL
The policy is one sentence — every failure path ends in a stamped terminal state on the node
(thread-graceful-error) — and the old settle did not meet it:
Idle + pending is not terminal, it is the exact state the submission watcher claims. So the
supervisor was not ending the failure, it was scheduling its next attempt, every 900 s, under a row
that said Failed. Three things changed, and their ORDER is the design:
- The input the supervisor gives up on is ANSWERED. Every pending message is materialised as its
user cell, and ONE response cell is created at the round's own deterministic id
(
ThreadSubmission.DeriveDeterministicResponseIdover the drained ids — the idDispatchRoundwould have minted for exactly this input),Status: Error, text*Error: {diagnosis}*,Summary: Error: {diagnosis}. The cells are created through the mesh's node operations, which reach the store without touching the thread's address, so nothing wakes the hub yet. - ONE thread write then drains the input as answered (
ThreadSubmission.IngestAsAnswered: ids intoMessages/IngestedMessageIds/UserMessageIds, out ofPendingUserMessages, the response id after them), resets execution, and records the verdict (SettleProjection). After it the thread is Idle with nothing pending — Healthy to the next sweep, however long it waits. The user sees their message and, under it, the error; a resubmit is their call, and it opens a new episode (below). - The hub's half. The thread write wakes the hub, and the hub's watcher claims the LOADED — still
pending — state before that write lands, so a fresh claim races the settle.
DispatchRoundnow reads the cell it finds "already exists": a fresh dispatch whose deterministic response cell is already terminal is a round that is over — it ingests its input as answered by that cell and settles Idle instead of running the turn into it and overwriting the verdict (the ruleClassifyResumeapplies to a resume, applied to a fresh claim). Both writers converge on one state whichever lands first; the owner's three-way merge sees identical values (the cell'sSummaryIS the settle's), and aConflicton the one instant both stamp (lastActivityAt) re-runs the supervisor's lambda against the drained node, where it lands.
Two consequences that were not the headline:
- A filed thread that still classifies Parked or Stale is SETTLED, not skipped.
Choose's filed arm answeredNone, which was right against relaunching and wrong against settling: every thread filed before the settle became terminal — the live LinkedIn thread among them — would have stayed in its cycle for ever, un-supervised. The filed arm now settles (terminal, idempotent, files nothing a second time), and its log line is a warning that names the earlier filing rather than a second "giving up on" — which is also why the give-up count stops measuring sweeps (#2262). - The four
supervisor*fields describe an EPISODE, and new input opens a new one.SupervisorRetrieswas never reset, so a thread once exhausted had zero relaunches left for any later input, and a thread once filed was never supervised again.Thread.ResetSupervision()runs inThreadInput.ApplyUserInput(the one funnel behindSubmitMessage/SubmitComposer) and inResubmitMessage; the supervisor's writes are guarded on the episode they diagnosed (SameEpisode), so a submission that lands between the sweep's read and the write is left to the next sweep instead of inheriting the old episode's verdict — and a sibling replica's settle is not overwritten either.
Pinned by ThreadSupervisorMeshTest against a Monolith mesh — where every hub is healthy and the
model answers, which turns the defect into a POSITIVE signal: on the unfixed code the settled
thread's response cell read Completed with the model's ack (the hub ran the round the supervisor
had just given up on); on the fix it reads Error with the diagnosis — plus the filed-and-still-
pending shape, the new-episode reset through SubmitMessage, and the hub's half on its own (a
seeded terminal cell, a read as the wake, the input ingested and the cell untouched). The pure half
is in ThreadSupervisorSettleLoopTest: the settled thread has nothing pending and is Healthy a day
later, the settle's response id equals the dispatch's, input drained meanwhile is not answered
twice, and a new submission opens a new episode.
D, still open — and the reading it forces on every "no change" sentence
Classify measures quiet off node.LastModified, and core's own-stream write path (UpdateOwn)
applies no audit stamp: the claim, a rollback and a commit bump Version and leave LastModified
alone, while every CROSS-hub write stamps it. So a hub oscillating claim → dispatch fault → rollback
at hub speed reads, to this gauge, exactly like a dead one — which is what the #2262 samples say
("no change to the node since T, 18 min ago" after a relaunch whose write landed at T and whose
woken hub then did a great deal, none of it visible). The Stale observation now says what was
measured and names both readings. Nothing here decides between them; the terminal settle ends the
loop in either case, because after it there is nothing left to claim.
A, fixed: the create IS the claim
A read-then-create across processes has no seam a guard can close: every replica's read happens
before any replica's create. So the submission's id is no longer fb-{now}-thread-{random} but a
pure function of the EPISODE (ThreadSupervisor.FeedbackIdFor): fb-thread- + 16 hex digits of
SHA-256 over the lower-cased thread path and the pending input — each pending message's id AND its
submission timestamp — or, with nothing pending, the active round's id. Every replica that gives up
on the same episode names the same node; the store admits one create, and the others are refused
NodeAlreadyExists, which FileFeedback reads as "another supervisor filed this episode" and
answers with that path. Every supervisor then settles naming the ONE submission that exists.
The key also carries Thread.SupervisionEpisode, a token ResetSupervision mints on every new
input: a round stale with no pending input and no active cell otherwise hashes the same active:-
in every episode, and a later give-up would be mistaken for the old claim. Both the winner and every
refused supervisor take the one read that activates the submission's hub, so a winner that dies
between its create and its read cannot leave a submission nobody handed to triage.
What the key deliberately leaves out: status and version (replicas read different snapshots of a thread whose claim oscillates — the production triple quoted two node versions) and the supervisor's own bookkeeping. What it keeps: the timestamp, because a resubmit reuses the message id and opens a new episode that must be able to file again.
Pinned by ThreadSupervisorFilingMeshTest — three supervisor instances, each with its own live
configuration watch, sweep one exhausted thread at once against a mesh where a submission can land.
With the random id restored it reads found 3; with the episode id, one. The pure half
(TheFeedbackId_IsTheEpisodes_NotTheSnapshots, ANewEpisode_GetsANewFeedbackId) is in
ThreadSupervisorSettleLoopTest.
Still open
D (above). And the question the supervisor cannot answer by construction: WHY the round never advances on the threads it settles — the supervisor sees only the node, never the process that owed it a round. The settle now leaves that question on the node, in the Error cell's text, instead of a pending message and a 900 s clock.
Delivery note: ThreadSupervisor is an IHostedService, one per process, so none of this arrives by
recycling a node — it needs a new image. A thread filed before the fix is settled terminally by the
first post-roll sweep that meets its bound.