The sync/ Hub Population

A 26 h production replica held 9 398 MessageHub instances, of which 6 925 were sync/ hubs at RunLevel=Started โ€” about 2.7 GB of a 3.5 GB live heap (#3432). This page is what the code says about where they come from and what is supposed to take them away. It is a narrowing, not a diagnosis: the sections marked ๐Ÿšจ NOT ESTABLISHED are still open.

๐Ÿšจ The cause was established on 2026-09-10, by experiment on a live replica: the READ path minted one sync/ hub per GetDataRequest. See The Read Path Minted a Hub Per Read, which also carries the live census that decomposes the population into its holders โ€” and re-scopes the evicted-stream mechanism to 1 parked stream in 475.

1. A sync/ hub is not a population โ€” it is one field of a stream

SynchronizationStream<T>'s constructor mints exactly one, unconditionally, for every stream:

// src/MeshWeaver.Data/Serialization/SynchronizationStream.cs
var syncHub = Host.RunLevel > MessageHubRunLevel.Started
    ? null
    : Host.GetHostedHub(
        SynchronizationAddress.Create(ClientId), ConfigureSynchronizationHub, HostedHubCreation.Always);
โ€ฆ
Hub = syncHub;
syncHub.RegisterForDisposal(_ => ReleaseHub());

And one cross-hub subscription builds two streams, hence two hubs:

side who builds it where
client JsonSynchronizationStream.CreateExternalClient, Host = the subscribing hub JsonSynchronizationStream.cs
owner Workspace.SubscribeToClient โ†’ CreateSynchronizationStream โ†’ ReduceStream(โ€ฆ WithClientId(request.StreamId)), Host = the owner's hub Workspace.cs ยท WorkspaceStreams.cs

Both carry the same ClientId, so the two hubs share an address string while living under different parents.

This settles the first of #3432's open questions. The dump's own arithmetic agrees:

6 925 Started sync/  +  1 495 Dead sync/  =  8 420
SynchronizationStream total                =  8 461     (0.5 % apart)

So there is no separate "sync hub leak" to explain. The question is entirely "why are there 8 461 live SynchronizationStreams", and each one costs its own Autofac ILifetimeScope and TypeRegistry (~390 KB, of which ~173 KB is framework metadata duplicated per hub).

2. What retires one โ€” three paths, and only one of them applies here

๐Ÿšจ This table is about the STREAM's lifetime. The registry-level account โ€” which collection roots a sync/ hub, what takes it out, and which (RunLevel, IsDisposing) states can persist there โ€” is ยง8 below, written against the code after the retire-and-replace change, and it is the half that says which of these hubs cannot be removed versus is removed only when X happens.

path where which population it reaches
the stream is disposed SynchronizationStream.Dispose() disposes its own Hub the 11
the parent hub is torn down (circuit ends, DisposeRequest, recycle) HostedHubsCollection โ†’ the RegisterForDisposal(_ => ReleaseHub()) hook (#3427) the 1 485 measured Dead
the cache releases the path MeshNodeStreamCache idle sweep โ†’ TryReleaseUnwatched โ†’ DetachUpstreams โ†’ TearDownEntry โ†’ stream.Dispose() the Started population โ€” and nothing else does

The middle row is #3321 step 3, and it is why the Dead bucket exists at all. It does not touch a hub whose parent is alive. So for a Started sync/ hub under a live parent โ€” which is this issue's entire 6 925 โ€” the idle sweep is the only reaper in the process.

MeshNodeStreamHandle.DetachUpstreams() is reachable from nowhere else, and it is also the only thing that collects the change-feed-parked predecessors described below.

3. The sweep's predicate, and why traffic defeats it

// src/MeshWeaver.Hosting/MeshNodeStreamCache.cs โ€” Entry
public bool TryMarkIdleEvicted(TimeSpan idleWindow)
{
    lock (gate)
    {
        if (evicted || subscribers > 0
            || Environment.TickCount64 - lastActiveAt < (long)idleWindow.TotalMilliseconds)
            return false;
        evicted = true;
        return true;
    }
}

Window 10 minutes, swept every minute (MeshNodeStreamCacheOptions). Two conjuncts, and the option's own summary states what refreshes the second: "Every GetStream subscription, unsubscription and Update on the path refreshes the window."

๐Ÿšจ The subscribers count is NOT the Rx-subscriber trap Workspace warns about. That warning ("the reduce chain subscribes to the stream ITSELF, so an evicted stream measures 2โ€“3 subscribers and never reaches zero โ€” implemented, measured and reverted") is about a different, abandoned mechanism at the workspace layer. Entry.subscribers is a purpose-built refcount moved only by TryAddSubscriber/RemoveSubscriber from the cache's own SharedView, and the internal hydration subscription bypasses it. It genuinely can reach zero. Anyone reasoning about #3432 from the Workspace comment alone will mis-attribute this.

What actually holds the sweep off is the other conjunct.

4. The change-feed orphan โ€” the write both creates it and postpones its collection

๐Ÿšจ Since #1174 this runs only for change events no mirror can be held to (a delete, a recreate, a version-less recycle broadcast); a versioned Updated keeps the stream โ€” see Live Mirrors and the Change Feed. The measurement below predates that.

Workspace.EvictForPath runs on every change-feed event for an owner path and removes every identity's cached stream for that owner:

_evictedRemoteStreams[removed.Value] = 0;
ReclaimIfUnheld(removed.Value);        // disposes ONLY if no declared lease remains

The mesh-node cache's hydration is a declared holder for the entry's whole life โ€” that is deliberate, and MeshNodeStreamHandle.Subscribe says so: "The lease lives exactly as long as this subscription. The shared mesh-node cache's hydration comes through here, so its entry IS the declared holder of the path's upstream for the entry's whole life โ€” which is what makes every OTHER (write-scoped) stream for the same path reclaimable on eviction."

So on the first write to a cached path:

  1. the entry's stream S1 is removed from _remoteStreamCache and parked;
  2. ReclaimIfUnheld(S1) finds the hydration's lease and does not dispose it;
  3. nothing re-points the hydration โ€” the sweep "only ever CLOSES idle entries; it never re-subscribes anything (the 2026-06-08 rule)" โ€” so the entry keeps reading S1 forever;
  4. the next reader or writer misses the cache and builds S2, which becomes the cached one.

From then on the path holds two live streams โ€” four sync/ hubs, ~1.5 MB โ€” where the design intends one. Write-scoped streams still churn correctly (an unleased evicted stream is reclaimed at once), so the accumulation is a floor of one orphaned pair per cached, written path, not growth per write.

๐Ÿšจ And the same act does both halves: the Update that fires the change-feed event which orphans S1 is also an Update that Touch()es the entry, pushing the 10-minute window out. A path written at least once every ten minutes can never be released, so its orphan can never be collected. The only reaper is switched off by the traffic that feeds it.

That is the answer to "is the retirement path missing, or present-but-not-firing?" โ€” present, and structurally unable to fire on exactly the paths that generate the garbage.

5. Bounding

Neither cache has memory-pressure eviction or any other TTL:

So the live-stream count is bounded only by (distinct paths ร— identities) touched since process start, with the per-path floor above added on top.

6. ๐Ÿšจ What is NOT established

๐Ÿšจ Partly superseded. The Evicted-Stream Retention closes the "what holds them" half for one named mechanism, with a controlled two-arm test rather than a dump: a change-feed eviction parks a remote stream, and ReclaimIfUnheld returns without disposing whenever the stream carries no lease entry at all โ€” so every unleased call site (including LayoutExtensions.GetControlStream, i.e. every rendered layout area) retains one stream, and two sync/ hubs, per change event. It also rules out static state and timers as roots, and replaces the referrer walk below with two field reads on the singleton Workspace. What stays open is the magnitude, not the mechanism.

7. The discriminator โ€” a ratio, from data the existing plan already collects

#3432's plan calls for a ClrMD referrer walk, which is the expensive half and needs a dump that (measured) restarts the replica. ยง1 makes a much cheaper reading decisive, and it comes off dumpheap -stat, which that plan already takes:

count MeshNodeStreamCache+Entry alongside SynchronizationStream.

reading what it means what to do
streams โ‰ˆ 2 ร— entries the ยง4 floor dominates: one change-feed orphan plus one live stream per cached path a fix with a named code path โ€” release the entry's lease on eviction (or re-point the hydration) so ReclaimIfUnheld collects S1. Halves the population without touching anyone's live view
streams โ‰ˆ 1 ร— entries no orphaning; the population is one stream per cached path the leak framing is falsified โ€” this is per-hub DI cost (#3432's third outcome row); close it and open the metadata-sharing issue
streams โ‰ซ 2 ร— entries something outside the mesh-node cache mints them look at the identity dimension of _remoteStreamCache's key (ยง5) before anything else

No referrer walk, no field traversal, no second dump โ€” one extra line of a histogram the plan already produces. ๐Ÿšจ It does not replace the referrer walk for the third row; it tells you whether you need one.

8. The registry account: what roots a sync/ hub, and what takes it out

Sections 1โ€“7 are about the STREAM. This section is about the HUB, read off HostedHubsCollection and MessageHub as they stand today. It answers the question the dump's histogram poses โ€” what can a Started sync/ hub be? โ€” by construction rather than by inference, and it distinguishes cannot be removed from is removed only when X happens.

8.1 One root, and it is the parent's registry

A sync/{clientId} hub is a HOSTED hub. There is exactly one creation site (ยง1), and every hosted hub enters its parent's registry through one method:

// src/MeshWeaver.Messaging.Hub/HostedHubsCollection.cs โ€” Track
private void Track(IMessageHub hub)
{
    messageHubs[hub.Address] = hub;
    hub.RegisterForDisposal(h =>
        messageHubs.TryRemove(new KeyValuePair<Address, IMessageHub>(h.Address, h)));
}

messageHubs is an instance ConcurrentDictionary<Address, IMessageHub> on the collection, whose lifetime is the parent hub's. There is a second root beside it, retiring, a reference-keyed dictionary of hubs taken out from under their address but still tearing down; Hubs enumerates both, which is why a census walking Hubs sees a retired hub too.

8.2 Three removals, all of them self-driven

# what removes the entry when it runs
1 Track's own registrant in the hub's ShutDown phase โ€” DisposeImpl is reached from nowhere else, and it is what walks the composite the registrant sits in
2 RetireCorpse only from a lookup for that exact address, only with HostedHubCreation.Always, only at RunLevel >= ShutDown, only while the collection itself is not disposing. Moves the hub to retiring
3 the retiring entry's own subscription on that hub's DisposalCompleted, i.e. at Dead

Nothing else removes a hosted hub from its parent's registry. There is no sweep, no idle eviction, no cap, and nothing anywhere asks "is this hub still needed". Removal 1 is the hub disposing itself; removals 2 and 3 only move a hub that is already disposing. So the registry is a pure consequence of who calls Dispose(), and for a sync/ hub that is: the stream's own Dispose(), the parent's teardown, or a routed DisposeRequest โ€” ยง2's table.

๐Ÿšจ Removal 2 is lookup-triggered, not self-triggered. Retire-and-replace made the hand-out safe; it did not make the registry self-cleaning. A hub wedged at ShutDown that no Always lookup ever asks for again stays in messageHubs for the parent's life. That is #4883's wedge showing through, deliberately left there.

8.3 The (RunLevel, IsDisposing) cross-tab โ€” which cells can persist

IsDisposing is disposalStarted, set as the FIRST statement of Dispose(), before it posts ShutdownRequest(Quiescing). The RunLevel only moves when the action block dequeues that request.

cell what it is can it persist?
Started, false nobody has asked this hub to die Yes, indefinitely โ€” by construction, not by defect. Removal 1 is the only route out and nothing has started it. Whether the hub is garbage is entirely a question about the stream that owns it
Started, true Dispose() ran; the turn was never dequeued Yes, unboundedly. Invisible to a RunLevel-only histogram โ€” it reads identically to the row above
Quiescing, true draining accepted work, off the block Bounded by the quiesce budget only if the block can dequeue the next phase
DisposeHostedHubs, true joined on the hosted subtree Yes, unboundedly โ€” the join carries no deadline by design, so one child that never completes holds it forever
HostedHubsDisposed, true โ€” No: empty by construction. The level is declared and never assigned; the phase sequence goes DisposeHostedHubs โ†’ ShutDown
ShutDown, true past the flip, not yet past the registrant walk Microseconds normally; unbounded if the ShutDown turn wedges, and then only a lookup takes it out (ยง8.2)
Dead, true terminal Should not appear in messageHubs at all: Dead is assigned strictly after DisposeImpl has run removal 1. Present โ‡’ the removal did not run
anything, in retiring retired predecessor Until DisposalCompleted โ€” so a wedged teardown keeps it, visible in Hubs and in the owner's disposal join

8.4 A second retention class the histogram cannot see: the registrant composite

A hub's registered cleanups live in one CompositeDisposable:

// src/MeshWeaver.Messaging.Hub/MessageHub.cs
private readonly CompositeDisposable disposables = new();
public IMessageHub RegisterForDisposal(IDisposable disposable)
{
    disposables.Add(GuardRegistrant(disposable));
    return this;
}

A plain registration is append-only. CompositeDisposable.Add never prunes, and nothing removes an entry added through RegisterForDisposal, so such a registrant is held โ€” with everything its closure captured โ€” for that hub's whole life. The ONE way out before the hub dies is a DETACHABLE registration (RegisterForDisposalDetachable, or SubscribeHeldUntilTerminal built on it): its handle removes the entry without disposing it, and it is to be used only once the registrant's disposal can no longer do anything (see the table below). A PLAIN registrant whose own subject is SHORTER-lived than the hub is therefore a monotone root, and it retains an object graph rather than a hub, which is why no MessageHub histogram can see it.

MessageHub.DisposalRegistrantCount is the reading that can, and it is printed in both disposal diagnostics as Registrants=. A hub whose registrant count climbs with the traffic it has served is retaining one object graph per unit of that traffic.

Two classes, both measured in MeshWeaver.Data.Test:

site growth status
JsonSynchronizationStream.CreateExternalClient registered the owner-protocol subscription on the SUBSCRIBING hub as well as on the stream +1 per remote stream ever opened, holding the whole SynchronizationStream graph (measured: floor 4 โ†’ 6 over two streams, still 6 after both were disposed) Fixed. The duplicate bought nothing, and by TWO routes rather than one: the stream's own composite rides its sync/ sub-hub (a hosted hub of the subscribing hub, torn down in its DisposeHostedHubs phase), AND Workspace.Dispose disposes every cached remote stream โ€” which its own comment says exists to release exactly this SubscribeRequest callback. A hub walks its OWN registrants only later, in ShutDown, so the duplicate was the third route and the last to fire. ๐Ÿšจ The second route is a correction owed to running the negative control: with the sub-hub hook removed the coverage test still passes, so it asserts the OUTCOME and not a route. Pinned by StreamRegistrantsLeaveTheSubscribingHubTest
the per-request handlers in DataExtensions and MeshDataSource register one cleanup per request on the OWNING hub โ€” HandleGetDataRequest, HandleSubscribeRequest, HandleDataChangeRequest, ApplyJsonMergePatchAndUpdate, ApplyMeshNodePatchInTurn, RegisterOwnerDisposingNack, HandleSaveMeshNode, the unified-reference handlers +1 or +2 per request served, measured over ten answered requests each from a floor of 3: DataChangeRequest โ†’ 23, PatchDataRequest (generic path) โ†’ 23, one-shot GetDataRequest โ†’ 13, UpdateUnifiedReferenceRequest โ†’ 13 Fixed. A registrant now leaves the hub at the moment its own disposal can no longer do anything, which changes no behaviour: a subscription at its TERMINAL (hub.SubscribeHeldUntilTerminal(source, subscribe)), a teardown NACK once its once-only answer gate is CLAIMED (IMessageHub.RegisterForDisposalDetachable โ€” the handle removes the registrant WITHOUT disposing it, one shared flag deciding against a racing teardown). A request still in flight when the hub goes down is torn down and NACKed exactly as before. Pinned by PerRequestRegistrantsLeaveTheOwnerHubTest (red 4 of 4 on the unfixed build with the readings above; green on the fix) and PatchAckTotalityTest's flush-leg assertion

๐Ÿšจ Still held by design, and NOT a leak this fix touches: a GetDataRequest over a LIVE workspace reference keeps shipping every change to its requester, so its subscription โ€” and its registrant โ€” live until the owner goes down. Whether anything should end such a read earlier (the requester's own Observe completes on the first answer) is a separate question, not measured here. A read that ends EMPTY on a healthy hub also stays registered on purpose: its caller is still owed the teardown NACK.