Module Reload
The rule (policy
module-reload-request, register). Any authorised caller can ask an instance to reload one module — or all of them — through ONE durable request. The instance resolves the newest COMPATIBLE published version, lands it, activates it (live when the module can be swapped in the running process, otherwise exactly one automatic restart through the self-update restart path), and reports the outcome on the request node. No governed activity, no approval, no person's signature, noconfirmation.
Why it exists
A module that ships is supposed to be in use (rule R3 of Module Adoption Policy).
When it is not, there was no single act that made it so. The shape that motivated this request: the
control instance kept MeshWeaver.AI 1.20.4 loaded while the Hosting module it also ran needed
1.21 (MissingMethodException: ThreadPreparation.set_Group). 1.21 was published and compatible; the
remedies were a manual sync, then a hand-filed Restart action that asked for a confirmation —
two people-steps for what is one fact the instance can establish itself. The unattended lane could
not take it either: the package's own update policy and its sync-owned partition both hold an
unattended landing (One Partition, One Bookkeeping), and a landed
generation waited for whatever restart came next.
The request
A ModuleReload node at Admin/_ModuleReload/{id}, content ModuleReloadRequest
(MeshWeaver.Graph.Configuration):
| field | written by | meaning |
|---|---|---|
module |
requester | the module's entry-assembly name (MeshWeaver.AI) or its package id (AI); blank = every installed module |
reason |
requester | required; carried into the restart announcement and every log line |
requestedBy / requestedAt |
requester | who asked — a user id, an agent, a watcher's name |
status |
executor | Requested → Landing → Activating | AwaitingRestart → Done | Failed | Faulted (open string constants, ModuleReloadStatus; Faulted is never final — see "A crash is never final" below) |
items[] |
executor | per module: runningVersion (what the executor's process ran), foundVersion + foundFloor + registry (what is published), targetVersion (what the activation record now names), landed, decision, failure |
activation / activationDetail |
executor | Live, Restart or NotNeeded, and the lane's own sentence |
liveSwapRequestedAt / restartRequestedAt |
executor | when each process was asked to swap, or when the ONE restart was requested |
replicas{process} |
each process | what that process has LOADED, per module — written under its own key, so reports merge rather than clobber |
failure |
executor | why the request is red (or faulted), by name |
attempt / faultedAt |
re-arm / executor | how many times a Faulted request has been re-armed (drives the backoff, never a cap), and when the current attempt faulted |
log[] |
executor | every step, in order |
The only writer is ModuleReload.Request, which writes the node as System after its CALLER has
authorised the requester. The executor acts only on a request whose createdBy is System; anything
else that can write under Admin/_ModuleReload is refused by name and nothing is landed — the same
trust rule as ActivationRecycle.
How it runs
The executor is the request node's OWN hub (ModuleReloadExecutor, registered as the node type's
initialization), so exactly one process drives a request however many replicas hear it. Every step is
a stream.Update on that node and every step is idempotent: an activation torn down mid-step simply
runs the step again when the node is next read.
- Resolve and land. The modules come from the install records (
Plugins/*that declare a compiled module). Per module,RegistryUpdateReconciler.ReloadModuleasks the configured registries in order — the first that serves the package answers — throughPluginBundleClient.AdoptModuleOutcomewithunattended: false: the reload is attended, like a Provision click, so the package's own update policy does not decline it. It runs on the reconciler's serialised lane and ends by proposing the module set (ModuleDependencyFloor.ProposeChecked); a refused set is the reload's failure, because neither a swap nor a restart would load the module. - Compatibility is the floor, never a seal. A bundle whose declared
minMeshVersionis above the running platform is not downloaded, the landed generation keeps serving, and the item readsdeclined — version X needs platform ≥ F; running R …(policypackage-min-mesh-version). Neither a seal for the running platform's identity nor a green platform build is consulted; whether the bytes link is measured by the landing's link probe, as for every landing. - Decide activation. A module whose
targetVersionis not the version THIS process has loaded needs activating. When every such module can be swapped (IModuleLiveActivation.CanSwap) the request goesActivatingand every process swaps; otherwise it goesAwaitingRestart. Nothing to activate isDonewithNotNeeded. - Exactly one restart. The executor stamps
restartRequestedAtFIRST, then callsIModuleActivationRestart.RequestRestartonce — the stamp is what keeps a resumed executor from asking twice. The portal'sSelfUpdateHostedServiceimplements it with its one restart path: a self-patch restart of the running image, orself-update-restart-pendinghanded to the control lane, which opens an unattendedRestartwith no approval and no confirmation (Self-Update on the Control Lane). The poller's roll floor does not defer a reload — it is one explicit request — and exactly one restart still holds against the poller's own checks: after a self-patch restart they read the fresh roll instant and defer, and on the control lane an open or freshly doneRestartalready delivers the announcement. - Report and decide. Each process reports what it LOADED (
ModuleReloadAgent, armed on the mesh hub of every process): after a restart only a process that BOOTED afterrestartRequestedAtreports (it lists open requests at boot, which also re-activates the executor), and after a live swap every process reports once it has swapped. A report is measured, not claimed: the generation directory each module was loaded from (ModuleActivationStatus.LoadedModuleGenerations) read against the activation record.ModuleReload.Evaluatedecides over the counted reports — reports from processes the cluster has recorded as gone are not counted — and the request isDonewhen every counted replica loads every target version, orFailednaming the replica and both versions. 🚨Doneis a verdict over the WHOLE roster: where the cluster can enumerate its running members (IClusterMembership.AliveMembers), every one of them must have a counted report first, so a pod still booting or still swapping keeps the request open (its swap may yet fail and need the restart fallback), and an old pod still running after a restart keeps it open until it is gone. A roster change re-evaluates the request (IClusterMembershipFeed), since no node write accompanies it. A swap failure or a version mismatch on a counted report is decisive at once. Without a roster (monolith, no cluster) the counted reports are all there is. A running member that never reports leaves the request open — visible on the node, never silentlyDone. A failed live swap leaves the previous generation serving and falls back to the one restart. 🚨 The agent reads the commit the feed ANNOUNCED, never whatever its mirror holds. It hears a request on the invalidation feed, which fires post-commit, while this process's mirror of the node receives the owner's echo on its own path — under load AFTER the feed. An unflooredTake(1)then readLandingfor the event that announcedActivating, swapped nothing, reported nothing, and nothing later woke it (the executor writes nothing until a replica reports): the request sat inActivatingwith nothing logged. So the read waits for a state at or past the event'sVersion— the floorRebaseSourceapplies to an announced version (#1174) — and a mirror that never gets there times out loudly. A path from the boot listing announces no commit and reads unfloored. The instance reboot's agent (Instance Reboot) reads the same way. Pinned byModuleReloadAgentReadsTheAnnouncedCommitTest, whose unfloored control reports nothing.
A crash is never final
The system repairs itself, so a step that CRASHED must never end a request. The executor tells two kinds of red apart:
| status | what produced it | final? |
|---|---|---|
Failed |
a DECIDED answer: a floor above the running platform, no installed module of that name, a refused bundle, a bundle the registry answers 404 / 401 / 403 for, a bundle whose bytes fail their digest, a replica that loads the wrong version after activation, a request not written by System, an install with no restart path (ModuleRestartKinds.Unavailable) |
yes — asking again would get the same answer |
Faulted |
a CRASH or an answer that may clear on its own: an exception in the landing or restart step, a registry index or a bundle download that answers 5xx (500, 502, 503, 504), 408 or 429, a transfer that timed out or whose connection was reset or refused, a module-set proposal that faulted, a reconcile lane that was shutting down, a restart attempt whose hand-over failed or whose call threw (ModuleRestartKinds.Faulted) — any item whose transient flag is set (ModuleAdoptOutcome.Transient) |
never — retried |
Which answers are transient is decided in ONE place, TransientRegistryFailure
(MeshWeaver.PluginCatalog): a status is transient when it is 5xx, 408 or 429 (the conventional
transient-HTTP set); a fault is transient when it carries such a status or never reached an answer
(a timeout, a reset or refused connection, a DNS or socket failure) and decided when it is a digest
mismatch or a definite refusal — and a definite refusal is the registry clients' own
RegistryRefusedException (no such manifest, blob or repository; a refused challenge or token; a
manifest with no bundle layer; a corrupt bundle), never a bare InvalidOperationException, which any
code the fetch pipeline runs can throw and which therefore reads as a crash, retried. The registry
feed read's retry (RegistryUpdateReconciler.ShouldRetryFeedRead) asks the same classifier. Every
download path — the HTTP bundle route and the OCI artifact path
— asks it; a second copy of the set is how a bundle download answering 503 once read as a final
Failed. Pinned by ModuleReloadTransientDownloadTest (503 → Faulted → the next pass retries and
lands; 429, 502, 504 and a timeout → Faulted; 404 and 403 → Failed, never retried),
TransientRegistryFailureTest (a RegistryRefusedException is decided, a bare
InvalidOperationException is a crash; the feed read asks the same classifier) and
PluginBundleArtifactFetchTest (the same split on the artifact path; a tampered layer → decided).
The restart lane draws the same line. ModuleRestartKinds.Unavailable means this install
CANNOT restart at all (no updater that can roll, no control inbox configured) — a configuration
answer, Failed. ModuleRestartKinds.Faulted means the path exists and this attempt failed (the
hand-over to the control lane was refused or unreachable, SelfUpdateOutcome.RestartHandoverFailed;
or the restart call threw) — Faulted, and the retry asks for the restart again
(ModuleReloadRestartAttemptTest).
A request with several modules is Faulted when ANY failing module crashed: that module's answer is
still unknown (ModuleReload.OutcomeOf).
The retry has no timer of its own. Every full reconcile pass of the plugin catalog
(RegistryUpdateReconciler: the boot pass, each safety-net tick, ReconcileNow) ends by calling
ModuleReload.RetryFaulted: a LISTING of Admin/_ModuleReload, a read of each open request from its
own node stream, and a re-arm (ModuleReload.Rearm, written through GetMeshNodeStream(path).Update
as System) of every Faulted request whose backoff is due. The backoff is the pass's own cadence,
doubled per attempt — faultedAt + interval × 2^min(attempt, 5), the interval being
PluginCatalog:ReconcileSafetyNetInterval (30 min); the doubling stops at 32 intervals, the retries
never do. With the safety net switched off, the boot pass is the retry.
The re-arm keeps the node and its log, increments attempt, clears every executor-owned field of the
faulted attempt (items, activation, replica reports, the restart stamp, the failure) and sets the
status back to Requested; the executor then runs the new attempt from the top. 🚨 The one-restart
rule therefore holds per attempt: a retry of a request that faulted after its restart was stamped
asks for a restart again.
Who can file one
| surface | authorisation | how |
|---|---|---|
MCP reload_module (MeshWeaver.Plugins McpMeshPlugin) |
IsGlobalAdmin — a platform admin on the Admin partition |
MeshOperations.ReloadModule(module, reason); returns {status, path, message} at once, read the node for the outcome |
| the package page — node menu 🔄 Reload module on an installed package that declares a module | IsGlobalAdmin, checked for the menu entry and again on the click |
the ReloadModule area (framework controls, en + de) files the request and opens its page |
| the platform's own watchers (the fleet-target intake, a release follow-through, a watchdog) | they run in-process as the platform | ModuleReload.Request(hub, …) directly |
What this replaces, and what it does not
RefreshModules self-filing is generalised into this request. The fleet-target intake
(MeshWeaver.Plugins Hosting/PlatformBuildInbox → FleetTargetIntake) used to file a Store
Maintenance task RefreshModules for a module of its OWN instance held behind by its sync. That
task re-runs the registry installer and reports "RESTART the deployment to activate them" — it lands
and stops. The intake now files a ModuleReload for that module instead, which lands AND activates
AND reports what loaded. The RefreshModules maintenance task itself stays: it is the operator's
content-and-module re-install of a whole package, a different act.
The live swap belongs to the live module loader (Doc/Architecture/LiveModuleUpdate, MeshWeaver#6121 —
slice 1 of that work puts every module in its own collectible load context; the swap is its slice 2).
This request does not fork it: IModuleLiveActivation is the reload's CALL SITE, registered by the
loader once it can swap. Until a host registers it, every reload that needs activating takes the one
restart — which is exactly the fallback rule 3 of that page prescribes.
Auto-update (policy packages-auto-update)
Every installed package updates by itself as soon as a newer compatible version is published. The unattended lane is the same machinery as an explicit reload, driven by the registry instead of a person:
- When it runs.
RegistryUpdateReconcilerreconciles on boot, on everyModulePublishedbroadcast the registry posts to this instance's inbox (minutes after a publish), and on its safety net (every 30 min). Each pass reads the registry's feed and bundle index — the package's OWN publish is the only event it needs. - What decides. Two inputs and nothing else: a newer version is served, and its declared floor
is met (
ModuleUpdateDecisionwithPackagePlatformFloorGate.HoldFor). An incompatible floor is declined by name — on the install record (heldUpdate) and in the log — and the landed generation keeps serving. - What activates. A pass that LANDED anything files ONE
ModuleReloadrequest for what it landed (auto-{hash of the landed set}, so replicas and re-runs file it once). The request activates it — live, else exactly one automatic restart. A second wave that lands before that restart has happened RIDES it (the executor finds the open request's restart stamp and does not ask for another). - Defaults. A fresh install is
Auto(PackageInstaller.SeedUpdatePolicy); the deployment-wideDefaultUpdatePolicy/AutoUpdateByDefaultare no longer consulted. At boot,PackageAutoUpdateMigrationmoves every record SEEDED with the old reminder-only default toAutothroughstream.Update(as System). It keeps — and names in one Warning per pass — a pin (None) and any policy a global administrator CHOSE on the catalog card (updatePolicySetAt, stamped bySetUpdatePolicyfrom now on).
What no longer holds an update — and what still does
| dependency | status |
|---|---|
| a platform image build, deploy, roll, CD or "arm the fleet" step | never consulted by the package lanes |
| a publication seal / a seal for the running framework identity | never consulted — the module lane decides on the floor and the link probe; the framework MVID is recorded, never a gate |
a green build of the platform or of the package repo's main |
never consulted — only the package's own publish to the registry |
| the partition's SYNC SOURCE (#4355 gate 1b: "its module waits for the same seal its content does") | retired for the module half: a published compatible module lands even when the partition's _GitSync still carries older content |
| the package's own policy | Auto by default; a pin or an administrator's explicit choice still declines, by name |
| the CONTENT half of a sync-owned partition | still the sync's — one writer per partition (#4355). The installer does not write content into a partition a sync source owns; that content follows its sync, and the sync source's own seal gate (SealedSyncGate) is not changed here. The code no longer waits for it. |
Measured on the control instance before this change (read-only search namespace:Plugins nodeType:Package): 80 install records, 0 of them not Auto (75 declare updatePolicy: Auto, 5
predate the field and carry autoUpdate: true). The reminder-only default was not what held its
modules; the sync-owned hold on the module lane was.
What is NOT established
A real cross-process run. The scenario tests (
Memex.Portal.Shared.Test→ModuleReloadByRestartTest,ModuleReloadLiveTest) run the real registry client, landing, reconciler, request node, executor, agent andSelfUpdateHostedServicein one monolith process; what they SIMULATE is which generation a process has loaded and the process that boots after the restart. A Kubernetes restart and a multi-replica Orleans cluster were not exercised."Newest compatible" across versions. The registry's bundle index advertises ONE version per package. When that version's floor is above the running platform the reload declines by name and N keeps serving; an older-but-newer-than-N version that WOULD be compatible is not discoverable from the index as it stands.
The sync-owned hold. An attended reload lands the module even when the package's partition is owned by a sync source whose content has not caught up — the same as a Provision click. The item's
decisionsays what was landed; the content half follows the sync.The install record is not re-stamped. A reload writes the ACTIVATION record only;
Plugins/{id}is written by the installer and the package reconciler/sync, because it also describes the content install. After a Done live reload to 1.2.0 the install record still reads 1.1.0. The instance report therefore carries each package row's running version separately (runningVersion, read off the activation record against the loaded generation) — see DeploymentInventory → Installed version vs running version.A restart that never comes. A restart handed to the control lane that the control plane never executes leaves the request
AwaitingRestartwith the hand-over sentence on it — visible, not retried (it did not crash; it is waiting).The Plugins self-update intake only files and logs.
FleetTargetIntake(MeshWeaver.Plugins,Hosting/PlatformBuildInbox/Source/FleetTargetIntake.cs) never decides a request's fate from its status.RequestRefreshonly files a request; the node idselfupdate-{module}-{version}is the key, so a second filing collides. On a collision it reads the existing request and passes it toExistingReload, which returns only a log level and a sentence:Failed→ Error.Faulted→ Warning, naming the attempt and the fault ("core retries it on its reconcile pass"). This branch is added by MeshWeaver.Plugins#2893 (open when this was written); until #2893 lands,Faultedfalls into the Information branch below — a log line, no behaviour.- An undated or overdue
Requested, or no status at all → Warning. - Anything else → Information.
Neither method marks a request finished or red or files a replacement, so core's re-arm of the same node is undisturbed. With #2893, the state table is pinned by the intake's
AFiledReload_NamesItsState_FailedAndUnclaimedAreLoudtest (itsFaultedcase). Either way the intake never changes the request, so core's retry does not depend on #2893.