Platform and Module Deploy

"after platform we always deploy control" · "we separate platform deploy 100% from module deploy" · "control will always be on latest platform and will always update to latest modules."

Three policies cover this. The register is Policy Not Prose:

policy the rule
platform-deploy-control-first Every platform build that passes the control image's own acceptance becomes the control instance's next image, with no approval, in the run that built it. The fleet is offered the build only after control is running it and healthy.
platform-module-deploy-separate The platform deploy ships the host image only: the core platform plus the boot-time infrastructure modules. It does not pack, bake or seal any module, and no module verdict gates it. Modules publish on their own lanes, and every instance updates them on its own (packages-auto-update).
control-always-latest The control instance runs the newest platform build it was given and the newest compatible version of every module it has installed. If it falls behind on either for longer than a declared bound, an alarm names the lag.

The two deliveries

PLATFORM (core main-cd)                               MODULES (each module's own lane)
core merge ─► gate ─► images ─► promote               module merge ─► pack ─► test ─► publish
                 │                                    (MeshWeaver.Plugins ci.yml, satellites)
                 └─► control image ─► acceptance ─► control-promote        │
                                                    │                       ▼
                                                    ▼             registry: newest version + floor
                              control-first: memex-control:<version>        │
                                                    │                       │  every instance, on its
         control-first POSTs self-update-available                          │  own schedule:
            → routed Roll for control (no approval)                         │
                                                    │                       │  newer version published
                                   control runs it, /health 200             │  AND floor ≤ running
                                                    │                       │  ⇒ lands + activates
   arm (any later run): ladder green + control runs it                      │  (live, or one restart)
                                                    │                       │
                         memex-portal-ai:<version>, release event           │
                                                    ▼                       ▼
                         the fleet rolls the platform              the fleet updates modules

Neither side waits for the other. A platform roll keeps the module bytes an instance already runs. This is the ladder in policy platform-backwards-compatibility. A module update never waits for a platform build, seal, pair tag or framework identity.

What guards a platform roll now

The fleet's arming lives in arm-promoted-set.py, function judge. It reads three things, and none of them comes from a module lane:

  1. The platform's own tests. A set is promoted only from a commit whose required check is green. This is gate in main-cd.yml.
  2. The compatibility ladder (platform-ladder-compat). It takes the module set the fleet runs and links it, unchanged, against the promoted portal image, checking types, assembly versions and every member by signature. The bytes are the published ones: published-modules reads them from the sealed plugins publication for this run's identity, and never builds them. A platform change that breaks a module's declared compatibility turns this step red. The fix is to restore compatibility or declare an epoch bump, never to re-bake.
  3. Control first. The control instance named in .github/control-instance.json must be running a build that contains the set's core commit. Containment is checked by ancestry, using the GitHub compare API on its public /api/version. Control's /health must also answer 200.

How each state is read:

On the module side, two things still guard what an instance runs:

Control first, in practice

Who rolls control (and why it is CD, not control's own self-updater)

The first version of this design said "control's self-update rolls to it". It never did, and no build after ci.9939 reached control without a hand-filed, approved Roll. Because the fleet is armed only once control runs a build, nothing reached the fleet either. Three independent defects stood behind that, measured 2026-10-06 on control.systemorph.com and memex.systemorph.com:

  1. The watcher looked at the wrong repository. The in-pod self-updater lists SelfUpdate:PortalRepository, which defaults to memex-portal-ai, and nothing renders it for control. memex-portal-ai:<version> is written only by arm, and arm waits for control to run the build. That is a circular wait: control could only ever follow the fleet.
  2. It could not list tags at all. Control's Admin/UpdatePolicy read check FAILED: CredentialUnavailableException on every hourly check. A record-provisioned instance gets the federated credential (hosting-control on the portal identity), but the rendered values never carry selfUpdate.azureClientId, so the pod has no workload identity.
  3. Its hand-over went to its own inbox. Control lists Hosting/PlatformBuilds with a secret and declares no Hosting:ControlInbox:Url, so its route is local. Its own mesh holds no Deployments record (a namespace:Deployments search answers 0), so a hand-over there could open nothing. Ops/Actions on the control plane held selfupdate-roll-* actions for memex, build and memex-cloud, and none for control.

Defects 1 and 2 are now fixed at the source (Systemorph/MeshWeaver.Plugins#2994), and reach the pod only with a Reconcile (below). Defect 3 needs no fix of its own while CD's hand-over goes to a plane that holds control's record — but see A frozen fleet is red: once CD's hand-over was pointed at control's own inbox, defect 3 became the whole outage.

The job that tags the control image is the one place that knows, deterministically, that control has a new build. So that job hands the build over. The record still decides everything that matters: which record is named, whether the instance takes rolls, the image repository, the line, the gate, and never-backwards.

The alarm (control-always-latest)

Platform half. The core workflow .github/workflows/control-always-latest.yml runs every 30 minutes and calls arm-promoted-set.py control-lag. It finds the newest promoted set whose Deploy control first job succeeded, and checks whether control's running commit contains that set. When it does not, it walks back through the earlier builds given to control until it reaches one control contains. The first build control was given and did not take starts the clock. The results:

🚨 The clock used to start at the newest build. With a platform build every 30–60 minutes that clock was reset on almost every tick. On 2026-10-06 control had been on ci.9939 for 22 hours while the alarm read "converging — 3.0.0-ci.10052 was given 29 min ago". A bound that every new build resets cannot catch the one state it exists for: control taking no build while builds keep arriving. When the walk reaches the end of the examined runs, or a containment it cannot read, the reading is a floor and says "at least".

Module half. This half runs in MeshWeaver.Plugins as FleetTarget.ControlBreaches (Hosting). It reads the control instance's own module inventory report. A module whose served version is newer than the installed one is a breach once it has been behind for longer than the bound. This applies when the served version's floor is at or below control's running platform, and it does not matter whether a hold reason exists. A module held because its floor is above control's platform is the floor working as designed, not lag. Each breach is filed through the existing FleetTargetIntake as a fleet-behind-target triage item that names the module, both versions and the hours behind.

A frozen fleet is red

Policy control-first-never-silent (register). Control-first makes the whole fleet wait for ONE instance. So a control that takes no build must be a failure on every surface where it is read. It must never be a green wait.

What happened (measured 2026-10-07/08). vars.CONTROL_WEBHOOK_URL was moved to control's own inbox (https://control.systemorph.com/api/hooks/Hosting/PlatformBuilds) at 18:04Z. The previous target, memex.systemorph.com, had answered 404 at 18:03Z, and core #6275 then made preflight require the declared control URL. From ci.10184 (run 37666055296) on, every hand-over answered 200 {"status":"accepted","signature":"verified"}. Control's own mesh holds no Deployments record, because cut-over steps 4–6 are unfinished: on control, namespace:Deployments and namespace:Ops/Actions both answer count: 0. The router therefore dropped every announcement (SelfUpdateRouting: "the record lookup IS the authorization"). A hand-filed Roll of control on memex.systemorph.com (Ops/Actions/roll-control-20261007-10184-handoff) and the identity Reconcile both failed at Launch operator job: that plane's operator is disabled. Control stayed on cac0664de, and control's own self-updater read check FAILED: CredentialUnavailableException … The requested identity has not been assigned to this resource on every check (Plugins#2994). memex and memex-cloud stayed on 57a6e5fd.

Meanwhile:

What now reads it as a failure:

  1. The arming goes RED past the bound. arm-promoted-set.py select (arming_frozen) fails the arm job when all of these hold:

    • nothing was selected and no override was given;
    • control's running build is readable;
    • that build does NOT contain the newest build control was given;
    • control has been behind for longer than platformLagBoundMinutes.

    The bound and the clock (the first build control missed) are the same as the alarm's, through the one control_lag rule, so the two can never disagree about when waiting stopped being a wait. alert-on-failure then files the run, with an arm line saying the fleet is frozen on control. The walk goes back through the whole examined history to the first build control contains, never a fixed window: a window of the newest ten would lose the first miss as soon as control missed more than ten. Behind past the bound freezes whatever /health says in the same reading, because the lag is what is clocked. Three readings have no clock behind them, so they stay the alarm's job and a single 503 never turns CD red:

    • an unreadable control;
    • a control that runs the newest build but answered one non-200;
    • a run whose own Deploy control first is still running. The remedy is control's roll. An arm override stays the maintainer's call, and nothing suggests one.
  2. Every sentence names control's own verdict. The arming's waiting line, its RED error and the control-lag issue body each quote control's self_update line off its public /health. When control publishes none, they say so and point at Admin/UpdatePolicy.lastCheckVerdict. An absent reading is never read as a clean one.

  3. A failing self-update is public. The one classification SelfUpdateVerdict.IsFailure covers a faulted check, a release that could not be applied or handed over, a refused migration, a stranded tag and a module that can never be activated. It feeds three things:

    • the Warning log level;
    • Admin/UpdatePolicy.lastCheckFailed, which the fleet console can flag without parsing a sentence;
    • the census-tagged self_update entry on /health (SelfUpdateHealthCheck over SelfUpdateCheckCensus).

    The entry reads self_update: Degraded — <first line of the verdict> [outcome …, trigger …, N min ago]. It is Degraded, never Unhealthy, and it carries no probe tag: a failing self-update costs delivery, and pulling the pod delivers nothing.

  4. The hand-over says only what a 200 proves. Deploy control first now writes "stored and signature-verified … that is not yet a roll". It no longer claims that a Roll opened.

Can control fix its own identity without a Reconcile? No. The fix (core #6227) changes the pod's ServiceAccount annotation, the azure.workload.identity/use label and AZURE_CLIENT_ID, and renders SelfUpdate__PortalRepository. All of these are helm-rendered. Only hosting-deploy applies them, and it runs as the Reconcile (or Provision) step "Re-apply the record". Neither image path can carry them:

A newer image on the old pod spec is still a pod with no workload identity.

The one-time remedy. These are maintainer acts, and none is taken by CD or by an agent:

  1. Finish cut-over steps 4–6, so the control plane that receives CD's hand-over holds Deployments/control and runs an enabled operator.
  2. Run one governed Reconcile of Deployments/control there.

From then on, CD's hand-over routes a Roll for every build, and control's own self-update check lists memex-control under its own identity. Until then the arming is red on every run, naming the reason.

What was measured before the change

The measurements were taken on the live system. They are evidence and are not to be updated.

Coupling points that were in main-cd.yml:

job / step what it coupled
gate → "Is plugins sealed for the identity this set resolves?" (plugins_seal_due) a platform reconcile re-attempted a MeshWeaver.Plugins seal
plugins-bake-image resolved digests for the Plugins bake
plugins-modules ("Plugins: pack the module bundles the bake composes") packed AI, Markdown.Collaboration, Maps and Payments.Stripe from Plugins source in core CD
plugins-bake ("bake + seal the publication for this identity") re-sealed the plugins publication on every platform build, although the identity (c003e001) is per epoch, not per build
report-plugins-seal + the cd-plugins-seal ledger judged the above
arm → "Mint a MeshWeaver.Plugins token" + select the fleet was armed only on a green Plugins dependent-suites verdict for the exact pair <core7>-p<plugins7>
control-arm memex-control:<version> followed the fleet's arming, so control received a build only after Plugins' verdict, which made control last
satellite-compat, ladder baseline read the module bundles that core CD packed
Memex control.json rollGate: {after: [memex-cloud, memex], soakMinutes: 120, approval: required} control rolled last, with an approval

How long a platform roll waited on module work. The figures come from CD run 9965 (260b3c4), which promoted at 07:12Z:

Image composition. Measured on memex-control staging-bd3ea73 (linux-x64, 10 layers, 1.20 GB compressed). The app layer is 472.6 MB uncompressed, of which modules/ is 251.8 MB:

module MB kind
MeshWeaver.Hosting.Cosmos 39.1 boot-time (storage backend)
MeshWeaver.Fleet.Control 37.8 boot-time (control only)
MeshWeaver.Hosting.Snowflake 25.7 boot-time (storage backend)
MeshWeaver.Hosting.Instance 14.5 boot-time (gates the registry it would arrive through)
MeshWeaver.SelfUpdate.Aks 11.8 boot-time (control only: rolls itself)
MeshWeaver.Blazor.Chat 23.5 not boot-time
MeshWeaver.Mcp 22.5 not boot-time
MeshWeaver.AI 22.2 not boot-time
MeshWeaver.Markdown.Export 14.0 not boot-time
MeshWeaver.Markdown.Collaboration 13.5 not boot-time
MeshWeaver.Blazor.Graph 13.2 not boot-time
MeshWeaver.Blazor.EntityViews 13.1 not boot-time

That makes 122.0 MB of the image module seeds that are not boot-time. The registry already replaces each seed in place when it serves a newer version.

Migration with no unguarded moment

Every step below keeps the platform roll guarded, and the steps run in this order:

  1. Core: control first, and the platform verdict. This is in this change.

    • control-first tags control on every accepted build.
    • arm judges ladder plus control instead of the Plugins verdict.
    • published-modules reads the sealed publication in place of plugins-modules.
    • plugins-bake, report-plugins-seal and the gate's seal probe are removed.
    • control-always-latest.yml starts alarming.

    No moment is unguarded. The ladder already ran on every build, so it has been part of the verdict from the first run. The control half can only make the arm stricter: until control rolls, the fleet waits and the alarm says so. Because the published-modules artifacts keep the module-bundle-<Module> names, the satellite baselines and every core pull request's ladder (fetch-deployed-plugin-set.sh) read them unchanged.

  2. Memex: the control record goes first. Drop control's rollGate and declare moduleUpdatePolicy: Auto. Until this lands, control-first tags are rolled only with an approval, so the fleet waits and control-always-latest stays red, naming the lag. That is loud, not unguarded.

  3. MeshWeaver.Plugins: the module half of the alarm (FleetTarget.ControlBreaches, the carried BehindSince clock, and FleetTargetIntake.AllBreaches).

  4. In flight — removing the non-boot seeds from the image (MeshWeaver.Plugins#2970, a draft). Until it merges, the published image still seeds all seven non-boot modules. Plugins #2893 (the reload_module/uninstall_package tools and the ModuleReload intake) has merged. Two decisions still hold the pull request: how a memex-local self-registry install, which has no registry to land the seven from, gets them; and confirming that the registry instance (memex.meshweaver.cloud) lands its own required modules from its catalog. That pull request makes Memex.Portal.Distributed carry only the modules declared in src/Memex.Portal.Distributed/image-boot-modules.txt, each with its reason:

    • both images: Hosting.Instance, Hosting.Cosmos and Hosting.Snowflake;
    • control only: Fleet.Control and SelfUpdate.Aks.

    Its guard, ImageBootModulesTest, holds the host's <MeshModuleClosure> rows equal to that declaration for both images, with a negative control for each of these mutations:

    • a registry module seeded back into the image;
    • a declared module with no row;
    • a control module leaking into the portal image;
    • a row under an unknown condition.

    The seven modules that leave become store-delivered: no baseline Modules:Assemblies entry, and all seven under Modules:Required. A baseline entry for bytes the image does not carry would make required_modules read "the image is supposed to ship it" (Unhealthy) and hold readiness on the registry. Without one, a module that has not landed is named as Degraded, the shape Radzen, Analysis and GoogleMaps already have. The DEV portal (Memex.Portal.Monolith) keeps its seeds. The manual, MeshWeaver.Plugins Hosting/ImageSeededModules.md, gains the boot-only section in the same pull request.

  5. Done — the pair tag <core7>-p<plugins7> is retired. main-cd no longer mints it on memex-portal-ai (promote phase A) or memex-control (control-promote). See "The pair tag, retired" below.

  6. Done — the promotion poller is retired. MeshWeaver.Plugins deleted promotion-candidate.yml (the 10-minute poll that measured every promoted pair), the green-pair nudge of core's CD and the held-pair ledger (core-release-attribution.py, the core-release-held issue). Core then deleted the pending command of arm-promoted-set.py that answered the poller, and the two core-candidate.yml rows in .github/lane-caller-grants.yml. Plugins went first, because a satellite asserts its roster row against core's main, and a pending: row excuses an absent caller but not an unrecorded one. core-candidate.yml stays for the advisory measurements a core pull request asks for (dependent-suites label or Pairs-with:) and for paired-core.yml. None of them gates a merge or an arming.

The pair tag, retired

<core7>-p<plugins7> named the MeshWeaver.Plugins commit a portal image's HOST was built from. After the separation it answered no delivery question, but three mechanisms still leaned on it. Each now reads something else:

What leaned on it What it reads now
gate's completeness probe: a set was "stale" when Plugins main had moved at all, so the portal was rebuilt on every Plugins merge (MeshWeaver#4688) Core's sha only. A host change reaches the image through MeshWeaver.Plugins' portal-image-rebuild.yml, which classifies each push (scripts/portal-image-relevance.py) and dispatches main-cd with rebuild: true only when a changed path can enter the image
arm's source for memex-portal-ai:<version>: the one tag that names THIS build after a rebuild of the same core commit moves the bare <core7> tag The build's staging tag staging-<core7>-<run id>, recorded as staging in the promotion record. It is unique per run, and arm refuses a staging tag that does not name the selected core
Recovering the core commit of the newest ARMED manifest (arm-promoted-set.py armed-base), resolving an unarmed set (resolve-platform.py), and a tagged release (release.yml) The bare <core7> tag or the staging tag on the same manifest. A legacy pair tag is still read, so sets promoted before the retirement resolve unchanged

The host commit is still recorded: the promotion record keeps plugins_sha, and the release event carries pluginsSha. Readers outside core (Memex scripts/image-contains.py) answer from the promotion record first and treat a pair tag as a legacy fallback.

Consumer sweep before the retirement (repositories at origin/main, the live mesh read through the control instance):

Where Readers of a pair tag
core main-cd.yml 4 writers/readers: promote phase A, control-promote, arm's source tag, gate's probe (via check-image-set.sh)
core scripts check-image-set.sh, arm-promoted-set.py (armed_commit), resolve-platform.py (portal identity), release.yml, lock-pinned-digests.py (protects <core7>-p* as part of a commit's closure; kept for legacy manifests)
MeshWeaver.Plugins 0 code readers. portal-image-rebuild.yml and portal-image-relevance.py mention it in comments; promotion-candidate.yml, core-candidate.yml and core-release-attribution.py use the record's verdict KEY pair-<core7>-p<plugins7>, which is a record field, not a registry tag
Memex 1 code reader: scripts/image-contains.py. Comments and historical notes in helm-release.yml, two values.*.public.yaml files and mesh/Deployments/memex.json
Education, Reinsurance, SocialMedia, Manufacturing, Crm 0
live Deployments/* on the control instance (5 Hosting/Deployment records) 0 pins; 1 prose mention in Deployments/memex

No Hosting/Deployment record pins an image tag of either shape, and none may.

Tests and self-tests

What is not established