Staged Pull Request Pipeline

๐Ÿ”€ Suites beside the review, the arm after both โ€” policy suites-parallel-with-review, the fleet-wide default again. The stage gate lets the suites start once stage 0 is green; they test the head merged onto the CURRENT main (Fresh Merge Under Test); and auto-merge is armed only once the review is answered AND every required check is green. Holding the suites until the review landed (review-then-suites, now retired) is the per-caller opt-in review-before-suites: true; the stage descriptions below read for that opt-in, and the history is in A day in parallel.

Historical decision, review-first (maintainer, 2026-10-04, verbatim; since released โ€” default is above): "should we maybe say that code review must pass and also other controls such as no client etc. must pass before we start test. and we arm only at end of test".

A pull request HEAD moves through four stages. With the opt-in review-before-suites: true they run in order, each starting only when the one before it is green for that head; on the default, stage 2 starts after stage 0 and stage 1 runs beside it, with the arm waiting for both:

Stage What runs Cost What moves the head on
0 โ€” static controls shape/validate, confidential terms ("no client names"), AGENTS.md shared-rule blocks, generated files and locks, repo policy gates, CI inputs (core: shared rules, closing keywords, package pins, interface additions, i18n mirror, CI shell, cross-repo pair) minutes, one light runner each they finish; a red one fails fast โ€” nothing heavy starts
1 โ€” automatic review the pull request has its review โ€” a landed Copilot review against the head or any earlier head (policies copilot-code-review, review-once-per-pull-request), or a completed internal-review (App systemorph-com) โ€” and every thread a reviewer opened has a reply from a person the reviewer's time, no runner Copilot: its review event starts no workflow, so the scheduled stage-advance sweep moves the head on (up to ~15 minutes later); internal: check_run: completed of the review; either way a person's reply โ€” see the event half
2 โ€” expensive suites core: build + test shards, doc gate, platform-compat, the dependent-suites request ยท Plugins: module bundles, compile-check, gate shards, portal hosts (every leg behind admission) the run's runner-minutes, almost all of them the suites finish
3 โ€” arming the control plane's babysitter arms auto-merge (Plugins PrArming, #2828) none the arm gate (check-review-answered.py โ†’ arm_readiness): the review answered (conditions 1โ€“3) AND every required status check of the base success on the head (condition 4, required_checks_green)

On the opt-in, the order is the point. Stage 1 is the stage most likely to send the author back to the keyboard, so it runs before the stage that costs the most. A finding answered with a fix push makes a NEW head, and a new head restarts at stage 0 โ€” the suites never ran on the head the review changed.

A day in parallel, and why it was reverted

On 2026-10-05 the suites briefly ran IN PARALLEL with the review (suites-parallel-with-review: the stage gate held on stage 0 only), because with ~50 reviews in flight the runners sat idle (09:00Z: core 31 jobs running / 0 waiting, Plugins 30 / 0). The maintainer corrected it the same day: the fleet-wide order is review first โ†’ suites against a fresh merge with the current main โ†’ arm only when the review is answered and the suites are green (policy review-then-suites); the maintainer later released that restriction and suites-parallel-with-review is the default again. What held across both:

Arming after green. Condition 4 of the arm gate (required_checks_green) refuses while any required status check of the base branch is missing, pending or red on the current head โ€” read from BOTH rulesets and classic protection, the review's own contexts excluded (they are conditions 2โ€“3), and an empty required set refused rather than read as "all green". MeshWeaver.Plugins' PrArming ports the same condition.

Measured โ€” why the review-first ordering saves runners (and when it does not)

Sample (REST only, 2026-10-04; bot and lock-settle pull requests excluded): the newest 40 merged pull requests per repository โ€” core 2026-10-02 09:44Z โ†’ 10-04 07:05Z (130 CI heads), Plugins 10-03 17:06Z โ†’ 10-04 03:30Z (163 CI heads); SocialMedia and Reinsurance 15 each (09-19 โ†’ 10-03). Review latency = the head's first CI run created โ†’ its first internal-review completed (success or failure). Runner-minutes = the sum of every job's duration, all attempts.

Core (dotnet-test.yml) Plugins (ci.yml)
Review latency, median / p75 / p90 16.6 / 52.3 / 118.9 min (n=49) 31.9 / 91.7 / 116.1 min (n=89)
PR CI wall time, median / p75 (all completed runs) 18.5 / 20.9 min (n=134) 59 / 92 min (n=176; successful runs 93 / 108)
Runner-minutes per run, median / p75 60 / 62 90 / 172 (successful runs 122 / 178)
Test run finished after the review, median / p75 +1.7 / +7.9 min โ€” they finish together +14 / +47 min โ€” the review usually lands first
Reviewed heads with โ‰ฅ 1 finding 46 / 49 (94 %) 79 / 89 (89 %)
Reviewed heads whose review was blocking (failure) 9 / 49 (18 %) 14 / 89 (16 %)
Reviewed heads followed by a new push after the review landed 33 / 49 (67 %) โ€” 31 after findings, all 9 blocking ones 57 / 89 (64 %) โ€” 49 after findings, all 14 blocking ones
Runner-minutes spent testing heads that a findings review then replaced 2,037 of 8,321 (24 %) 4,547 of 18,705 (24 %)
โ€ฆof which after a BLOCKING review 560 (7 %) 940 (5 %)
Heads superseded by any later push (any reason) 90 / 130 (69 %), 66 % of runner-minutes 123 / 163 (75 %), 68 % of runner-minutes
"Reviewer unavailable" degradations in the sample 9 (2026-10-03 14:11Z โ†’ 20:16Z) 18 (10-03 13:19Z โ†’ 20:16Z, ~7 h) + 1 "agent not found"

Satellites are small and recently reviewed (the internal review reached them ~09-28): SocialMedia 7 reviewed heads, 0 re-pushes; Reinsurance 9, 4 re-pushes (34 runner-minutes). Their CI is 10โ€“16 runner-minutes per run, so the saving there is small and adoption is not urgent.

What it saves. Up to ~24 % of pull-request runner-minutes in both repositories โ€” the tests run on heads that a review then sent back for a fix push (2,037 min in core and 4,547 min in Plugins over the sample; the sampled pull requests were opened from 09-27, so the sample does not convert honestly into a per-day rate). That is the upper bound: it counts every push that FOLLOWED a findings review, and some of those pushes were merges of main or work the author would have pushed anyway; the blocking-review subset (7 % / 5 %) is the floor. Stage 0's fail-fast saving comes on top and was not measured separately.

What it costs. Time to a green wall grows by the review latency where the two used to overlap: in core the tests and the review finished together, so a head now waits a median ~17 min longer (p90 ~2 h when the reviewer is queued); in Plugins the review usually landed first, so the added wait is the median ~32 min review latency minus the overlap that already existed. Time to MERGE moves less than that, because merging already waited for the review AND every answer (Automatic review answered, the arm gate). The answer time is now on the critical path too โ€” by the maintainer's decision: a review must PASS, answers included, before the tests start.

What the numbers also say.

Stage 1's predicate โ€” one implementation, three questions

Stage 1 asks the same question the arm gate and the merge gate ask, at an earlier moment. check-review-answered.py holds all three, and stage_readiness reuses the arm gate's internal_review_runs, listing_incomplete and reviewer_threads verbatim, so the three can never disagree about what "reviewed" or "answered" means:

Merge gate (Automatic review answered) Stage gate (--stage-gate) Arm gate (--arm-gate, control)
Asks may this merge? may stage 2 start for THIS head? may auto-merge be armed now?
Review must be on any head of the PR the current head โ€” or carried over a clean base merge the current head โ€” or carried over a clean base merge
Reviewer unavailable (degradation) releases releases, loudly refuses โ€” a person merges
No review after the fallback red releases, loudly refuses
tests-before-review label โ€” releases, loudly ignored
Unanswered thread red holds โ€” never falls back refuses
Draft โ€” holds refuses

An unanswered thread never falls back, because that wait belongs to a person, not to the infrastructure. Every RELEASE is loud: a ::warning:: titled "Stage 2 released without a completed review", a line in the job summary, and โ€” because none of them arms โ€” a head released that way still needs its real review before stage 3.

What moves a head on โ€” events, not polling

A hold is the stage gate failing. Something must ask again when the answer changes, and the doctrine already has the shape: review-answered-on-degradation.yml (core) listens for a check run and re-runs the read-eligible pull_request run instead of judging from a default-branch run (whose check runs would land on main's commit, where branch protection never reads them). The staged pipeline generalises it:

Trigger What it means
check_run: completed (job filter: name internal-review, App id 4918443) the review landed โ€” or degraded. The happy path.
pull_request_review_comment: created, pull_request_review: submitted a person answered a finding
schedule every 15 minutes the bounded fallback timer โ€” a review OUTAGE raises no event at all

rerun-failed-jobs re-runs the gate, the front door it held (admission in Plugins, collect-results in core) and every job that skipped behind them, on the SAME head; the green stage-0 jobs are not repeated. Runs created by GitHub Actions raise no check_run workflow events, so there is no loop.

Why not the alternatives. workflow_run cannot see a check run an App posts. A repository_dispatch from the control plane would make stage 2 depend on the control plane being up โ€” exactly the component whose outage the fallback must survive โ€” and would need a second credential path. The re-run uses only the repository's own GITHUB_TOKEN (actions: write) and the remedy the doctrine has already sanctioned.

The three loud releases โ€” an infrastructure blocker never freezes the fleet

Blocker What happens Arming
Reviewer unavailable โ€” the steward's neutral internal-review "Reviewer unavailable โ€ฆ" (posted only for an infrastructure cause, after a re-kick) its check_run: completed re-runs the held run; stage 2 starts at once, warning still refused; a person merges and a post-merge review is owed (unchanged contract)
Review outage โ€” nothing posted at all (2026-10-04: the FlattenMarkdown fault; on the ten newest open Plugins pull requests that morning no head carried any internal-review run) after fallback-minutes (default 120 โ€” the measured p90 review latency, below) from the head's run creation with no completed review, the next sweep re-runs it; stage 2 starts, warning "REVIEW UNAVAILABLE โ€ฆ that is the incident to chase" still refused until the real review lands and is answered
tests-before-review label (a draft that wants test feedback, an author who prefers it) stage 2 starts immediately, warning unaffected

The fallback is measured from the run's created_at, which a re-run keeps, so re-evaluation never restarts the clock. The sweep fires every 15 minutes, so the effective release lands between 120 and ~135 minutes (plus GitHub's own schedule jitter) โ€” bounded, and every minute of it loud in the gate's log.

Why 120 and not less. The fallback exists for an OUTAGE, not for a slow review: the measured p75 review latency is 52 min (core) and 92 min (Plugins), so a 60-minute fallback would have released a quarter of all Plugins heads into stage 2 before their review โ€” exactly the spend the staging removes. 120 sits at the measured p90 (119 / 116 min). It costs little during an outage, because arming waits for the review anyway: the fallback buys an early test READING, never an earlier merge. A caller may pass fallback-minutes to both lanes (they must agree), and tests-before-review is the per-PR override.

Generated-only pull requests owe no review. main's own jobs propose generated files as pull requests โ€” settle-locks (every manifest.lock) and stamp-floors (mesh-floor.lock plus each root's minMeshVersion). There is nothing to review, and the settle job rewrites the head on every main merge, so a per-head fallback clock restarted forever: during the 2026-10-04 review outage the held settle PR (Plugins #2860) stopped module publishing. Such a pull request skips stage 1 on PROVENANCE (generated_only): authored by the App that writes them (meshweaver-cloud[bot], by id), every commit that App's (the commit listing must cover the pull request's own commit count), every changed file a lock or a root index.json whose changed lines are each NOTHING BUT a "minMeshVersion": "โ€ฆ" key/value. A person's pull request that touches a lock is staged like any other, a generated-only DRAFT is held as a draft, and for any pull request the App authored the fallback clock keys on the PULL REQUEST's creation, not the head.

The required verdict agrees (Plugins #3044). Automatic review answered now consults the same generated_only rule. For a generated-only App pull request, condition 1 ("the review landed") is NOT OWED once a TERMINAL signal says that no review is coming. The log then says NOT OWED: โ€ฆ โ€” nothing to review (generated_only). There are two such signals:

It is never released while the reviewer could still answer: a pass on the opened evaluation could merge before a late thread arrives, and a thread that opens after the merge can block nothing. Every thread the reviewer did open still needs a person's reply, and a draft is still held. Until this, the reviewer REFUSED every lock-only pull request ("Copilot wasn't able to review any files in this pull request" โ€” locks are in its default exclusions), so the verdict was RED โ€” UNREVIEWABLE and only a person's review-waived label let a settle or a floor stamp merge. Measured 2026-10-08 on Plugins: settle PR #3165 sat red from 13:38Z; every main run in between read NOT settled (n lock(s) would move) and published, sealed and tagged NOTHING, so between 07:16Z and the next waiver no module reached the registry โ€” while every portal's sync installed each newer source from HEAD and declined the older prebuilt, holding Governance/Activity and Governance/Standard at StaleAdopted 0.9.16 over installed 0.10 (#3044, the MeshWeaver#3583 delivery hold).

The provenance it rests on reads GitHub's resolved commit author, which follows the author email โ€” so it is as strong as the App's branch is closed to other pushers (the App's commits are unsigned today, so a signature cannot be required until the producing jobs commit through the API). What it admits is bounded by the file rule: only locks and minMeshVersion lines, whose CONTENT is held exact by Validate node repos and the settle job's own --settle check, and every required test context still has to pass.

Stage-0 blockers keep their own rules. A control whose INPUT is absent (the confidential-terms denylist secret on a fork or a Dependabot run) already skips with a notice by its own design and so does not hold stage 2. A control that goes RED on infrastructure (an unreadable API, a lost runner) is a red like any other: retry-known-transients.yml re-runs the recognised shapes, and the head waits in stage 0 โ€” because "the static controls could not run" must never read as "the static controls passed".

A push during stage 2

Unchanged supersession rules, applied per head:

A clean merge of the base branch carries the review

Policy review-carries-over-clean-base-merge. Measured 2026-10-05 22:38Z: 25 open non-draft pull requests across core, Plugins and Memex, and 7 merges in 2.5 hours โ€” six of the 25 DIRTY at the same moment. The loop that kept them there: main moves โ†’ a pull request goes DIRTY โ†’ main is merged into the branch โ†’ the new head has no internal-review โ†’ this gate holds it at stage 1 โ†’ a full review round queues behind admission โ†’ main moves again. A merge of the base branch that applied cleanly changes nothing the reviewer read, so it must not cost a review round.

The rule. The review of an earlier head carries over to the current head when ALL of these hold, each computed from the repository through REST โ€” never from a commit message or a branch name (check-review-answered.py find_reviewed_ancestor + carry_over, read by read_carry):

  1. Walking back from the head through MERGE commits only reaches a commit whose newest internal-review run from the reviewer's App is a review โ€” its title is on an ALLOW-list (No blocking findings, N blocking finding(s), Review carried from โ€ฆ), so the Reviewer unavailable degradation, the steward's Review not completed exit, and any title Plugins rewords later are NOT carried: the list fails safe, a fresh round. Every merge walked has exactly ONE parent on the pull request's side (the other is in the base branch), so a merge of another feature branch does not carry.
  2. Every commit on the pull request now that was not on it at the reviewed head is a merge. One non-merge commit is new content and owes a fresh review.
  3. The pull request's OWN diff is byte-identical: compare(base...head) and compare(base...reviewed head) โ€” each merge-base..head, as GitHub computes it โ€” have the same patch-id. The patch-id hashes every file's status, names and patch with each hunk header reduced to @@ (line numbers move when main changes elsewhere in the file; the change does not); every other byte is hashed โ€” context lines and trailing whitespace included (two trailing spaces are a Markdown line break). A file sent without a patch (binary, too large) contributes its blob id.

Anything unreadable does NOT carry: a short commit listing, the compare API's 300-file cap, a file with neither patch nor blob id, a failed read. The head is then reviewed fresh, exactly as before.

What carries. The reviewed head's run stands in for the current head's in the STAGE gate (mode carried โ€” stage 2 starts at once) and the ARM gate; the pull request's threads are pull-request-wide already, so a thread answered on the earlier head stays answered. The head's own real review, when it has one, always wins; a carry replaces only an absent, running or degraded own run. The merge gate (Automatic review answered) never asked per head and is unchanged.

What is logged. The stage gate's job summary and the arm gate's line name the reviewed head, its check run and conclusion, the merges since it, and the patch-id with both merge-base ranges:

โœ… Stage 1 green for #N by CARRY-OVER โ€” no new review round: review CARRIED from head d4815d432e (internal-review check run 111882997610: success "No blocking findings") to 4aff97fd03: the 1 commit(s) since it are all merges of the base branch (4aff97fd03), and the pull request's own diff is byte-identical โ€” patch-id โ€ฆ over N file(s), <merge-base>..d4815d432e = <merge-base>..4aff97fd03

A refused carry is logged too (no carry-over from reviewed head โ€ฆ: the pull request's own diff changed โ€” patch-id X (โ€ฆ) vs Y (โ€ฆ)), so a head that waits says why it waits.

No round is spent. The PR steward (MeshWeaver.Plugins Hosting/Deployment/Source/ReviewCarryOver) applies the same rule before it claims a review slot: a carried head gets an internal-review run posted with the reviewed head's conclusion, titled Review carried from <sha>, and no reviewer thread โ€” the required internal-review context of a plugin repository is satisfied by the evidence, not by a model.

Measured on the history (replaying every reviewed-head โ†’ merge-of-main pair on the last 60 core pull requests, base = the merged main commit): 7 of 11 carry. The 4 that do not are each a change to the pull request's diff โ€” a conflict resolution, or main editing a line inside a hunk's context (the policy register's last row, which every policy-adding pull request appends after: the added row is unchanged, its neighbour is not; strict by design). The same replay showed why the walk refuses Review not completed: three further clean merges sat on that exit, which is not a review.

Negative control. --no-carry (stage gate and arm gate) disables the carry; the self-test runs the pure-merge fixture both ways โ€” carried โ†’ stage 2, disabled โ†’ waiting โ€” and a mutation of each condition (patch content ignored, non-merge commits allowed, hunk headers kept, either gate ignoring the carry) turns the self-test red.

Required contexts stay satisfiable โ€” a hold is RED, never skipped

A skipped required context counts as satisfied, so the hold is a failure at every level:

Every hold's error MESSAGE starts with STAGE 2 HELD (stage 1, <mode>) or STAGE 2 HELD (stage 0 red) โ€” the stage gate, Plugins' admission and core's Consolidate test results all print it โ€” so a hold is recognisable from a check's annotations alone. The control plane's PR babysitter keys on it (PrBabysitter.IsStageHold): a stage-1 hold with no open thread is the class waiting-for-review (never sent to the author, never interrupted, never handed to the fixer, never re-run from there โ€” the event half owns it), with open threads it is review-unanswered, and whatever is red INDEPENDENTLY of the hold (the stage-0 control itself, a leg that does not wait for the front door) is classified exactly as before. Without the marker, every hold would read as a real defect on the front door โ€” and the fixer would cancel the very run the review is about to release.

Not staged at all (the gate answers not-staged in green): a fork's pull request (the internal reviewer does not review forks, so a hold would only wait out the fallback under a warning blaming an outage โ€” the merge gate still applies to it), and core's green-tree reuse path (it spends nothing).

push, schedule, workflow_dispatch and merge_group are never staged: the gate answers not-staged in green at once. Trunk never waits for a review, and core's merge queue re-tests a PR whose review was already required to merge.

Stage 3 โ€” arming waits for stage 2

Arming moves to the control plane (Plugins #2828, PrArming), and core #6063 makes auto-arm.yml disarm-only. The arm predicate gains "every required check is green on the head" in front of arm_readiness, so the order is enforced at both ends: nothing heavy starts before the review (opt-in only), and nothing arms before the tests. A head released by any of the three loud releases is never armed by that release.

Under policy copilot-code-review, PrArming is to take the same review as arm_readiness: a landed Copilot review of the pull request, on the current head or an earlier one (copilot_review_of_pull_request, policy review-once-per-pull-request) first, then a real internal-review run. That port is owed (Systemorph/MeshWeaver.Plugins#3088): until it is merged and deployed on the control instance, PrArming still requires an internal-review run, and with the internal reviewer retired it arms nothing, so a Copilot-reviewed head is not armed by the control plane yet.

Adoption, per repository

Repository Stage gate Event half State
Systemorph/MeshWeaver dotnet-test.yml โ†’ stage-gate (local lane, scripts-ref: github.sha) stage-advance.yml default: suites beside the review
Systemorph/MeshWeaver.Plugins ci.yml โ†’ stage-gate โ†’ admission stage-advance.yml default: suites beside the review (unless its caller opts in)
the satellites not yet โ€” they run as before โ€” adopt with the same two edits: a stage-gate job (node-repo-stage-gate.yml@main) that their heavy jobs need with result == 'success', and the thin stage-advance.yml; plus their rows in .github/lane-caller-grants.yml

A repository that has the gate but not the listener would hold forever, which is why the two are adopted together and why the Plugins pin (check-build-queue-admission.py) names the lane.

Trunk outranks pull requests โ€” through the CI queue, not a lane

Staging stops a pull request from STARTING expensive work too early. It does not decide who gets a runner when the pools are full. That is the CI queue's job, under policy ci-two-pools-priority-queues (Runner Pools and Dispatch Queues). There are two shared pools, and runs are dispatched in tier order: express, trunk, gate, pr.

๐Ÿ—„๏ธ Superseded: the trunk lanes. On 2026-10-04, with both ARC sets at their cap and a label's queue first-come-first-served, a main run's worst queued job waited like a pull request's (median 1.9 vs 1.4 min, p90 4.7 vs 4.4 min over 19 main and 133 PR runs). The first answer was two reserved lanes, aks-silos-trunk (8) and aks-silos-dind-trunk (12), with a higher PriorityClass. Measured overnight after they went live, trunk jobs on them waited p90 6.6 / 3.8 min against 0.6 / 1.0 min for pull requests on the shared sets: the reservation was too small for its own work. The lanes are retired, and every runner gate of the fleet refuses their names.

What it does not do (residue, stated)

Reading a held pull request

  1. The red front door (Consolidate test results / Admitted by the build queue) names the stage.
  2. Open Stage gate: may the suites start (formerly Stage 1: review landed and answered) โ€” it says waiting (with minutes of the fallback used), unanswered (with the first thread), draft, or unreadable.
  3. Do nothing for waiting; reply to each thread for unanswered. The suites start by themselves.
  4. If the review is down and you cannot wait for the fallback: tests-before-review. It buys the test reading early; it does not buy the merge.