The Merge Queue

๐Ÿšจ THE QUEUE IS OFF on main. Measured 2026-09-27: enqueuePullRequest answers "No merge queue found for branch 'main'", and the ruleset main pr protection (id 2128472) carries no merge_queue rule โ€” main merges on its required checks plus auto-merge. The same day the dependent-suites wait left the merge path (policy core-merge-never-blocked), and it came back on the PULL REQUEST, bounded to what the diff can reach (policy dependent-suites-affected-gate, Cross-Repo Pair Gate) โ€” a core merge waits on exactly that one verdict. What follows is the queue's design and history, kept for the day it is re-enabled; the steward acts only on dequeued events, so it is dormant while no queue exists.

A merge queue builds the combination that is about to land, before it lands. With strict: false branch protection every pull request is tested against the main it branched from, so a burst of merges lands a tree no run ever compiled. That happened twice in one week โ€” CS0246: MeshOperations on 2026-08-26, and on 2026-08-30 a new guard (#2782) plus an independently-landed gate primitive (#2792) that were green alone and red together, holding main red for five consecutive runs and catching a documentation-only PR in the blast. The queue is the structural fix (#2412): each entry is built on top of the entries ahead of it, and the tested commit is the commit that lands.

This page is the operating manual: what the first outing taught, the settings that answer it, and the steward โ€” the lane that acts on every ejection so a person does not have to.

What the first outing measured (2026-08-30 โ†’ 09-01)

The queue was enabled on 2026-08-30 by #2799 (merge_group trigger on Build and Test) plus a merge_queue rule on the main pr protection ruleset โ€” ALLGREEN, up to 3 entries built and merged, SQUASH, 60-minute check timeout โ€” and removed on 2026-09-01. Two failure modes, both measured:

  1. The churn window. With three entries built speculatively, every membership mutation โ€” an enqueue, a dequeue, a push to a queued branch โ€” rebuilt the speculative stack and restarted every in-flight build. Several sessions each acting reasonably kept the stack in permanent rebuild: for over an hour nothing landed, and nothing failed. Zero reds, zero merges โ€” the run list showed only cancellations, which nothing alerts on (the same shape as Reading CI Signals ยง "Delivery can stop for hours with every dashboard green").
  2. Hand re-queues. An entry a flake ejected simply left the queue. Nothing re-queued it; the maintainer did, by hand, after noticing. Of the 83 red Build-and-Test runs between 08-30 and 09-02, 51 were merge_group runs โ€” most of them ejections that then needed a person.

Two things did work and are kept: queue merges are ordinary pushes, so main's push lanes ran on every one of them (40/40 on 08-31 โ€” the GITHUB_TOKEN non-triggering trap does not apply to the queue); and Build and Test already fires on merge_group, with Consolidate test results produced unconditionally on that event (collect-results is if: always() with no event filter, and the one gate that needs a pull-request body โ€” cross-repo-pair โ€” is exempted on the event, so its fail step cannot fire on a queue run).

The settings, and why each one

Parameter Value Why
max_entries_to_build 1 Removes the speculative stack, and with it the churn window: with one entry building there is nothing for a membership mutation to rebuild except that entry. The cost is serial throughput โ€” one group build at a time โ€” which is affordable now that a PR run takes 6โ€“19 minutes, and often ~2: a single entry whose PR run tested against the same main tip has an identical tree, and Build and Test reuses that green (refs/ci-green/<tree>) instead of re-running it.
max_entries_to_merge 3 The upper bound on how many consecutive green entries land as one push. Bounds the blast radius of a red main to three PRs.
min_entries_to_merge 1 A green entry is never held hostage to a second one arriving.
min_entries_to_merge_wait_minutes 3 Only delays a group smaller than the minimum, so with a minimum of 1 it cannot bind today; it is set short so that raising the minimum later never introduces a long wait by accident.
grouping_strategy ALLGREEN Every entry in a merge group must be green on its own build; nothing lands on the strength of a later entry's green. (With one entry built at a time, HEADGREEN would be equivalent โ€” ALLGREEN states the intent.)
merge_method MERGE The repo's convention (gh pr merge --merge), keeps each PR's commits and Co-Authored-By trailers, and the commit the queue built is the commit main fast-forwards to โ€” sha-identical, so verifying an image by commit needs no detour through the tree.
check_response_timeout_minutes 45 Matches the fleet's hard per-job cap (check-workflow-timeouts.py): a queue build that has not reported in 45 minutes is stuck by the same doctrine, and the steward re-queues it once. 60 was the 08-30 value; the slowest honest PR run is 19 minutes.

The refs/ci-green/<tree> marker is written only after Consolidate test results passes every collector gate verdict, including documentation, compatibility, policy, shell, and client checks. A successful build and test matrix alone cannot mint the marker: a red gate leaves the tree unverified, so the queue must run its checks instead of reusing that result (#5988).

Two properties of dotnet-test.yml matter for the queue and were checked rather than assumed:

๐Ÿšจ The dependent's suites lengthen a queue build โ€” and the table above has drifted

Every queue entry now also runs Dependent suites (MeshWeaver.Plugins) (policy dependent-suites-gate; The Cross-Repo Pair Gate ยง "The dependent's suites run against the candidate"): MeshWeaver.Plugins builds its reachable suites against the entry's commit and core waits for the verdict AFTER its own tests. A queue build is therefore core's run (~20 min) plus a waiter of up to 45 โ€” about 65 minutes end to end, with each JOB still under the fleet's 45-minute cap. check_response_timeout_minutes must cover that: the 45 in the table above would eject every entry whose Plugins run is slow.

Measured live on 2026-09-25 (gh api repos/Systemorph/MeshWeaver/rulesets/2128472), the ruleset is NOT the table above: max_entries_to_build: 8, max_entries_to_merge: 8, check_response_timeout_minutes: 120. The 120 covers the gate. The 8 means up to eight entries โ€” eight Plugins candidate runs, each up to ~12 legs on aks-silos-dind โ€” can build at once, which that pool (โ‰ˆ24 runners) cannot serve alongside Plugins' own CI. Whether to return to a small max_entries_to_build is a maintainer's ruleset edit; merge-queue-steward.py status prints the drift.

๐Ÿšจ What that drift did: the queue froze on a waiter with NO verdict (2026-09-27)

The paragraph above predicted it, and on 2026-09-27 09:00โ€“11:00Z it happened: nothing merged for two hours. Three entries in a row (core runs 36307979031, 36308367885, 36311194922) failed on ONE job, Dependent suites (MeshWeaver.Plugins). Each time its waiter hit its 42-minute verdict deadline (await-dependent-verdict.py --deadline-minutes 42, inside the job's 45-minute timeout-minutes) with no verdict. Not one was a test failure. The Plugins candidate runs were green, and each landed 3โ€“10 minutes after core had stopped waiting (candidate run 36308382966 published at 10:18Z; its core waiter gave up at 10:08Z).

measured over REST value
a candidate leg's run time 5โ€“14 min
its wait for a runner 20โ€“50 min
aks-silos-dind registrations 24 (the cap), all busy. The 6 listed "offline" were ephemeral runners still starting, not zombies: each picked up a job minutes later.
the set's runner-minutes 08:00โ€“10:50Z 4,477, of which 27% went to candidate legs and the rest to Plugins PR/push lanes and satellite bakes
a candidate run whose core run had already finished 36311218290: up to 8 runners for 74 min after core run 36311194922 failed, while four newer entries' legs queued behind it

Root cause. The candidate legs shared the aks-silos-dind label with every Plugins PR and every satellite bake. GitHub hands a scale set's queued jobs out first-come-first-served, with no job priority. So the one piece of work that gates EVERY core merge waited behind whatever PR work was queued before it. The set's cap is the hardware (the CI pools take the whole spot and Dsv6 quota), so raising it on the same label would only lengthen the same line. What was missing was an ORDER, not capacity. The run that kept going after core stopped waiting was a second, smaller defect on the same path.

The fix, in the repos that own each half:

๐Ÿšจ A Dependent suites NO-VERDICT is infrastructure, and the steward re-queues it. The steward used to reject it with every other non-shard job ("a build or gate failure is never a flake"), which left five entries (#5789, #5791, #5792, #5793, #5795) queue-rejected that morning for starvation alone. Silence and a verdict are now two different reds. await-dependent-verdict.py exits 3 on silence. dotnet-test.yml fails that case on its own step, No verdict in time: the dependent's suites did not report (infrastructure). A Dependent-suites job that failed on that step and nothing else is re-queued as infra, capped at 2 per head sha, with the usual marker comment. A red verdict (the candidate broke a dependent suite) still fails the wait step itself and is still rejected, and a no-verdict next to a build failure or an uncatalogued assertion is rejected too. Before you touch an entry that hit the cap, read the candidate run's leg WAIT times: a leg that waited longer than it ran is starvation, and the gate lane above is where to look.

The remaining lever is max_entries_to_build: 8. Up to eight concurrent candidates of up to 12 legs each is more than the 12-runner gate lane can serve inside core's 42-minute verdict deadline. Whether to return to a small value is still the maintainer's ruleset edit, as the paragraph above says.

Enabling it

The rule is added to ruleset 2128472 (main pr protection) with the REST rulesets API. PUT replaces the whole ruleset, so the existing rules are read back and the queue rule appended:

gh api repos/Systemorph/MeshWeaver/rulesets/2128472 \
  --jq '{name, target, enforcement, conditions, bypass_actors,
         rules: (.rules + [{type: "merge_queue", parameters: {
           merge_method: "MERGE", grouping_strategy: "ALLGREEN",
           max_entries_to_build: 1, max_entries_to_merge: 3,
           min_entries_to_merge: 1, min_entries_to_merge_wait_minutes: 3,
           check_response_timeout_minutes: 45}}])}' > /tmp/ruleset-2128472.json
gh api -X PUT repos/Systemorph/MeshWeaver/rulesets/2128472 --input /tmp/ruleset-2128472.json

Read it back โ€” mergeQueue(branch:"main") is null while disabled:

python3 .github/scripts/merge-queue-steward.py status --repo Systemorph/MeshWeaver

status prints the live configuration, warns on every parameter that drifts from the table above, and lists the entries currently queued.

The steward โ€” the hand that re-queues, on evidence only

.github/workflows/merge-queue-steward.yml fires on pull_request with action: dequeued; the event carries a reason. The workflow is thin; .github/scripts/merge-queue-steward.py decides, mints nothing itself, and proves its own decision table with --self-test before every real decision. It uses two tokens, split by direction: every read (the failed run's jobs and artifacts, the PR head's check-runs, commits, comments) goes through the job's own GITHUB_TOKEN with actions: read + checks: read; every write (comment, label, enqueue) goes through the App installation token minted the way auto-arm.yml does โ€” the org grants that App exactly contents: write + pull_requests: write, so it cannot read Actions, and a re-queue performed with GITHUB_TOKEN would merge as the bot and start no run on main (#2916). The queue-rejected label is a repository fixture (creating a label needs issues: write, which the App does not hold; applying an existing one needs only pull_requests: write).

Reason What the steward finds Action Cap per head sha
CI_TIMEOUT โ€” re-queue 2
CI_FAILURE a job other than a test shard failed (build, a gate) reject โ€” never a flake โ€”
CI_FAILURE every failed assertion matches an active catalogue entry re-queue 2
CI_FAILURE a shard failed on an infrastructure step (download, upload, setup) and left no test evidence re-queue 2
CI_FAILURE the only failed job is Dependent suites (MeshWeaver.Plugins), on its no-verdict step: the dependent never reported inside the deadline re-queue (infra) 2
CI_FAILURE an uncatalogued assertion, the group held more than one PR, and this PR's own run was green re-queue alone โ€” the culprit's solo group fails and stays out 1
CI_FAILURE anything else โ€” an uncatalogued assertion, a dead host with no recorded failure, no artifact to read reject: comment the assertion and the run, label queue-rejected โ€”
MANUAL, QUEUE_CLEARED, ROLL_BACK, BRANCH_PROTECTIONS, GIT_TREE_INVALID, INVALID_MERGE_COMMIT, MERGE_CONFLICT, UNKNOWN_REMOVAL_REASON โ€” comment once, no action โ€”
MERGE, ALREADY_MERGED โ€” nothing โ€”

How it reads the failure: the newest merge_group run of MeshWeaver Build and Test whose head branch is gh-readonly-queue/main/pr-<N>-โ€ฆ; its failed jobs; for each failed shard the testResults-shard<N> artifact โ€” the same evidence Consolidate test results downloads โ€” parsed for <UnitTestResult outcome="Failed"> with message and stack, and for non-TESTFAIL exit markers in test-results.log. The group's PR list is the first-parent chain of the queue commit back to a commit on main; "own run green" is Consolidate test results on the PR's head commit.

Every attempt is recorded in a hidden marker on the PR โ€” <!-- steward: requeued=N head=<sha> kind=<kind> --> โ€” and caps are counted from those markers per head sha: a new push is a new question. A rejection spends nothing.

๐Ÿšจ The steward never re-runs a workflow. A re-run of the same tree hides the bug the failing run found (and the run's failing transcript is the control arm for the flake). A re-queue builds a new tree against the main that has moved โ€” a different measurement.

The flake catalogue

.github/known-flakes.json is the only thing that turns a red assertion into a re-queue, so each entry is a temporary, evidence-bearing allowance:

{
  "id": "graph-late-nack-timeout",
  "assertionPattern": "TimeoutException : The operation has timed out\\.[\\s\\S]*LateNackReenqueueTest\\.cs:line 131",
  "testName": "MeshWeaver.Graph.Test.LateNackReenqueueTest.LateOwnerDisposingNack_AfterOptimisticEmit_ReenqueuesAndLands",
  "issue": "https://github.com/Systemorph/MeshWeaver/issues/NNNN",
  "evidence": ["https://github.com/Systemorph/MeshWeaver/actions/runs/33630685580"],
  "addedOn": "2026-09-02",
  "expires": "2026-10-02",
  "addedBy": "rbuergi"
}

The catalogue was seeded empty on 2026-09-02, and deliberately. The 700 Build-and-Test runs since the queue was first enabled hold 83 reds; the repeated signatures were MeshWeaver.Hosting.Monolith.Test.HOST_CRASHED (ร—5, exit=124 at the 8-minute cap) and CompileFinishAndDisposeTest (ร—4) โ€” both in a project that has since moved to MeshWeaver.Plugins โ€” and TestTimeoutLiteralRatchetGuard (ร—6 on 09-01), which was the queue working: two PRs each moved a ratchet count, green alone and red together. Every remaining candidate was a single occurrence, or its subject changed after the failing runs (ScopeTeardownRenderTest, #2877 on 08-31). A catalogue seeded on weaker evidence than that would be an automatic re-runner with a JSON file in front of it.

Guards

๐Ÿšจ Auto-merge waits on the BASE's protection โ€” so a stacked PR is armed to merge NOW

--auto does not mean "merge when this pull request is green". It means "merge when the contexts branch protection requires of the BASE are satisfied" โ€” and in this fleet protection is configured on the default branch and nowhere else. A stacked pull request, opened against another pull request's feature branch, therefore has a base that is protected by nothing, an EMPTY required set, and a condition that is met the instant GitHub first reads it.

Measured 2026-09-11. MeshWeaver.Plugins#1685 was opened against the feature branch of #1681. auto-arm.yml armed it and GitHub merged it at 21:03:23Z โ€” 61 seconds after it was opened, before its own CI had started a single job. Nothing failed and nothing was bypassed; the stack collapsed unreviewed, exactly as configured. auto-arm.yml is the fleet's only copy of the arm lane and every satellite reaches it through workflow_call, so the hole was open in every repository at once.

๐Ÿšจ But "the satellites get it for free" is true only where the caller pins @main, and one does not. Measured 2026-09-11 over the contents API (never a local clone โ€” satellite checkouts here run days stale): all seven callers exist, and six pin auto-arm.yml@main โ€” MeshWeaver.Plugins, .Reinsurance, .SocialMedia, .Education, .Crm, .Manufacturing โ€” so a core fix lands there on merge. Systemorph/Memex pins a SHA (@c7fef7a2, 971 commits behind core's main when measured), so it picks up nothing until that line moves. The lesson generalises past this fix: before claiming a reusable-lane change reaches the fleet, read every caller's uses: ref โ€” a fleet-wide fix and a six-of-seven fix are indistinguishable from core, and the odd one out is silent, not red.

The lane now reads github.event.repository.default_branch off the event โ€” never a literal main, which would be a copy of a fact every repo happens to share today and the exact drift the single-copy design exists to end โ€” compares it with github.event.pull_request.base.ref, and arms only on a match.

๐Ÿšจ One self-clearing window, named so it is not misread as the fix failing. The lane runs on pull_request_target, and that event runs the workflow file as it exists on the pull request's BASE branch โ€” not on main. So a stacked pull request opened onto a feature branch that was cut before this fix landed still runs the old, unguarded copy and is still armed immediately. Nothing in core can reach that: the branch carries its own copy by the event's definition. It clears as soon as the branch is cut from, or catches up with, a main that has the fix. If you see a stacked pull request merge instantly in the days after this lands, check the age of its base branch before concluding the guard is broken.

๐Ÿšจ The decision is announced, and the JOB is never what skips. Moving the base test onto the job's if: is the tidy-looking version of this fix and deletes the message with it: a skipped job renders like a passed one and carries no warning, no summary and no comment, so "the lane declined" and "the lane is broken" become the same page. Instead the job runs, a ::warning:: plus a step summary name the base, and one idempotent comment says so on the pull request itself. Draft and fork stay on the job condition, because for those there is genuinely nothing to say.

This section is the HISTORY of the retired arm lane, kept for the lesson it taught. auto-arm.yml no longer arms anything, so the base-chain guard it once carried is gone with the arming step: the control plane (Plugins PrArming) is the only armer and applies the arm gate, which includes the base-is-the-default-branch condition. What still holds the property in core is ArmedMergeMustTriggerMainsPushLanesGuard.NoWorkflowArmsAutoMerge: no workflow in the directory arms auto-merge at all, so a stacked pull request can no longer be armed by a workflow here. The other lanes were checked structurally, not by luck: arm-credential.yml mints and inspects a token and arms nothing; the steward acts only on a dequeued event; release.yml and node-repo-platform-ref-bump.yml OPEN pull requests (--base main) without arming them.

Working with the queue

๐Ÿšจ Disarming auto-merge is not a hold โ€” DRAFT is the only durable one

Measured: MeshWeaver.Plugins#1683 was deliberately disarmed during a merge window and merged anyway. Nothing malfunctioned: at that time auto-arm.yml armed on synchronize, so the next push re-armed it. That lane has since been retired, and the shape of the lesson changed with it.

Today auto-arm.yml runs on synchronize only and disarms; it never arms. The control plane arms a head once it is reviewed, every finding is answered and the required checks are green, so a disarmed pull request stays disarmed until the control plane judges a new head ready. That is still not a hold: any such head is armed again without anyone asking, and a pull request someone is still pushing to is exactly one that will be. Draft remains the durable hold, because the control plane never arms a draft.

act what it is how long it lasts
gh pr merge --disable-auto a state GitHub owns until the control plane next judges the head ready and arms it
convert to draft a property of the pull request the control plane reads until you mark it ready

(GitHub also disables auto-merge when a pull request is converted to draft, so the conversion does both halves in one act. That clause is GitHub's documented behaviour rather than something measured here.)

So: to hold a pull request, convert it to draft. Never rely on --disable-auto โ€” and if you find a PR merged that you thought you had stopped, look for the control plane's arm after the disarm.

#4057 does not change this. That fix stopped the retired lane arming a pull request whose base is not the default branch; the control plane's arm gate carries the same condition now.

The option considered and NOT taken

(Written when the lane still armed; kept as the design record. The re-arming agent is now the control plane, and the same two options would apply to it.)

Making the lane refuse to re-arm a pull request a human explicitly disarmed is attractive and is not small, so it was left alone rather than half-built:

Either way the rule that must survive is the one this document already states: a decision not to arm has to be SAID, on the pull request. A hold that works by something quietly not happening is the same defect as a gate that skips on a missing input โ€” silence and "still running" look identical.

๐Ÿšจ Before concluding a pull request is queued โ€” or that it is stranded โ€” read Reading CI Signals โ†’ "mergeable_state: clean + auto_merge: false is a TWO-POLE ambiguity". The two states demand opposite actions, they are byte-identical over REST, and a gh-readonly-queue ref proves nothing in either direction; only the entry list answers it.

Reading CI Signals ยท The Continuous Delivery Contract ยท The Cross-Repo Pair Gate ยท Writing Tests