Bug triage — one global queue, one dedicated agent, one permanent thread per defect

Maintainer, 2026-09-14: bug triaging should be in Essentials — re-use globally — only one bug triage queue with dedicated agent. So: this concept ships with Essentials (pre-installed for everyone), there is exactly one queue — on the control instance, fed by the pooled inbox every portal and repository already posts to (Hosting/Triage) — and exactly one agent, Essentials/Agent/bug-triage, works it. No module, portal or repository runs a triage of its own; a module that finds a bug hands it to the queue (/feedback for a person, a signed ci-failure / incident event for a machine, /bug-fix-thread for an agent already on it).

The general model is Governance/Design; the catalogue is Governance/Catalogue. This page is the Essentials package's part for defects: a red main, a production incident, a technical feedback — anything that ends in a GitHub issue today (Hosting/Triage, the mechanism) and then, too often, in silence.

The one idea

A bug gets one thread, and the thread outlives every session. It is anchored on the node that raised the defect — the Hosting/TriageItem (a ci-failure, a feedback) or the Hosting/Issue (a FleetWatch incident) — and it is closed by exactly one thing: a bug.verify activity that read the fixed behaviour on the target production instance. Not a merged PR, not a green build, not a comment saying "should be fixed". The maintainer's rule (2026-09-09): "Done", "fixed", "shipped" means production behaviour was verified.

The thread is where the bug-triage agent works, where every proposal it raises is linked, and where a person reads the whole story cold — six months later, in one place.

The chain of standards under a bug

red main · incident · feedback ──▶ Hosting/TriageItem / Hosting/Issue ──▶ ONE thread: Hosting/Triage/_Thread/{t}
                                                                             │
  bug.thread     (unattended)  open the thread, file / link the GitHub issue, name the owner repo
  bug.diagnose   (unattended)  the fixer's finding as a record: symptom · cause · evidence · fix direction
  dev.merge      (governed)    the fix PR — Role.Reviewer reviews the diff, a HUMAN maintainer signs the merge
  hosting.roll   (governed)    the deploy — Scope.Production: passkey-backed human signature   ─┐ one of the two,
  module.publish (unattended)  a node module: main IS the approved state; the lane publishes    ─┘ by what the fix touched
  bug.verify     (unattended)  the read-only proof on the target instance — closes the issue and the thread
Step Standard Proposer → Reviewer → Signers → Executor (identity) Ends with
Open bug.thread the inbox (a signed event) / a person → none → none → Propose + CreateIssue (App) the thread, the issue, the owning repository named on the item
Diagnose bug.diagnose Essentials/Agent/bug-triage → none → none → Record (Control) diagnosis on the item: symptom, root cause, evidence (log lines, node paths, run ids), the fix direction, the module page the finding goes to
Fix dev.merge the fixer (its PR) → Role.Reviewer (the diff against the diagnosis; Copilot's review is input) → a human in the repo's Role.Maintainer, never the author → MergePullRequest (App) the merge sha, linked on the item
Deploy hosting.roll — or module.publish for a node module the build server (trunk green) → Role.Reviewer → Role.ProductionApprover, passkey → Build(deploy) / Workflow (App) — or nothing for a module: the lane publishes and the portals adopt the tag the target instance runs, or the module version it adopted
Verify bug.verify the fixer → none → none → Probe (Control: a read — a render, a Sample, a log query with the fingerprint absent) verified: true with the evidence, the issue closed by the activity, the thread closed
Reopen bug.thread again the inbox (the fingerprint recurs — LogIncidentFiler's reopen) the same thread, a new round

Three rules:

  1. Nothing but bug.verify closes. A merged PR moves the item to fixed, unverified; the thread stays open. A fixer that cannot verify (no access to the instance, the behaviour is not observable) says so on the item and hands the verification to a person — it does not close.
  2. The diagnosis is conserved where the module lives. bug.diagnose names the page the finding goes to ({Module}/…md, a WhatsNew entry for a user-visible fix); the fix PR carries it. A finding that lives only in the thread is lost to the next session.
  3. The fixer never merges, never deploys, never holds a token. It writes code, opens the PR, proposes dev.merge, follows, proposes bug.verify. The merge is a human's signature; the deploy is the production scope's.

How a bug reaches its thread — the intake's hand-over

The pooled inbox (Hosting/Triage) classifies and files; it does not work the defect. The hand-over is the ISSUE: a defect gets its thread when its GitHub issue is a classified bug — bug plus exactly one sev: (what the release gate counts) — and that is recorded on the issue's own item, Hosting/Triage/issue/{repo}-{number}. Three events do it (TriageIntake.EnsureBugThread, MeshWeaver.Plugins#2566):

Event Where Effect
a bug issue is opened already classified — including the triage agent's own FileIssue, which comes back through the organisation webhook as an issues/opened of systemorph-com[bot] TriageIntake.HandleIssue the item is recorded Done (triage has nothing to do) and the bug thread is started
triage classifies an opened issue as a bug TriageActions.LabelIssue, after the labels are applied the bug thread is started on that item
a bug issue is reopened TriageIntake.HandleOpen a message on the SAME thread — a new round; its first thread if it never had one

The backlog from before the hand-over — adopted by the pool itself

A bug whose item was recorded BEFORE this hand-over existed is Done with the note "Classified when it was opened — nothing to triage." (or was labelled by the triage agent then) and carries no bugThreadPath and no bugQueuedAt — so the pool, which visits only the items it owns, would never see it. Every pool pass therefore also looks for that backlog (BugFixPool.LegacyBugOf): an issue item, Done, with neither field, whose recorded labels (labels ∪ appliedLabels) are a classified bug of a member or fleet bot (TriageIntake.BugThreadRefusal — a stranger's issue is never a candidate and is not even read). For at most five candidates per pass, oldest first, it reads the issue through the App: a closed issue is recorded (issueClosedAt) and never adopted; one that is open but no longer a classified bug is recorded (bugAdoptionRefusal); an unreadable one is left for the next pass (fail closed); an open classified bug goes through the SAME hand-over as the intake — TriageIntake.EnsureBugThread — so the pool's assignment and bounds, EU-only (no EU variant served → queued, never another model) and the one-thread-per-issue path all apply. At most one bug is adopted per pass, and never in a pass that started a queued bug, so an old backlog trickles in rather than flooding the pool. The item records bugAdoptedAt, bugAdoptionAttempts and a note saying it was adopted from before the hand-over (Open while queued, Done once its thread exists); a hand-over that fails three times is left for a person. An issue GitHub gives no answer for is neither adopted nor settled: the item is stamped bugAdoptionUnreadAt and goes to the back of the order, so one unreadable issue (deleted, or in a repository the App no longer covers) cannot hold the head of the backlog behind the five reads a pass makes. The ledger counts bugPoolAdopted per day and shows bugPoolBacklog (still waiting). Tests: BugFixPoolTests (TheBacklogRule_*, TheLiveHalf_*, TheBacklog_*) and, on a real mesh, src/MeshWeaver.Fleet.Control.Test/ABacklogBugFromBeforeTheHandOverIsAdoptedTest.cs (controls: a stranger's, an already-threaded and a closed issue; a thread that already exists is recorded, never a second; with no EU variant served the adopted bug is queued).

🚨 Rounds are capped — the work is a chain of rounds, re-driven by the pool

An agent round is aborted at 30 minutes and an aborted round records nothing (see MeshWeaver.Plugins#2568 for the reviewer's instance of the same cap). The chain above is therefore worked ONE STEP PER ROUND. Each round reads the item, does one step, records it and stops. The Dispatch (below) posts the next round whenever the thread is at rest, so a bug moves forward around the clock. No round is ever one long round, and a bug never waits for a person unless the agent said it is blocked.

Dispatch — 24/7, per estate, several models

Bug fixing is done by agents around the clock, and several models work side by side so the comparison shows which works best. Dispatch is one design, instantiated per ESTATE by configuration. An estate is a control instance with its own triage scope, its own repositories, its own model account and its own Dispatch. Our estate's Dispatch runs on the Systemorph control instance. A client estate's Dispatch runs on that estate's own control instance: it has its own repositories (its CRM-read triage scope), its own OpenRouter account (its own Provider/OpenRouter key) and its own variants. The code is the same on every estate. Only the node below differs.

The name is Dispatch; the wire identifiers keep the old one

Everything a person reads calls this Dispatch — the page title, the node's name, the editor labels, the status notes and these docs. It used to be called the "bug-fix pool". The identifiers that code, configuration and deployment records bind to keep the old name on purpose:

Kept as is Why
NodeType Hosting/BugFixPool, node Hosting/BugFix in-mesh code and every estate's stored node reference these paths; a rename moves data and breaks callers no dotnet build sees
config keys Hosting:BugFix:* (Hosting__BugFix__Enabled, …) deployment records in Systemorph/Memex mesh/Deployments/*.json set them; a renamed key is silently ignored, so the old value would stop applying
JSON field names (lastBugPoolAt, bugPoolRuns, …) and C# types/members (BugFixPool, BugFixPoolContent, …) stored content and in-mesh sources read them by name
policy id bug-fix-eu-only, lane id bug-pool ids that registers, ledgers and tests match on

Renaming any of these is a separate, breaking change with its own migration — not a text edit.

An estate whose Hosting/BugFix node was created before the rename still carries the old seeded name ("Bug-fix pool") and description. Dispatch's own pass moves them to the new name the next time it runs (BugFixPool.Renamed): only a value still equal to its former seed is replaced, so a name or description a person chose is kept.

The switch — On · Drain · Stop

Dispatch has ONE live switch, the mode field of Hosting/BugFix. A global admin flips it on the Dispatch page (the Mode select at the top, a framework editor bound to the node) or with an ordinary patch of content.mode. Every pass and every intake reads it off the node's own stream, so a flip takes effect on the next pass and the next event, with no restart, Reconcile or recycle.

Mode New work (queued bugs, a red main's worker, backlog adoption, a redo one tier up, bundling) New triage intake (webhook issue / feedback / red run / digest, the hourly issue sweep) What is already running
On starts starts its thread runs
Drain waits, queued (bugQueueReason: "Dispatch is not starting new work…") held: the item is recorded, open, with no thread gets its rounds and its review, and finishes
Stop waits held every running round of a fix thread or its review thread is cancelled

Pinned by DispatchSwitchTests (in-mesh: mode resolution, the stop record, the pending-stop decision, the queue reason, held intake, the page) and, on a real mesh, DispatchSwitchTest (Drain queues / On starts; Drain holds intake / On resumes it), DispatchStopCancelsARunningRoundTest (Drain leaves a running round alone; Stop cancels it and gives it back; a pending stop is cleared for a round that completed and recorded for one that was cancelled) and AGlobalAdminFlipsTheDispatchSwitchTest (an admin sets the mode; a non-admin is refused; the admin is refused on another node of the same type).

Who runs on which model — routing by role

Every model Dispatch uses is declared ONCE, per role (DispatchRoles, an open vocabulary), in the process (BugFixProcess.Roles). The pool node overrides any role in place (roles: role, model, effort, contextLimit), with no redeploy. Every role names an explicit reasoning effort; none is ever left blank. All of them are EU routes (Provider/OpenRouterEU/…).

Role What it does Model Effort
triage classify, rate the difficulty, park or go — one bounded round Claude Sonnet 5.5 medium
fix the main fix thread: plans, edits, decides — rung 1 Kimi K2.7-code (EU: one provider; 262k context, declared) xhigh
(rung 2) the main fix thread after a failed attempt, a Kimi that serves no tokens, a context beyond 262k, or sev:B / sev:H from the start Claude Sonnet 5.5 xhigh
escalation the main fix thread's last rung — only after the SECOND failed attempt Claude Opus 5.5 xhigh
read a sub-thread that only reads: get, search, read files, summarise logs and CI output, extract verbatim GPT-6 Luna medium
write a sub-thread that does ONE mechanical write: one patch, one governed write, one status GPT-6 Luna medium
self-review reviews every attempt of the fix (another model family than the author) GPT-6 Sol xhigh

Pinned by DispatchRoutingTests (in-mesh: every role and effort, the ladder and its escalation, severe start, the no-tokens skip with its negative controls, infra rounds outside the budget, the spend by role, the one-time upgrade and the re-pin). The older four-rung routing's mechanics stay pinned on LegacyRouting.

🚨 EU only — fail closed (policy bug-fix-eu-only)

Every round of a bug thread runs on an EU-routed model, and only on one:

The configuration is a node: Hosting/BugFix

Hosting/BugFix (NodeType Hosting/BugFixPool) is runtime state on the estate's control instance. It is ExcludeThisOnly, so a Hosting import never prunes it. The pool's sweep creates it on its first pass and seeds it once. After that it is edited in place: adding, dropping, disabling or re-ordering a variant takes effect at the next assignment and the next pass, with no deploy and no restart.

Field Meaning
enabled false stops the pool on this estate: every bug stays queued and nothing is re-driven
cooldownMinutes (20) how long a thread at rest waits before the next round
waitingCooldownMinutes (180) the same for a bug whose agent said awaiting-merge / awaiting-deploy
maxRedrives (24) rounds per bug before the pool stops and names the bug for a person
variants[] id, agentPath (default Essentials/Agent/bug-triage), model (a model NODE path on this estate), effort (xhigh for every coding variant), dataResidency (only Eu takes bugs), enabled, maxConcurrent (N), order, note
flow the triage agent and model, the fix agent and model, the escalation ladder, maxConcurrent, escalateAfterBlocked — when set it replaces the variants (below)
table[] the model table: kind (bug · ci-failure · feedback · *), difficulty, fixer (a rung of the ladder, or auto:…), reviewer, rounds (below)
ciWorkers (true) false: a red main gets no fix worker on this estate; bugs still do
ciMaxAgeHours (24) a red run older than this is never handed to a worker
calibration[], workLedger[] written by the pool, never by hand: outcomes per (task kind, model) for the Auto router, and per (work kind, difficulty, model) for refining the table
supervision enabled (true), reviewAgent (Essentials/Agent/bug-review), briefSkills[] (Essentials/Skill/bug-fix-thread), diffBudget (changed lines per difficulty: 300 · 1000 · 3000), toolLoop (6), maxReviewPosts (2) — the brief, the early stops and the review (below). Absent on a flow pool: these defaults

🚨 Every one of these is TYPED on BugFixPoolContent — the node's own hub re-serialises the record on every write, and a field it does not declare is dropped by the next one (pinned by TheSeed_CarriesTheModelTable_AndTheTypedContentKeepsIt).

A variant takes bugs only when all of these hold (BugFixPool.Ineligible):

An unknown residency is never EU.

Every item kind gets its own worker — bugs, CI failures, feedback (maintainer, 2026-10-02)

"For all issues we find in the triage, we spawn a sub-agent where we find the most suitable model to use." Decided the same day: it runs in the mesh triage on the control instance (durable, not in a local session), and the items that get their own worker are bugs, CI failures and feedback — a feature request or a question stays classified only; no agent writes a feature unasked.

Until this change only a classified bug ISSUE got a worker. A red main and a feedback got one thread with Hosting/Agent/triage on a fixed tier, which classified, commented or filed, re-ran a transient once and marked the item Done — nobody FIXED a content or platform break, and a feedback defect was worked only if it happened to come back as a bug issue. Now the SAME flow — the scope round, the model choice, the escalation, the ledger — runs for three work kinds, recorded on the item as workKind and shown as a column wherever the pool shows work:

Work kind What raises it Gets a worker when Never when
bug a GitHub issue it is a classified bug (bug + one sev:) of a member or a fleet bot — as before a stranger wrote it (Held); it is a red-main ledger issue (label ci-main-red — the red run's own item carries the worker)
ci-failure a signed ci-failure event — a red main / push run triage's verdict is content or platform, the repository is in the triage scope, the run is fresher than ciMaxAgeHours (24), and main is still red the verdict is transient (triage's single re-run, nothing more) or missing; the repository went green (greenAt, or its ledger issue is closed); the scope is UNKNOWN or does not cover the repository
feedback a signed feedback event triage's verdict is defect: it is filed as a bug with one sev:, and the worker hangs on that filed issue's item the verdict is idea or question (filed WITHOUT bug: classified and nothing more), praise or not-actionable (Done with a note, never filed); the feedback was submitted outside the trusted instances (Held)

The routing is enforced by the control plane, not left to the prompt. FileIssue refuses a feedback draft whose labels contradict its verdict (TriageActions.FeedbackVerdictGap): a defect without bug would silently get no worker; an idea or question with bug would start a worker on a feature; praise / not-actionable are not filed at all; and a feedback with NO verdict from the vocabulary is not filed until one is written — otherwise the gate could be passed by saying nothing. A red run is handed over by the pool's own pass from the item's recorded verdict (BugFixPool.CiRefusal names every reason it is not).

One red main, one worker — red runs are bundled by episode. A repository's main usually stays red for several runs, and each run is its own item. The first one owed a worker becomes the lead: its thread is Hosting/Triage/_Thread/ci-{repo}-{ledger issue} (the repository's ci-main-red ledger issue is opened on the first red, appended per run and closed on green, so its number IS the episode; with no ledger issue named, ci-{repo}-run-{run}). Every later red run of that repository joins it — workLead on the run, workBundle on the lead, no second thread — as long as the lead is not finished. The hand-over is done by the ONE sequential pool pass (never by the red runs' own hubs), so two runs cannot race into two workers; the pass reads the leads as it has moved them. A red that flaps back inside the ledger's window is the same episode: the new run adopts the same thread. main green again ends the work — the ci-green event stamps greenAt on every worked red item of the repository (TriageIntake.WorkedRedItemsOf), and the pass also reads the ledger issue's state through the App before it starts a worker and on every visit.

A red PULL REQUEST is not a ci-failure. The lane posts the event for main only; a red on a pull request — the worker's own included — is the PR babysitter's (Hosting/PrBabysitter).

🚨 The Held rule covers feedback too. A feedback is text anyone who can sign in to that portal can write, and the issue filed from it is opened by our own App — which the stranger rule trusts. So the control plane states the origin on the FIRST line of every issue it files from a feedback (TriageIntake.OriginMarker: the feedback item, and trust=trusted or trust=held), and the hand-over reads that line on an issue a fleet bot opened (TriageIntake.HeldReason). Trusted means: submitted on this control instance itself, or on an instance the record names in Hosting:Triage:TrustedFeedbackInstances. Anything else is Held — the issue is filed and classified, its item stays Open with a HELD note, and no model that writes code reads it. The line is the control plane's: any marker-shaped line in the reporter's text is dropped before filing, only the first line is read, and a person who types it into their own issue has stated nothing. This tightens what ran before: a stranger's feedback, filed as a bug by the triage agent, used to get a worker because the issue's author was our bot.

The queue order when a slot frees (BugFixPool.Rank): a red main first (it stops every module's delivery, the fix of any bug included), then sev:B, sev:H, sev:M, sev:L, then what carries no severity; the oldest first within a rank.

Every existing guard is unchanged and applies to all three kinds: EU only and fail closed (no served EU rung → queued, never another model), the concurrency bound (a red main's worker takes a slot like a bug's), maxRedrives, pull requests opened and never merged, production only through governed activities. ciWorkers: false on Hosting/BugFix switches the red-main workers off on an estate without a deploy.

The model table — kind + difficulty → model (maintainer, 2026-10-01)

The table and the stronger-rung review described in this section and the next are the routing of process version 1. They still apply where a node opted out of the upgrade, and as the fallback when no self-review model is served. The shipped routing is "Who runs on which model" above.

"Then we choose a model we consider best for the job — let's start with a configuration we guess of what may be good for what. We can then still refine later, e.g. by measuring."

The scope round rates the work; the table on Hosting/BugFix (table, edited in place) then names the model. A flow pool whose node carries no table rows — one created before the table existed — runs on the seed table, so the starting guess does not wait for an edit; to put a difficulty on Auto, give it a row whose fixer is auto:coding. The seed (BugFixPool.SeedTable) is the maintainer's starting guess, the same for every kind until measurement says otherwise (a row whose kind is a work kind wins over a * row):

Complexity Fixer Fixer $/M in · out Reviewer (one tier up) Reviewer $/M in · out Rounds per attempt
scope round (diagnosis, scope, difficulty) Opus 5.5 (high), once per piece of work 4.40 · 22.00 — — 1
simple Kimi K2.7-code 0.67 · 3.35 GLM-5.3 (max) 1.40 · 4.40 3
standard GLM-5.3 (max) 1.40 · 4.40 Sonnet 5.5 (xhigh) 2.20 · 11.00 4
hard Sonnet 5.5 (xhigh) 2.20 · 11.00 Opus 5.5 (xhigh) 4.40 · 22.00 5
hardest Sonnet 5.5 (xhigh) 2.20 · 11.00 Opus 5.5 (xhigh) — which also takes it over after a redo; Opus's own work goes to a person 4.40 · 22.00 6

🚨 Opus manages; for a fix it is only the LAST rung (maintainer, 2026-10-03). Opus is the model of the triage (manager) round. Implementation and fix attempts START on a cheaper rung and reach Opus only after the rungs below were reviewed as bad (a reviewer's redo) or blocked twice. ChooseModel enforces it whatever a table, a host's override or the learning record says (StartCeiling: one below the top), the declaration refuses it (BugFixProcess.WithStart), and the Opus model node is not labelled coding, so Auto(coding) never picks it either. The seed table is the declared process (BugFixProcess.Default.ToTable()), not a second copy.

(EU prices per million tokens as read from the provider's EU endpoints on 2026-10-01; the nodes carry them, the table carries node paths.)

The choice (BugFixPool.ChooseModel, pure, every case tested) is made ONCE, on the first fix round after the scope round rated the work, and recorded on the item:

The row's fixer The fix runs on workModelReason says
a rung of the ladder, served on this estate that rung, pinned (bugLadderStep) table */standard → …/glm-5.3-coding (rung 2/4)
a rung that is NOT served the next served rung ABOVE it — never one below: the row said this much model is needed … not served on this estate — the next served rung up: …
an auto: selection Auto among the EU models carrying the label, as before table … → auto:coding
no row · not a rung · nothing served from the row's rung up the flow's own selection (auto:coding) why

What is served is recorded and shown. The node holds the DECLARED ladder; every pass resolves each rung onto this estate's node of the same model on the rung's own provider (ResolveOnEstate — the provider carries the residency and retention the governed rung was declared with, so an OpenRouterEU rung never resolves onto the global OpenRouter route) and writes the result onto the pool node: resolvedLadder (each rung as resolved here), servedLadder and missingLadder. The model table on Hosting/BugFix shows, per row, the declared fixer and reviewer next to Starts on (served here) and Reviewed by (served here) — the next served rung above the start, or a person when none is — and names the missing rungs under the table. The reviewer recorded on an item (workReviewer) is only ever a served rung. Measured 2026-10-04 on the control instance: matched by wire id alone, the Kimi rung resolved onto Provider/OpenRouter/moonshotai/kimi-k2.7-code (global — it sorts before …OpenRouterEU/…) and a fix round ran off the EU route, although Provider/OpenRouterEU/moonshotai/kimi-k2.7-code existed.

Spend is booked to the round that spent it. A round that delegates (the scope round handing its node patch to Worker, say) spends on the delegate's model, and that spend belongs to the delegating round's process: the scope round's delegation is triage spend, and only a top-level fix round names the fix model (bugChosenModel). Before 2026-10-04 a scope round's delegation was booked as a fix on the delegate's model before any fix round had run.

The item records workModel (the choice), workModelReason, workRung (n/N), workReviewer and workRoundBudget — the reviewer and the budget are what the review and the over-budget stop below run on; bugChosenModel stays the model a round actually RAN on (read off the thread). The pool page shows, per item: kind, difficulty, model, rung and why — and the item's Detail page the same. An escalation moves up from the chosen rung and re-states workRung; bugEscalatedFrom says where it came from.

Refining it. Each pass also writes workLedger on the pool node: one row per (work kind, difficulty, model) with finished fixes, merges, cost and median rounds — the page shows success rate and cost per success. That is the measurement a table edit is made from ("simple red mains merge on Kimi at a fifth of Sonnet's cost"). The table itself is a person's edit; what moves by itself is where a class STARTS — the learning record (below).

The flow — triage, then the table's model, then escalation (maintainer, 2026-10-01)

The seed no longer spreads bugs over fixed model buckets. A NEW pool node starts with a flow (BugFixPool.SeedFlow, flow on Hosting/BugFix), which replaces the buckets:

  1. Triage — one round per bug. Essentials/Agent/bug-scope (modelLabels: [triage]) runs on auto:triage. Claude Opus 5.5 is the one model labelled triage, so Auto takes it without a call. It records diagnosis, bugScope, bugRepro, bugDifficulty (simple · standard · hard · hardest) and bugDifficultyReason on the item. Until a difficulty is recorded, every due round is a triage round.
  2. Fix — on the model the table names. Essentials/Agent/bug-triage (modelLabels: [coding]) runs on the rung the model table gives this kind and difficulty (above). Where the table says auto:coding, has no row, or names nothing served, the round runs on auto:coding as #2650 built it: Auto filters the EU-admitted models labelled coding, picks by difficulty (triage's rating is the task hint in every fix message), price and the ledger's calibration, and records its decision on the round (ThreadMessage.ModelRouting). The pool copies the model that ran and Auto's reason onto the item (bugChosenModel, bugAutoReason).
  3. Escalation. A fixer that ends blocked gets ONE more round on the same selection. Blocked again, the pool pins the next rung of the ladder above the chosen one: Kimi K2.7-code → GLM-5.3 (max) → Sonnet 5.5 (xhigh) → Opus 5.5 (xhigh), cheapest to strongest by price. Blocked twice on Opus, the work is PARKED as a person's (bugEscalationExhaustedAt) with the whole trail on the item and the pool page — it is never re-chosen and never re-driven.
  4. A tooling block is not a model failure. A fixer blocked ONLY because it cannot open a pull request (Plugins #2624) records bugBlockKind: no-pr-tooling. Such a bug is never escalated and never counted in the ledger. 🚨 Retired 2026-10-04: every fixer HAS the working tree (ReadCode / PushFix / OpenPullRequest, Plugins#2629), and the round messages say so instead of offering the claim. A no-pr-tooling park is read as the control-plane limit pr-tooling at generation 1, which the working tree lifted — re-driven ONCE on the same rung and stamped bugToolingLiftedAt; a fixer that records it again after that is a person's. The survey that day found 13 open bugs (~$196) parked on this claim or on prose saying the same.
  5. 🚨 A control-plane LIMIT parks the bug until the limit changes. A cap, a missing permission or a missing standard of the control plane is bugBlockKind: control-plane-limit, with bugBlockLimit naming it. The working tree's own gates stamp it on a refusal, with the limit's GENERATION (bugBlockLimitGeneration, from BugFixPool.ControlPlaneLimits: the file cap, the commit cap, a governance file, no GitHub App, a write GitHub refused the App); a fixer that meets one the control plane cannot see records it in words. Such a bug is never re-driven, never escalated, never counted against a model — in any pool, flow or not. The change that LIFTS a limit bumps its generation, and the next pass re-drives every bug parked on the older one ONCE, on the same rung, not counted as a block; a limit named in words is lifted by a person moving the bug. Why: until 2026-10-04 every block that was not no-pr-tooling read as a MODEL block, so MeshWeaver#5758 — refused by the PushFix file cap (Subscriptions.cs, 79 917 characters, over the then 60 000-character full-text cap) — climbed the ladder for 12 rounds and ~$70 against a limit none of the models could move, and #5555 (MessageService.cs, 227 113 characters) hit the same wall. The same change lifted that limit (generation 2 of pushfix-file-cap): reads deliver a large file whole in chunks, and PushFix changes it by span edits.
  6. 🚨 Waiting on a PERSON parks the bug until the person acts — never a timer. A fixer that only a person can move on (an approval, an owner's decision, a grant, an account setting) records bugPhase: blocked, bugBlockKind: awaiting-person, bugAwaitingAction (ONE exact question or action, for the person who owns it) and bugAwaitingNode (the node whose change IS that act — the governed request they approve, the decision node they write). The pool re-admits the bug only when that node changes AFTER the round that parked it ended (PersonActed), or when a person moves its phase; the next round resumes on the same rung. 🚨 A claimed artefact is read back: a park on a node that does not exist is a FAILED round — recorded in workLastFailure and un-parked so the next round writes it. Why: MeshWeaver.Crm#160 was re-driven 18 times (~$19) while waiting on a global admin's approval and two owner decisions, and its last round said the request "awaits approval" for a request it had never created. 🚨 The park is DELIVERED to the person (Plugins#3037). The fixer also records bugAwaitingPersons — the user ids of the people who must act, Admin standing for "a global admin" (none named: the platform operators' bell). Once the parking round has ended the pool sends each of them bugAwaitingAction through NotificationService.Dispatch — their bell, and email or Teams as their own notification rules decide — linked to bugAwaitingNode, and reminds them every 24 h while the park lasts (BugFixPool.AwaitingNoticesDue). Each notice is CLAIMED on the item's notified ledger inside one atomic update before it is sent (…:{person}:claimed@{at}), so two passes never send twice, and then SETTLED: a delivered notice becomes the receipt awaiting-person:{bugAwaitingSince}:{reminder}:{person}, a failed one is released and retried by the next pass, and a claim whose sender died expires after 10 minutes. A delayed pass never rolls the ledger back past a later reminder. Only a receipt for EVERY person counts as delivered. Why: Crm#160 then sat three days on an approval nobody had been asked for — notified: [] — while the stuck detector re-woke the coordinator, which can neither approve nor decide for anyone, 12+ times. Once every parked bug records its delivery the stuck detector routes limits/limit-park/awaiting-person to the persons (routedTo: persons) and the coordinator is not woken (Hosting/StuckDetection). 🚨 A governance activity is not a person. A park whose bugAwaitingNode IS a governance activity (Governance/Activities/{id} — an approval node below one is still a person's act) is re-classified as the control-plane limit governance-activity: escalated by the stuck detector as a system block, never delivered to a person, and resumed when the activity changes. Unless the activity itself waits on a person: the pool reads it first, and an activity Proposed with no review verdict under a standard that REQUIRES a review (it lists acceptance criteria — the pool reads the standard too; one with none, e.g. bug.verify, is admitted by the processor alone, so such an activity still in Proposed is the processor's), or Gating on a Signatures / Manual / ExternalSignature gate not yet green, is a real person-wait (the verdict and the signature are written onto the activity) and stays one, delivered as usual. An activity — or its standard — the pool could not read is not re-classified. Why: MeshWeaver#5936 "waited on a person" at an activity stuck in Gating with 0 gates armed — only the governance processor moves that.
  7. 🚨 A masked read is never written back. The model route's PII guardrail (an OpenRouter account setting, Plugins#2567) splices [PERSON_NAME]/[ADDRESS] into what the model READS; the repository text is clean. PushFix therefore refuses any write that would ADD such a placeholder to a file — an edit whose oldText carries one does not match the real file and is refused saying so; an edit or a full text that would write one is refused before anything is committed (a full text carrying one is checked against the file at the head). The fixer changes only spans it read unmasked, with edits. Lifting the guardrail itself is the account owner's setting (#2567).
  8. 🚨 The pool keeps starting work. The pass's budget (20 actions, one start) pays only for work that RUNS: a round with no variant to run on and a start with no free slot take nothing (BugFixPool.Admit). A bug WAITING on a merge, a deploy or a person (awaiting-*) holds no concurrency slot (HoldsASlot). A BLOCKED bug with no difficulty gets one more triage round, then it is a person's. Why: measured 2026-10-04 on the control instance — 67 bugs queued and no new bug thread for 11 hours: ~17 legacy blocked bugs (on variants the flow no longer serves) spent every pass's 20 actions finding no slot, the bug that held an auto slot (Plugins#2819, sev:M) was never reached, a bug waiting on a person's approval (Crm#160) held another, and so the bound of 3 never freed; blocked bugs with no difficulty had been re-triaged on every pass (MeshWeaver#1186: 21 rounds).
  9. 🚨 No absolute count — relative load only. Dispatch admits EVERY ready bug while the model pool has headroom; only a MEASURED exhaustion (a 429, a quota or credit refusal, a round cap, a throughput collapse — the pressure the review rounds record on Hosting/Triage/Status) bounds it, and then to the bug-fix lane's allowance of the measured pool. That is the ONE shared admission (ThreadDispatchPool.Allowance(Lanes.BugFix) → BugFixPool.BudgetOf for the pass, Admit per round; AI/AgentAdmission), not a rule of the pool's own. A variant's maxConcurrent and the flow's maxConcurrent no longer bound anything (BugFixPool.Assign). Why: maintainer, 2026-10-04, "why doesn't it dispatch more?" — 67 bugs queued behind an absolute bound of 3; the standing rule is "we can configure relative load but not absolute number", "until pool exhausted ⇒ unlimited, then up to 10% of pool".
  10. 🚨 A pull request closed WITHOUT MERGING re-opens its bug for triage — never silently Done. When the steward sees a closed event with merged: false (PullRequestIntake.Transition), every bug item recording that PR (pullRequest or prUrl) that has not finished is re-opened: bugPhase goes back to diagnosing and bugDifficulty is cleared, so the next round is a TRIAGE round; any park is lifted; the PR pointers (pullRequest, fixHeadSha, workPrHead, …) are cleared; and the closed PR is recorded on bugClosedPullRequests (url, head, when, the difficulty it was fixed at), with bugPrClosedUnmergedAt and a line on note. It is a NEW ATTEMPT with a fresh budget — the fresh-attempt subset of an issue re-open (TriageIntake.ReopenedFields: the exhaustion markers, the redrive count, the rung's rounds, a joined workLead, the last review), or the pool would record it re-opened and schedule nothing. A PR named in EITHER field matches. Both new members are declared on TriageItemContent — the live probe measured the typed read DROPPING them before they were. A bug that already finished (bugMergedAt, fixSha, or phase done) is never touched by a later close of an old PR. Each candidate is re-judged on its live content inside its own update, so a stale listing cannot re-open a bug that moved on. This is the CONSERVATIVE DEFAULT, and the maintainer may override it: a person who closed the PR because the bug is superseded or moot moves the bug's phase themselves. Why: Plugins#3039 — #5758's item kept pullRequest: #2847, bugPhase: blocked for a day after #2847 was closed as superseded, because the steward marked its own PR item Done and touched no bug item.

The rules and the tests that hold them (Hosting/BugFixPool/Test, Hosting/TriageItem/Test — run in each type's Tests area):

Rule Where Test that fails when the rule is removed
A control-plane limit parks the bug; it resumes once, on the same rung, when the limit's generation moves BugFixPool.Decide / LimitLifted / PlanRound BugFixPoolTests.AControlPlaneLimit_IsParked_UntilTheLimitChanges, BugFixActionTests.ALimitRefusal_ParksTheBug_NotTheModel
A refusal that is a limit stamps it on the item, with its generation; a landed push lifts it BugFixActions.RefusalVerdict / AtLimit / Unparked BugFixActionTests.ALimitRefusal_ParksTheBug_NotTheModel
Span edits change a file over the full-text cap, applied to the file at the merged head BugFixActions.EditedFile / Resolve BugFixActionTests.AFileOverTheCap_IsChangedBySpanEdits, BugFixActionTests.ASpanEdit_IsAppliedToTheMergedText
A span edit must match exactly once (overlaps count); CRLF kept BugFixActions.ApplyEdits BugFixActionTests.ASpanEdit_MustMatchExactlyOnce
A large file is read whole, in chunks; path#offset reads on BugFixActions.Bound / Chunks / ReadEntry BugFixActionTests.ALargeFile_IsReadWholeInChunks
A masked read is never written back (no write adds a scrubber token) BugFixActions.IntroducedTokens / EditedFile BugFixActionTests.AMaskedRead_IsNeverWrittenBack
Waiting on a person parks the bug; only the awaited node's change after the frozen baseline re-admits it BugFixPool.Decide / PersonActed / AwaitingBaseline BugFixPoolTests.AWaitOnAPerson_IsParked_UntilThePersonActs
A park on an artefact that does not exist is a failed round, recorded and un-parked BugFixPool.MissingArtefactPatch (read: AwaitedOf) BugFixPoolTests.AWaitOnAPerson_IsParked_UntilThePersonActs
The no-pr-tooling park is retired: re-driven once, then a person's BugFixPool.RetiredToolingBlock / ResumePatch BugFixPoolTests.TheRetiredToolingPark_IsRedrivenOnce
The pass's budget pays only for work that runs BugFixPool.Admit BugFixPoolTests.ThePool_KeepsStartingWork
A bug waiting on a merge, a deploy or a person holds no slot BugFixPool.HoldsASlot BugFixPoolTests.ThePool_KeepsStartingWork
A blocked triage gets one more round, then it is a person's; a lifted park resumes triage and is cleared BugFixPool.PlanRound(Flow, …) BugFixPoolTests.ThePool_KeepsStartingWork
No absolute count: no variant count bounds a bug; every ready bug starts while the pool has headroom; exhausted → the lane's allowance BugFixPool.Assign / BudgetOf (shared Allowance) BugFixPoolTests.ThePassBudget_IsTheSharedAdmissionsAllowance, BugFixPoolTests.TheBucket_IsStable_SpreadsBugs_AndHoldsTheBound, BugFixPoolTests.ReTargetsInOnePass_HoldTheBound

Each row was falsified on 2026-10-04: the rule was removed from the source and the named tests turned red (and only restored source is committed).

One rule was added after that falsification, with its own negative control (Plugins#3039 / #3107):

| Rule | Where | Test that fails when the rule is removed | |---|---|---| | A PR closed unmerged re-opens its bug for triage, the PR recorded — also when no steward item exists; a merged close, another PR, or a finished bug is untouched | PullRequestIntake.ReopenOnUnmergedClose / ReopenBugsOfClosedPull (Transition, both branches) | PullRequestIntakeTests.ABugWhosePrClosesUnmerged_IsReopenedForTriage (pure, Hosting/Deployment/Test); ClosedPullRequestProbe.AnUnmergedCloseReopensItsBug (live, the Deployment Tests area) | Not covered by a unit test (impure wiring, exercised only on a live mesh): the live read of the awaited node (AwaitedOf), and the placement of Admit inside the pass's Routine. A pass-level test belongs in src/MeshWeaver.Fleet.Control.Test, next to ABugWhoseRoundDiedWithItsReplicaIsReDrivenTest.

Rung / role Model node (created by its governed standard) Labels Effort on the wire EU providers $/M in · out
1 Provider/OpenRouterEU/kimi-k2.7-coding (bug-pool.model.kimi-coding) coding none sent — reasoning built in Inceptron, Nebius 0.67 · 3.35
2 Provider/OpenRouterEU/glm-5.3-coding (bug-pool.model.glm-coding) coding, review max (the fixer's xhigh, clamped up) Mistral (ZDR), Inceptron 1.40 · 4.40
3 Provider/OpenRouterEU/sonnet-5.5-coding (bug-pool.model.sonnet-coding) coding xhigh Vertex europe 2.20 · 11.00
4 (last, by escalation only) + triage Provider/OpenRouterEU/opus-5.5 (bug-pool.model.opus) triage, review — not coding xhigh Bedrock eu-west-1, Vertex europe 4.40 · 22.00

Effort per model. The fixer declares xhigh, and the wire clamps it to what each model accepts (OpenAIReasoningEffortWire, from the node's supportedReasoningEfforts or the built-in table): the nearest supported level at or above. GLM-5.3 (low · high · max) gets max. Sonnet and Opus 5.5 get xhigh. Kimi K2.7-code declares builtin, which means it reasons on its own, so no effort is sent at all.

Model nodes generated from a deployment record's ai.openRouterEU.models carry no labels, so these four are created through the sanctioned path for a model node on the control instance: one governed activity per standard, run as System, signed by a global admin who is not the proposer (as pr.review-model.bind was).

The outcome ledger — calibration for Auto

Each pass writes calibration on Hosting/BugFix: one row per (kind, model) with tasks, successes, total cost and median rounds (BugFixPool.OutcomesOf → LedgerRows).

The page shows the rows with success rate and cost per success. With Ai__Router__CalibrationNode: Hosting/BugFix on the instance, every replica's Auto router reads the same rows (NodeModelCalibration, a live query of that one node). The replica that runs the pool's pass is not the replica a round routes on, which is why the rows go through a node.

Per work kind. The same finished fixes are also written as workLedger: one row per (work kind, difficulty, model). A red main that went green without the worker's fix merging counts as neither a success nor a failure there, and a tooling block counts nothing, as above.

The brief, the stops and the review — the process around a worker (maintainer, 2026-10-01)

"First we triage how we want to bundle the issues together … then assess complexity, make detailed instructions including inlining of skills … then we choose a model … We must also detect when it is going in the wrong direction. Let a stronger model review and decide if it is up to quality or needs to be re-done by a stronger model. Do until you are in top-most tier."

Six steps. The scope round does the first three, the table the fourth (above), and WorkSupervision (Hosting/Deployment/Source/WorkSupervision.cs) the last two. All of it is ON for a flow pool unless its node says supervision: { enabled: false }.

1 · Bundle — one bundle, one thread, one branch, one pull request. A bug's first round is its scope round, and its first message offers the other open bugs of the same repository that have no pull request yet (BugFixPool.BundleOffer — paths and titles, the titles as data). The scope round names in workRelated the ones with the SAME root cause or the same feature. Before the lead's first fix round the pool joins them (JoinsOf, once — workBundledAt): each joined item records workLead and is not worked itself any more (no round, no review, its slot released); the lead records workBundle, and its brief names what its one pull request closes. The claim is honoured ONLY for what was offered — another repository's item, one that is already scoped (it has a fixer of its own), one with a pull request, a finished or parked one, a path the model made up is not joined — and never after the brief was given. A joined item is its own work again when it is reopened, or when its lead ended without a fix (parked, out of rounds, closed by hand: ReleasesOf), so no issue stays open with no worker. Red runs are bundled without a model, by episode (above).

2 · Complexity is the scope round's bugDifficulty with its reason (above).

3 · The brief (WorkSupervision.Brief) is the work package: the diagnosis and fix direction, the scope and the PATH LIST the fix may touch (workScopePaths), the reproduction, the acceptance check (workAcceptance), the bundle, the rules of the pool, the model and its round budget — and the procedure INLINED: the instructions of the skills the node lists in supervision.briefSkills (Essentials/Skill/bug-fix-thread by default), read through the index once per pass, so the fixer never loads them. A field the scope round left empty is said to be empty, never invented. 🚨 It is posted ONCE, in front of the first fix round (workBriefAt), never per round: a thread re-sends its whole history every round, so a brief repeated per round would be re-sent N times. Given once it is the long, unchanging head of the context that every later round and every redo is appended to — the order a provider's input cache needs. What is NOT built: the engine sends no explicit cache breakpoints, so whether a round pays the cached rate depends on the provider caching by itself — not measured.

5 · Supervise — a thread going the wrong way is stopped early. Each pass first RECORDS what it sees on the item (WorkSupervision.Signals, written only where it changed) and then judges the item alone (StopOf — pure, most specific first):

Stop The item shows Where the signal comes from
out-of-scope workPrOutOfScope — files changed outside workScopePaths the worker's pull request, read through the App (GET /pulls/{n}/files; a list that cannot be read in full is NOT read as "nothing outside")
ci-red-twice workCiRedHeads — its own checks red on two different heads GET /commits/{head}/check-runs; the newest run per name decides, so a re-run that went green clears its red
same-failure workLastFailure equals workPrevFailure the fixer records ONE line per round naming the test or gate that still fails; the pool moves it to workPrevFailure when it posts the next round
tool-loop workLoopCalls ≥ 6 the longest run of identical tool calls (name + arguments) in the last round's cell
diff-beyond-brief workPrLines over the difficulty's bound (300 · 1000 · 3000; hardest unbounded) the pull request's additions + deletions
over-budget workRungRounds ≥ workRoundBudget fix rounds posted in this attempt; a round posted while the work only waits for a merge or a deploy is not counted

Nothing stops work that is not being fixed (not scoped yet, awaiting-*, verifying, done). A stop posts no further fix round: it calls the review, with the reason (workStop, workStopWhy). An attempt on a stronger model starts clean: a redo works on the SAME pull request, so what the pull request shows (out-of-scope, diff-beyond-brief) is judged only after the attempt's second round and only on a head no reviewer has judged yet (workReviewedHead) — the old attempt's files are the stronger fixer's to repair, not a reason to stop it at once. The signals are judged one pass after they are recorded.

6 · A STRONGER model reviews; a redo is done BY that model and reviewed by the next stronger one; repeated to the top; then a person (maintainer, 2026-10-03: "implement and review by better model ⇒ if bad, redo by this model and review by even better model. repeat until converge"). A review is due when the fixer says its pull request is ready (bugPhase: awaiting-merge with a pullRequest) or the supervisor stopped the attempt — and it is taken BEFORE the cooldown, so "ready" is reviewed at the next pass, not after the hours an awaiting-merge item rests. The reviewer is Essentials/Agent/bug-review on the table's reviewer for that difficulty when it is a served rung ABOVE the fixer, else the next served rung up — always a stronger model, never the fixer's own (ReviewerFor; BugFixProcess.ReviewTierOffset = 1). It works in a thread of its own, …/_Thread/{work thread}-review-{attempt}: a fresh context that reads the item and the pull request, not the fixer's history. It records ONE verdict on the item (WorkSupervision.NextReview — pure, every case tested):

attempt 1 on Kimi ──ready──▶ GLM reviews ──redo──▶ attempt 2 on GLM ──ready──▶ Sonnet reviews ──redo──▶ attempt 3 on Sonnet
   ──ready──▶ Opus reviews ──redo──▶ attempt 4 on Opus ──ready──▶ a PERSON reviews (humanVerdict, by feedback)
   any ──accept──▶ the steward's review and a person's merge
The item shows The pool
workReview: accept records it (workAcceptedAt, and the head it covers: workAcceptedHead) and CONSUMES the verdict; the pull request goes on to the steward's review and a person's merge. The same head is not reviewed again; if the fixer pushes more and says ready again, that is a new attempt and it is reviewed
workReview: redo hands the work to the REVIEWER's model (the next served rung up only when that model is no longer a served rung above the fixer; a fix that ran off the ladder goes to the model that judged it; nothing served above → parked) — same thread, same pull request, the reviewer's findings verbatim in the message (RedoMessage); a new attempt with a fresh round count, whose reviewer is the NEXT stronger rung (workReviewer; none from the top); the model whose attempt was sent back is counted in the ledgers (workRedone)
ready (or stopped) on the TOP rung no model is stronger and none reviews its own work: PARKED for a PERSON's review (workParkKind: person-review, not a failure — workSettledRung = the top) with the whole trail on the item (workTrail: one line per attempt's end — model, stop, verdict, findings) and on the pool page. The person's verdict comes back as feedback (below). Never re-driven, whatever phase it was left in; its slot is released
redo with no stronger rung left (a stale pending review on the top) PARKED — the ladder failed (workParkKind: ladder-failed, settled past the top)
no verdict yet the fixer WAITS — no fix round while a review is pending; a reviewer at rest past the cooldown (or whose thread cannot be read past the round cap) is asked once more; asked twice with no verdict → parked for a person, never waved through
a stronger rung exists but none is served right now the attempt goes on unreviewed FOR NOW and says so (workReviewSkipped); nothing is consumed, so the review starts as soon as a reviewer is served. The steward's review and the person's merge still stand

A review is also taken before the item is named out of rounds (maxRedrives): work whose last round ended in "ready" is reviewed, not written off. A REOPENED issue is a new attempt — no earlier verdict or accepted head covers it, and a park is lifted: a person said "again".

The answer to the one question the design left open: the top tier is never reviewed by itself — its work goes to a person with the full trail; the same thing the ladder does when its top rung is blocked twice. Every verdict records where the work SETTLED (workSettledRung, workSettledBy, workSettledAt) — the learning record's input.

The learning record — step 3, fed by models AND people (maintainer, 2026-10-03)

"self learning process. can be also fed by humans."

Built ON the existing record (WorkLearning, Hosting/Deployment/Source/WorkLearning.cs — pure, every rule pinned and falsified by WorkLearningTests), never beside it:

The declared process — zero configuration, fluent overrides (maintainer, 2026-10-03)

"job must be configured properly. use fluent api. put all defaults without bothering user."

Every job — triage, fix, review, redo, learning — is declared in BugFixProcess (Hosting/Deployment/Source/BugFixProcess.cs) with its trigger events, its agent and the paths that agent needs installed, its queue and priority, its round budget, its retry cadence and its timeout; the process adds the ladder, the start rung per difficulty and the stop rules. BugFixProcess.Default is complete — a tenant sets nothing beyond the instance basics — and every default is pinned by BugFixProcessTests. The pool reads its seed, table, flow and supervision from it and nowhere else. An override is a fluent call a compiled host registers (BugFixProcess.Default.WithConcurrency(5).WithQueue("bugs")), or the same setting on the pool node. Queue and priority sit behind the declaration (Job.Queue, PriorityOf: a red main first, then by severity) and are recorded on the item (workQueue, workPriority), so the jobs move to the multi-queue model without touching the job logic. The step by step for a new estate is Hosting/BugFixProcessSetup.

On an existing estate, two things make the declaration run without an edit: a pool node from before the flow (variants only) is UPGRADED once to the declared flow, table and supervision (BugFixPool.UpgradePatch; opt out Hosting:BugFix:UpgradeToFlow=false), and every pass resolves the declared rungs onto the estate's own model nodes of the same model (ResolveOnEstate, by wire id — Provider/OpenRouterEU/opus-5.5 and a record-generated …/anthropic/claude-opus-5.5 are one model), pinning the manager job to the top rung where no node is labelled triage.

On the build queue — the target execution model (designed, NOT built)

"we must install the agent's paths for the execution. execution must be triggered by an event and then running in build queue on control instance." (maintainer, 2026-10-03)

Today a timer pass (BugFixPool.RunPass, every 10 minutes) finds each piece of work's due step and posts the round, bounded by the pool's concurrency. The target replaces the timer with the queue:

  1. The event enqueues the job. Each BugFixProcess.Events name (an issue classified, a red main, a feedback defect, triage's rating, "ready", a stop, a verdict, a person's verdict) writes ONE job node into the operations instance's queue (Hosting/Builds, beside the git builds), keyed by work item + job kind + attempt so a redelivery is the same job.
  2. The queue admits it by the rules it already has (BuildQueueLogic.Order / Admissions / Expired), with the job's declared queue and priority and a per-queue concurrency — and a budget gate: a job whose queue spent its daily budget waits, it does not fail.
  3. A runner claims it (the CAS Queued → Running the build runner uses), installs the job's InstallPaths where it runs, posts the round on the agent and the rung the declaration and the learning record give, and records the job's cost on it.
  4. Its end is the next event — the fixer's "ready" enqueues the review; the verdict enqueues the redo or the learning job.

What it needs that does not exist yet: a job KIND beside the git build in BuildContent, a runner for it on the operations instance, and the multi-queue model (first-class queue nodes with tiers and a default queue per tenant — designed separately). The supervision, the review, the learning record and the declaration above are unchanged by it: they are the job logic, already pure.

The PII guardrail on the EU route — the operator's choice, stated

A provider-side PII guardrail on the key behind the EU route masks names in every prompt — and can rewrite code identifiers too (measured: Blazor → [PERSON_NAME]). A fixer then reads masked code. Whether it is on is the operator's configuration of the account key; the process does not switch it off, and an estate that runs the fixer on a masked key should expect fixes that touch masked identifiers to fail review. The EU route now masks personal data in-process, precisely and with code intact (OpenRouter's EU route), so the provider's beta person-name and address presets can be switched off without losing that protection.

What is deliberately NOT done (yet)

Stated so the next session does not look for it.

The legacy variants

The table below is the bucket seed this flow REPLACED. A pool node created before 2026-10-01 keeps its variants until that node is edited; edits are system-only. The flow takes effect where a pool node starts fresh (the control instance at the cut-over).

The seed variants — EU only

id model residency takes bugs
opus-eu Provider/OpenRouterEU/anthropic/claude-opus-5.5 Eu (Bedrock eu-west-1, Vertex europe) once the EU route serves it on the estate
glm-eu Provider/OpenRouterEU/z-ai/glm-5.3 Eu (Inceptron, Mistral) once the EU route serves it on the estate
gpt-luna-eu Provider/OpenRouterEU/openai/gpt-6-luna Eu (Azure EU) once served; the light tier, here to measure whether a cheaper model fixes bugs

The instance switch. Hosting:BugFix:Enabled=false now stops the pool on that instance COMPLETELY: the re-drive sweep does not arm, AND the intake queues every new bug with that reason instead of starting a thread (BugFixPool.SwitchedOff). Before, the switch stopped only the sweep and the intake still started first rounds on the instance's own pool node. On the triage inbox the switch is never silent: the portal records a Disabled pass on Hosting/Triage/Status (outcome Disabled, the note naming the key) and logs it at Warning (BugFixPool.DisabledInbox), re-asserting that entry every Hosting:BugFix:Interval while the switch stays off; a Disabled entry moves lastBugPoolAt and the outcome but is not a pass, so bugPoolRuns stays frozen — on 2026-10-03 a pool switched off for two days still read the last pass's "Ok", and its only trace was an Information line the log store never kept. To move the pool to another instance, move the triage inbox with it (webhook targets and Hosting:PlatformWebhookSecret); switching it off on the inbox alone leaves it running nowhere.

All of them run at effort xhigh. The EU candidates are the EU-routed tiers of AI/ModelDataResidency. An estate whose EU route is not configured yet serves none of them, so its bugs stay queued — "queued — no EU variant served on this estate" — until it is.

Assignment and bounds

The rule is BugFixPool.Assign, applied when a bug gets its thread:

  1. Take a stable bucket of the thread path (FNV-1a) over the eligible variants in order. The same bug always lands on the same variant, and bugs spread evenly.
  2. A variant at its bound N (the bug threads it is WORKING, read from the items — a blocked or exhausted bug releases its slot) passes the bug to the next variant with a free slot.
  3. When every variant is full, the bug is queued on its item (bugQueuedAt, bugQueueReason), and the sweep starts it when a slot frees.
  4. With no eligible variant — none configured, none EU, none served — the bug is queued with "no EU variant served on this estate"; with the switch not On (Drain or Stop, see "The switch"), queued with "Dispatch is not starting new work on this estate (its mode is Drain or Stop)". It never runs on another model.

A Dispatch node that cannot be read queues the bug with the reason. A bug is never started outside the bounds.

The thread's composer carries the variant's agent, model and effort. Every later round runs on the same selection without naming it again. The item records bugVariant, bugModel, bugEffort and bugAssignment.

The re-drive — what makes it 24/7

BugFixPool.RunPass runs on the always-activated Hosting/PlatformBuilds hub of the triage inbox portal, every Hosting:BugFix:Interval (10 min; opt out with Hosting:BugFix:Enabled=false). Each pass reads every open bug item and its thread and applies BugFixPool.Decide:

Item / thread Step
verified, bugDone, bugPhase: done, or the issue closed on GitHub (read through the App) done
queued start its thread if a slot is free — one per pass: the bound is read from the index, which does not list a thread started this pass yet
bugPhase: blocked left for a person
maxRedrives rounds posted named for a person once (bugThreadError, bugExhaustedAt); a reopen gives the bug a fresh budget
a round running, or input already pending left alone
Executing with no activity for 45 minutes (past the cap) the round died: re-drive. The message waits in the pending input until the thread supervisor settles the dead round
at rest, inside the cooldown (the long one for awaiting-*) cooling
at rest, past the cooldown post the next round
due a round (either of the two above), but its variant is not an allowed, served EU variant re-target the thread to one that is, then post; none → no round, the item says "no EU variant served on this estate" (counted as waiting)

A crash or a roll therefore loses nothing. The item and the thread are in the store. A thread left Executing is re-driven. A queued bug is started by whichever replica runs the next pass. This is pinned on core's fault-injection harness by src/MeshWeaver.Fleet.Control.Test/ABugWhoseRoundDiedWithItsReplicaIsReDrivenTest.cs: the round runs on replica 0, replica 0 is killed, and the pass on replica 1 re-drives the bug. The negative control is a Decide that trusts every Executing round, which posts nothing.

The EU-only rule is pinned by src/MeshWeaver.Fleet.Control.Test/NoEuVariantServedDispatchesNoRoundTest.cs on a real mesh: with no EU variant served, a LIVE hand-over queues the bug and starts no thread, and a pass posts no round on a default-model thread. Its negative controls — the old fallback in Assign, an unchecked variant in RoundVariant — each turn it red.

The protocol the agent follows is in the next-round message and /bug-fix-thread. It records bugPhase (diagnosing · fixing · awaiting-merge · awaiting-deploy · verifying · blocked · done), diagnosis, pullRequest and fixSha on the item. Every field is typed on TriageItemContent, so the item's own hub never drops it.

The bug's working tree — how a fix reaches a pull request (Plugins#2624)

The agent runs on the control instance with no shell, no git and no token, so until this existed every bug thread stopped after the diagnosis (the threads of MeshWeaver#5932 and MeshWeaver.Plugins#2626 recorded blocked: "NO-GIT"). The working tree is now the bug's own issue item, in the same shape as every other triage action and as the PR babysitter's fixer (PrFixer): the agent patches what it wants read or written plus a requestedAction, and BugFixActions on the item's hub validates it and performs it through the GitHub App (systemorph-com) the control instance already holds.

Action What the control plane does
ReadCode reads fixReads (files, directories, "" = root, path#offset = a file from that character on) and an optional fixSearch (code search, always repo: the fix repository) at the bug's branch — or the default branch before it exists — into fixFiles: a file comes in chunks of 60 000 characters, each naming its characters, until the read's 240 000 are spent, so a file over the chunk size is seen WHOLE; where the budget cuts, the note names the path#offset that reads on
PushFix commits fixChange (≤ 10 files, each a full text, a deletion, or span edits — exact oldText → newText replacements) as ONE commit on the bug's branch: cut from the default branch's head on the first push, fast-forward only afterwards, and only over the head the change names (expectedHeadSha) — a moved head is refused as stale. Span edits are applied to the file AT the head the commit goes on, after the merge of the latest default branch: each must match exactly once there or nothing is pushed, so they never revert what the merge brought in (the overlap guard holds for full texts). A full text stays capped at 60 000 characters; an edited file at 1 000 000
OpenPullRequest opens the pull request from the bug's branch into the default branch, not a draft, with the control plane's footer (Refs the issue — never Fixes; Bug-Thread:; the merge rule); adopts one already open only when the App itself opened it; on a recorded one (fixPullNumber) reads it first and edits the title and body only when it is the App's own open pull request from this bug's branch. The whole POSTED text is held to the directive rule, not only the draft: the footer quotes the bug thread (which must be the path derived for this issue) and the variant and model only when they are plain identifiers

Isolation. Every bug works on its OWN branch, bugfix/{repo-slug}-{number}, derived from the issue — the key of its permanent thread — and never named by the agent. Two bugs never share a tree, and nothing is ever written to a default branch. fixRepository sends the fix to another repository of the triage scope (a Plugins issue whose root cause is in core).

The same gates a person's pull request passes. The pull request is the App's, but its code is the agent's, so the steward reviews it: PullRequestIntake.IsOwnAppWrite holds every OTHER App-authored PR out of review and admits a bugfix/ head (and PullRequestSweep sweeps it). Every required check runs on it; the internal review posts on it; the merge is dev.merge, signed by a person. Nothing in the working tree merges — there is no merge action, not a refused one — and nothing calls an admin bypass.

Refused before anything is sent: an item that is not a classified bug's (no bug thread — a stranger's issue is never given one), a finished bug, a fix repository outside the scope, a repository governance file (anything under .github/ — workflows, CODEOWNERS, dependabot, templates — and any CODEOWNERS: a person's), a file larger than a read returns whole (60 000 characters), a GitHub closing keyword followed by an issue reference in the agent's own commit message, title or body (Fixes #n would close the bug's issue at the merge, unverified — Refs, never Fixes), an unsafe path (.., .git, absolute, backslash), a conflict marker, a commit message or PR text addressed to an agent or a reviewer (the Crm#161 rule), more than 30 commits on one bug, and a request written by anyone but the bug thread (system, a Hosting/Triage/_Thread hub) or a global administrator (PullRequestActions.RequesterGate).

What it needs on the estate: the App installed on the repository with Contents: write and Pull requests: write. A refusal from GitHub is recorded on the item (error), naming the repository — the agent sets blocked and says so.

Not covered yet: the agent cannot build or run tests (CI is its build: it reads the PR's steward item, Hosting/Triage/pull-request/{repo-slug}-{number}, for lastChecks and the posted review), and it cannot answer a review thread on GitHub (Automatic review answered wants a person's reply). The git-data commit leaf mirrors PrFixer's; once that lands on main the two converge on one helper.

Guardrails, in every round's message and in the agent's instructions:

Measured: which variant works best

The pool stamps each measure the first time a pass sees it, so each is accurate to one pass interval:

The Comparison area of Hosting/BugFix shows one row per variant: whether it takes bugs and why not, open/N, bugs, diagnosed, PRs, merged, verified, reopened, needing a person, median hours to diagnosis and to a PR, rounds, tokens and review findings. It then shows one row per bug. It uses framework DataGrids with English and German labels. Not measured yet: CI reds on the fix's pull request, and cost in currency. Tokens are the cost measure.

The pool is counted on Hosting/Triage/Status: bugPoolRuns, bugPoolStarted, bugPoolRedriven, bugPoolOrphans, bugPoolExhausted and bugPoolAdopted (the backlog from before the hand-over) per day, plus bugPoolBacklog, lastBugPoolAt, lastBugPoolOutcome and lastBugPoolNote.

What the thread carries

Who does what

Role Human / agent Does Never
Role.BugTriage Essentials/Agent/bug-triage — the ONE agent on the ONE queue classifies the red, files the issue, opens the thread, diagnoses, writes the fix and opens the PR through the item's working tree (ReadCode / PushFix / OpenPullRequest), proposes merge / verify, conserves the finding merges, deploys, holds a token, closes without verification
Role.Reviewer the reviewer agent + a person for production accepts / declines dev.merge on the diff vs the diagnosis signs
the pool's reviewer Essentials/Agent/bug-review, always a stronger model than the fixer, in a thread of its own checks the fixer's result against the brief: accept, or redo with findings — its own model then takes the work over fixes, opens or merges a pull request
a person on the learning record anyone on a trusted instance gives verdict: accept|reject (and difficulty:) as feedback on a work item — it outweighs every model —
Role.Maintainer a human signs dev.merge reviews their own PR
Role.ProductionApprover a human, passkey signs the deploy —