Agent admission — one rule for every lane
The rule (policy
ci-two-pools-priority-queues): no limit. Every round is admitted at once, in every lane, whatever is running. The ONLY throttle is a live provider refusal — a 429 (rate limit) backs off that one (model, upstream) pool; a daily limit, credit or rejected key (402/403) closes that provider. A slow, stalled or round-capped round holds nothing back.
Every agent round in a process asks ONE admission before it starts: the round-admission pool
(ThreadDispatchPool, src/MeshWeaver.AI/Supervision/ — not the bug lane called Dispatch,
Hosting/BugFix, which is one of its clients). Its state is AdmissionState (immutable,
pure transitions) and its rule is LaneScheduler.Decide (pure). The lanes that start work — Dispatch
(the bug lane), the triage issue sweep, the review steward — read the same admission's Allowance(lane): null
(start everything) unless every pool the lane's work runs on is held back by a live refusal, then 0.
There is no number of concurrent rounds, no share, no floor and no weight anywhere in configuration:
what is configured is a declaration (alternatives, pins, lane overrides).
What #2874 had, and why it is gone. #2874 admitted everything while a pool had headroom and, once
a pool was measured exhausted, gave each lane a relative share of its estimate (floors 15/15/25/25/10 %,
borrowing by weighted fair share, a probability that fell as the pool filled, a daily cost guard and a
budget reserve spent by priority). "Exhausted" included a stalled stream, the round cap and a slow
review round, so a single slow round rationed every lane on that pool for 35 minutes. That ration is
removed; so are the review ledger's own share (ReviewAdmission, which bounded reviews to 25 % of the
measured pool) and the fleet coordinator's free share and lane floors (Hosting/FleetCoordinator).
laneFloors / laneWeights on Admin/Threads and Hosting:Triage:PullRequestReview:TargetLoad,
Hosting:Coordinator:LaneFloors are read by nothing; the review key is warned about.
The layers
1. Bulkheads — one per model pool
A bulkhead is the measured state of one model pool (the model a round runs on). Exhaustion is
measured per bulkhead, so an exhausted Opus pool never holds back work on GLM or Kimi. A bulkhead
carries its estimate K̂ (the concurrency it serves now), the carrying capacity K (the load at
which it last said "too much"), when and why, Thompson's success/failure counts, and the mean round
time and cost.
What holds a pool back — reported by the engine where it happens:
| Signal | Effect | Where |
|---|---|---|
A provider rate limit (429, too many requests) |
that ONE bulkhead backs off: BackoffBase (30 s) × 2^(refusals − 1), at most BackoffMax (15 min), counted since its last clean round |
ThreadExecution error path → ThreadDispatchPool.FailureOf → ReportFailure → AdmissionState.Exhausted |
| A daily limit, credit, a rejected or forbidden key (402/403) | every pool of that provider CLOSES (§8) | same |
A 403 naming redaction_context_lost |
only that request fails; the provider and key stay available to other rounds | ProviderHealthRule.Classify returns null, so FailureOf applies no pool pressure |
| The provider's daily budget reached at the platform's gate | the provider closes until UTC midnight | ThreadExecution budget gate → ReportClosed |
| A stalled stream, the round cap, a slow review round | nothing — shown on the page, holds nothing back | FailureOf returns null for them; ReviewAdmission.Measure names them without a pressure |
A clean round reports ReportClean (and resets the bulkhead's refusal count); every round's price
reports ReportCost (kept per lane per provider, shown only).
2. Seed-and-grow — the estimate of one pool (SHOWN, never consulted to admit)
The estimate is kept because it is a useful reading of what a pool served — the page shows it and
Plan reports it — but no admission reads it.
| Rule | What it does |
|---|---|
| Seed | A new pool starts at Seed = 2 rounds — a starting point, not a cap. |
| Slow start | While no carrying capacity is remembered and demand exceeds a utilized estimate (at least half running), K̂ doubles every SlowStartEvery (15 s) — 2, 4, 8, 16, 32 within a minute. A 429 answers within seconds, so that is evidence enough, and an idle pool does not crawl. Each clean round that ended at or above the estimate adds one more. |
| Multiplicative decrease | A refusal sets K = the load that ran and K̂ = Decrease × K (0.5) — once per burst (DecreaseHold, 30 s): thirteen 429s of one burst are one decrease. |
| Logistic growth | Inside the exhaustion window, at most every GrowEvery (1 min), (ClearAfter, 35 min = a round's whole life) K̂ grows logistically toward K: K̂ += r·K̂·(1 − K̂/K), never past it. |
| Recovery | Past the window K is forgotten and the pool probes upward again by doubling. |
| Spreading | A pool first used because its neighbour was full starts from a quarter of the neighbour's estimate instead of the bare seed. |
| Memory | The bulkheads are written to Admin/Threads and re-read on start: a restart does not forget what was measured, so a pool measured at 18 is used at 18 at once. |
3. Lanes — names, not shares
| Lane | Who |
|---|---|
express |
Blocking work: a red main (a bug-fix thread whose id starts ci-) |
babysitter |
The babysitter, the PR fixer, platform builds |
reviews |
Pull-request review rounds |
bugfix |
Bug threads and their fix rounds |
triage |
Issues, feedback, incidents |
interactive |
Everything else — a person's own chat |
A thread's lane is its Threads-app group (Lanes.Of, AI/ThreadGroups); laneOfAgent on
Admin/Threads overrides it per agent. The lane names where a round is counted on the page and which
Allowance a starter reads. It carries no floor, weight or share. Order between lanes is the job of the
queue the work was dispatched from (Queues: express > trunk > gate > pr), never
of this admission.
4. Aging inside a lane
Where a lane chooses WHICH item to start (Dispatch), urgency is
(1 + severity) × (1 + age / AgingUnit) (LaneScheduler.Urgency); Dispatch weights a red main
32, sev:B 16, H 8, M 4, L 2, unlabelled 1, with an aging unit of a day — so an eight-day-old
low bug comes before a fresh blocking one, and nothing waits for ever.
5. Starvation signal
A closed provider with work waiting (provider-closed, the top item) and every hard-pinned round
waiting longer than StarvedAfter (15 min) are listed under starving on Admin/Threads. The stuck
watchdog reads it as data.
6. Fallback to an alternative, and the hard pin
A round its own pool holds back (backing off or closed) moves to a declared equivalent model (modelAlternatives: same
tier, route rules respected — declared, never guessed) that admits it now. The same model on another
provider wins over a different model: when one of those admits the round it is chosen, and a different
model of the tier is chosen only when none does. Within that group, Thompson
sampling picks: each alternative's success rate is drawn from Beta(successes + 1, failures + 1),
discounted by its mean cost; an unexplored alternative draws from Beta(1, 1), so the choice keeps
exploring. The move is written onto the round's drained messages in the claim write, so the round
runs on the alternative.
A lane may exclude models from its fallbacks (LanePolicy.LaneExcludedModels, model-id
fragments, LanePolicy.AlternativesOf(model, lane)): the review lane never falls back to Kimi,
whether the alternative is a tier default or an explicit modelAlternatives entry — the gate review
is the GLM-5.3 + GPT-6 Sol ensemble, and a half moved onto Kimi would
post findings of a model the review was never measured with. Other lanes keep Kimi in the standard
tier (DefaultAlternativesTest.ReviewLane_NeverFallsBackToKimi_OtherLanesStillDo).
A hard pin (pins: a lane or an agent) disables the fallback: the round waits for its model,
the wait says hard-pinned to … on the pool's census row, and past StarvedAfter it is a
pinned starvation signal.
7. Fallback is ON by default — default alternatives per tier
"Very tolerant — always find an alternative", a hard pin the exception (maintainer, 2026-10-04).
With nothing configured, a model in one of the default tiers (LanePolicy.Equivalences) falls back
to the other models of its tier that the deployment's catalog actually has (every
LanguageModel under Provider/, read by the supervisor at start and again as the first step of every
sweep — ThreadSupervisor.RefreshCatalog → ThreadDispatchPool.SetCatalog; a failed read costs one sweep
and is logged, the pool keeping its last catalog, never a feed that dies on its first fault):
| Tier | Members (token order) |
|---|---|
| heavy | claude-opus-5.5 ↔ claude-sonnet-5.5 ↔ gpt-6-sol |
| standard | glm-5.3 (a token, so it also matches the configured variants -review, -coding, …) ↔ deepseek-v4-flash ↔ mistral-large |
| light | gpt-6-luna ↔ mistral-small ↔ gemini-3.5-flash-lite |
🚨 Kimi K2.7-Code is in NO tier. It is a dispatch model only — the fix ladder's rung 1 and the fleet
coordinator's hard pin, chosen by name — never a stand-in. It used to be GLM-5.3's first standard-tier
alternative, so every pull-request review (glm-5.3-review) whose GLM pool refused ran on Kimi; the review
gate is GLM-5.3 (GPT-6 Sol its ensemble partner), never Kimi.
Order (LaneScheduler.DefaultAlternatives): the same model on a DIFFERENT provider first (a
direct Anthropic or Azure Foundry model beside an OpenRouter route — a key's daily cap is per key, so
another provider is what keeps the lane alive), then the tier's other models in token order, own
provider before others, and the same model's own variants on the SAME provider last (a
glm-5.3-review round refused on OpenRouterEU tries deepseek-v4-flash before glm-5.3-coding there:
the variants share the refused key's pool and daily cap, so the refusal almost always holds them
too — pinned by DefaultAlternativesTest). The admission honours that order (§6: the same model
elsewhere wins whenever it admits the round — LaneScheduler.IsSameModelElsewhere). A model in no tier has no defaults; an explicit
modelAlternatives entry on Admin/Threads replaces its defaults. A hard pin still waits.
🇪🇺 A fallback never weakens what the round was promised (LaneScheduler.KeepsMarks, applied by
LanePolicy.AlternativesOf to the defaults AND to an explicit entry). It is decided on each model's OWN
node marks — dataResidency and dataRetention, the values the round's residency gate reads
(Where a model processes prompts) — never on a provider's name:
| Rule | Effect |
|---|---|
The instance's requirement (AI:RequiredDataResidency) binds every target |
on an EU-only instance a Global or Unknown model is never offered — the round would be refused before sending anyway |
| An established residency on the source is kept | an Eu round moves only to another Eu model: never the global router, never a maker's public API (api.anthropic.com is Global), never an undeclared Azure route (Unknown) |
| A declared retention on the source is kept | a ZDR round moves only to a ZDR model — not to an EU deployment under Azure's default abuse monitoring |
| A source with no established marks | constrains nothing beyond the instance's requirement |
The supervisor reads the catalog with each model's marks (CatalogModel) at start and every sweep
(ThreadSupervisor.RefreshCatalog), and hands it to the pool with the instance's requirement
(ThreadDispatchPool.SetCatalog); the requirement half is also kept live between sweeps — a
configuration reload that changes it re-hands the current catalog with the new value
(ThreadSupervisor.ResidencyWatch).
🚪 A pool first seen after its provider closed is born closed (AdmissionState.BulkheadOf): the limit
is the KEY's (§8), so an alternative on the same exhausted key is not admitted only to fail with the same
403. So when the key that closed is the only compliant route a tier has, the round waits — it never
leaves the route to keep running. That is the state of an instance whose only Eu/ZDR provider is one
OpenRouter key; the fallback only helps once a second compliant provider is in its catalog
(ModelDataResidency → A second compliant provider).
8. The provider's DAY — closed pools
On 2026-10-04 at 14:47Z the control instance's provider key hit its daily limit (HTTP 403: Key limit exceeded (daily limit)), and every lane stopped at once — reviews, triage, bug fixes,
embeddings. Two rules answer it:
- A refusal no load change cures CLOSES the provider (
AdmissionState.Closed,ThreadDispatchPool.FailureOf): the key's daily limit until the next UTC midnight; credit or a rejected/forbidden key forClosedFor(35 min); the provider's daily budget reached at the gate — until midnight. Every pool of that provider closes (the limit is the key's, not one model's). Nothing is admitted on a closed pool: each round moves to its declared alternative on another provider (within the declared, EU-respecting equivalences) or — hard-pinned — waits. A closed provider with work waiting is the top item ofstarving(provider-closed). A closed pool lets one probe round through everyProbeEvery(10 s) while nothing runs on it; a probe that ends cleanly reopens it — a topped-up key does not wait for midnight, and a refused probe costs a fast 403. A rate limit (429) backs off the one pool it hit and closes nothing; a stall or the round cap does nothing. 🔑 When the round's KEY is known, only that key closes (One key per function): each seat carries the credential it runs with (Seat.Credential— the key holder, and the function when its own key serves), and the refusal closes that credential (AdmissionState.ClosedCredentials) — its rounds wait or move to an alternative on another key, one probe at a time reopens it — while every other key's rounds go on starting on the same pools. With one shared key every lane carries the same credential, so nothing changes; with the credential unknown, every pool of the provider closes as described above. - The day's budget is READ and shown, never reserved: the budget gate's reading (spent / daily
limit) is reported on every round (
ReportBudget) to every pool of that provider and shown per bulkhead. #2874 stopped lanes by priority as the day neared its end (BudgetReserve); that reserve is removed with the share — what stops work is the provider's own refusal at the limit (a closure).
Measured the same day: from 16:31Z to ~20:01Z every review round of every pull-request head on the control instance ended within a second on that 403 (this PR's own seven rounds included) — the reviewer is bound to one model on one key, with no declared alternative to fail over to. Default alternatives per tier, including another provider first, are §7.
Not wired into admission yet: OpenRouter publishes a key's own limit and remaining credit (GET /api/v1/key:
limit, limit_remaining, usage_daily) — the Daily budget & health section now SHOWS it per
function key (One key per function), but admission reads the platform's own
ledger against the provider node's DailyLimit instead. A key whose limit is set only at the provider is therefore
learned from its first refusal (and closed at once), not ahead of it.
A waiting round says so ON ITS NODE — the pool is per process, the supervisor is per replica
Measured on the control instance (3 replicas), 2026-10-07 (Plugins#3062 and the wave
#3063–#3078): triage, PR-owner and review threads queued on the OpenRouter key closed by its daily
limit. Each thread was held — correctly — by the pool of the replica its hub lived on. But the
thread supervisor runs on EVERY replica and asked only its OWN pool
(PoolHolds): for the two replicas that did not host the thread, a queued thread with a queuedAt
marker read exactly like one left by a process that is gone, so they called it Parked. The trail of
Hosting/Triage/_Thread/triage-f9e721ce24b7 (Loki, replicas labelled A–D): created on replica A 07:13:38Z;
relaunched 1/2 by the supervisor on replica B 07:18:27Z (the recycle tore the hub down on replica A);
relaunched 2/2 by replica C 07:20:34Z — the fresh activation was refused on the same closed key and
re-stamped queuedAt at 07:20:35Z, the instant every report of the wave quotes as "last changed";
settled by replica D 07:23:15Z — its input answered by an Error cell reading "nothing ran far enough
to write one" and filed into triage, whose new triage thread queued on the same key in turn. A
relaunch cannot revive a thread whose key is spent: it can only re-queue it.
The rule now: a refused round writes the refusal's reason and its end onto the thread
(Thread.QueuedWhy, Thread.QueuedUntil, from PoolAdmission.Why/WaitUntil — a closed key's or
pool's reopening, a 429 backoff's end), next to queuedAt, and the claim clears all three. The
supervisor's Classify reads a future queuedUntil as Queued whichever replica asks — reported on
the queue page with what it waits for, never woken, relaunched or settled — and, past the end, counts
the parked bound FROM the end, and a park it then reports names the refusal it last met. A marker with
no end (a process that is gone) keeps the old reading. Pinned by
ThreadSupervisorMeshTest.AThreadWaitingOutAClosedKey_IsQueuedForAnotherReplica_NeverRelaunchedOrSettled
(a real refusal on the mesh's pool, a second supervisor on its own pool; its control — a marker with no
end — is woken by the same sweep; with the old Classify the subject is woken too) and
ThreadSupervisorClassifyTest.AQueuedThreadWithAFutureEnd_IsQueuedOnEveryReplica_NotParked /
PastTheEnd_TheParkedBoundCountsFromTheEnd_AndTheParkNamesTheRefusal.
What this does not change: a closed key still lets one probe round through every ProbeEvery;
a probe on a key that is still spent ends in the provider's 403, written on that thread's response cell
as an Error (Plugins#2980) — a named terminal state, not a park. How often a DAILY limit should be
probed is a separate decision this change does not make.
How often a spent DAY is probed is a SETTING, not code (MeshWeaver.Plugins#3076)
Whether a key closed by its daily limit is probed every 10 s or held until its reset is an open
policy question: each refused probe consumes one queued round as a 403 Error cell. The code does not
decide it. A closure remembers whether it is a daily limit (CredentialClosure.DailyLimit,
Bulkhead.ClosedByDailyLimit — set by FailureOf for the provider's key-limit 403 and by the
platform's own daily-budget refusal, and kept on the status node's bulkhead rows), and its probe
interval is LanePolicy.DailyLimitProbeEvery, declared on the status node as
dailyLimitProbeEverySeconds:
dailyLimitProbeEverySeconds on Admin/Threads |
behaviour |
|---|---|
| absent / not positive (the default) | today's behaviour: a daily limit is probed every 10 s, like any closure |
86400 |
hold until the reset: no probe goes through before the closure ends (a daily limit ends at the next UTC midnight) |
| anything between | one probe per that many seconds |
Every other closure (credit, a rejected key) keeps ProbeEvery. The status node is read live, so the
decision is one field write. Pinned by DailyLimitProbeIntervalTest — its hold case carries the
negative control (a credit closure under the same policy still probes after 10 s); with
ProbeEveryOf returning ProbeEvery unconditionally the hold case fails.
9. Delegated rounds
A sub-thread of a round that holds a slot ({parent}/…) is part of that round's work: it is
admitted at once on its parent's lane, even while its pool backs off — refusing it could only stall the parent (a parent holding a
slot while waiting on a child that waits for a slot is a nested-resource deadlock). It does NOT ride
the parent's slot: it is COUNTED as a running round of that lane, because the provider sees it, and it
is marked (Seat.DelegatedBy; delegated per lane × bulkhead on Admin/Threads) so a fan-out is
visible. Nothing in admission bounds the fan-out: a bound there would reintroduce the deadlock. A bound
belongs where the delegation is made (the delegate call refused visibly to the parent), not here.
10. The read model and the dispatch entry point (for a coordinator)
Admin/Threads carries, per bulkhead: capacity (K̂), carryingCapacity, exhausted,
exhaustedAt, why, slowStart, free, running, waiting, successes, failures,
meanSeconds, meanCostUsd; per lane × bulkhead: running, waiting (guaranteed 0, borrowed 0 and
bound null — the columns stay, there is no share), oldestWaitingSince, spentTodayUsd (on that bulkhead's provider), delegated; per bulkhead also closedUntil, budgetSpentUsd,
budgetLimitUsd, budgetAt; and starving (a closed provider first). A coordinator decides what to
start and how much: ThreadDispatchPool.Plan(lane, model, wanted) answers how many of wanted
rounds the admission would start now (all of them, unless a live refusal holds the pool; pins and
alternatives honoured, nothing seated). Enforcement stays in the admission: each round still asks.
The audit — every cap the lanes had (2026-10-04)
Read on origin/main fd3dbd1e1. (a) = a throughput throttle, replaced by the shared admission;
(b) = a safety bound, kept, with why.
| Where | Cap | Value | Class | Now |
|---|---|---|---|---|
src/MeshWeaver.AI/Supervision/ThreadDispatchPool.cs:34, ThreadSupervisorStatus.cs:21 (Admin/Threads) |
maxConcurrentAgents — agent rounds per process |
50 | (a) | Retired. Admission is per measured pool (this page). A configured value is warned about. |
Hosting/Deployment/Source/BugFixPool.cs:2083 |
MaxActionsPerPass — starts + re-drives per 10-min pass |
20 | (a) | Removed. BugFixPool.BudgetOf(Allowance(bugfix)): every due action with headroom, the lane's allowance once exhausted. |
Hosting/Deployment/Source/BugFixPool.cs:2088 |
MaxStartsPerPass |
1 | (a) — its reason was the per-variant bound reading a stale index | Removed. A start now counts on its variant at once (PassBudget.Move), so several starts in one pass cannot read one slot as free twice. |
Hosting/Deployment/Source/TriageIssueSweep.cs:73,91 |
Hosting:Triage:IssueSweep:MaxPerRun — items created/re-opened per hourly run |
20 | (a) | Retired. A run acts on every actionable issue with headroom, on Allowance(triage) once exhausted; a configured value is warned about. |
Hosting/Deployment/Source/ReviewAdmission.cs:77 |
review share of an exhausted pool | 10 % | relative already | Removed with the share: the ledger bounds nothing; review rounds are held only by the shared admission's live-refusal backoff (ReviewAdmission.Combine(PoolFree, live)). |
Hosting/Deployment/Source/BugFixPool.cs:69, BugFixProcess.cs:173 |
per-variant maxConcurrent |
2 / 3 | (a) | Not changed here — removed by the bug-pool change in flight; it adopts BudgetOf/Allowance. |
Hosting/Deployment/Source/PrFixer.cs:73 |
MaxFixesPerDay — fix COMMITS pushed fleet-wide per UTC day |
20 | (b) | Kept: a blast-radius guard on automated pushes to other people's branches, not a model-pool throttle (no model call is saved by it). |
Hosting/Deployment/Source/PrFixer.cs:81 |
MaxFixThreadsPerPass |
1 | (b) | Kept: correctness — every fixer answers on one page list that a patch replaces; two answers before the claim would lose one. |
Hosting/Deployment/Source/PrFixer.cs:70 |
MaxAttemptsPerHead |
2 | (b) | Kept: once-per-head loop guard. |
Hosting/Deployment/Source/PrBabysitter.cs:1512 |
MaxProposalsPerHandoff |
20 | (b) | Kept: the size of ONE validator thread's task, not how many threads run. |
Hosting/Deployment/Source/PrBabysitter.cs:2318,2324 |
MaxLogTails, MaxChangedFiles |
2 / 60 | (b) | Kept: evidence size per proposal. |
Hosting/Deployment/Source/BugFixPool.cs:1673,1676 |
MaxAdoptionsPerPass, MaxAdoptionReadsPerPass (pre-hand-over backlog) |
1 / 5 | (b) | Kept: GitHub reads of a legacy backlog; a migration trickle, not agent throughput. |
Hosting/Deployment/Source/PullRequestSweep.cs:59 |
MaxPerRun — kicks per run only while the slot ledger cannot be read |
20 | (b) | Kept: the fail-safe of an unreadable ledger; with a readable ledger the kick budget is the admission's. |
src/MeshWeaver.AI/Supervision/ThreadSupervisorStatus.cs |
MaxRetries 2, LookbackLimit 500 |
(b) | Kept: relaunch loop guard; scan page size. | |
PullRequestIntake.MaxReviewAttempts, ReviewAdmission.MaxOutageRounds |
3 / 6 | (b) | Kept: once-per-head loop guards. | |
Hosting/Build/Source/BuildQueueLogic.cs:24 |
Hosting:Builds:Concurrency — CI heavy legs admitted |
3 | different resource | Not the model pool: GitHub runner capacity. Left as it is; a runner-pool admission is its own change. |
Hosting/Queue/Source/QueueLogic.cs |
the job queue's measured Bound (AIMD, TargetUtilization 0.1) |
relative | relative already | Off by default for the agents queue (Bounds false — "no limit"); it holds a job back only for a live refusal of its provider key. |
src/MeshWeaver.AI/Supervision/LaneScheduler.cs (#2874) |
lane floors, weights, the admission probability, the cost guard, the budget reserve — applied once a pool was measured exhausted | relative | (a) | Removed (policy ci-two-pools-priority-queues): the only throttle is a live refusal's backoff per bulkhead. |
Hosting/Coordination/Source/CoordinatorPolicy.cs |
LoadView.Bound/Free — the coordinator's share of an exhausted pool; lane floors |
relative | (a) | Removed: it dispatches everything ready; a refused provider's work moves or (pinned) waits. |
The Monte Carlo comparison — retired with the share
#2874 chose seed-and-grow over "step" (everything admitted until a refusal, then a tenth of the load)
with a Monte Carlo test over a provider model fitted to the control instance's measured rates. That test
(AgentAdmissionSimulationTest) compared two RATIONS; with no ration left it measured nothing that
ships, so it is deleted rather than kept green. Its provider model is the reason to watch the first
burst after this change: it rewarded holding load back (fewer 429s, fewer rounds cut at the cap). Now
the backoff is the only answer to a burst, so a burst costs its 429s and each refused pool waits out
its backoff — measured, not simulated, on the page's refusals per bulkhead.
What this does not establish
- Multi-pod. The pool is per process; on a multi-silo mesh every silo admits its own rounds and
Admin/Threadsshows whichever silo wrote last. The review steward's ledger stays the cross-pod record for reviews. A thread's OWN wait is cross-pod: it is on its node (queuedWhy/queuedUntil, §8). A closure is not: a replica that never met the 403 does not know the key is spent until its own round meets it. - The provider model's shape beyond the three measured points (1, ~5 and ~20 rounds in flight).
- Equivalence beyond the default tiers. The three default tiers name today's models; a new model
joins a tier only when its id carries a tier token or
modelAlternativesnames it. - The burst behaviour without a ration. Removing the share was decided, not simulated: how many 429s and cut rounds a 35-head burst now costs is a production measurement still to take.
- Upstream granularity. A bulkhead is keyed by the model a round ran on. Where one model node is pinned to one upstream (the review models), that is the (model, upstream) pair; where a model node lets the router pick the upstream, a 429 backs off the whole model node.
- Cost guard inputs. The guard reads each lane's priced spend; the provider's daily limit is
applied by
ProviderBudgetGuarditself (a reached budget is an exhaustion signal here).
The tests — one per rule
Each test below fails when its rule is removed — checked by mutating the rule and running the suite: rationing an exhausted pool again turns the negative controls red, dropping the backoff turns the backoff tests red, and keeping the refusal count past a clean round turns the reset test red.
| Rule | Test |
|---|---|
| 🚨 No limit — an "exhausted" pool with no refusal is NOT rationed (negative control) | LaneSchedulerTest.NoLimit_AnExhaustedPoolWithNoRefusal_IsNotRationed, AdmissionStateTest.ExhaustedPool_WithNoRefusal_IsNotRationed, ReviewAdmissionTests.NoLimit_EveryHeadIsAdmitted_AnExhaustedLedgerRationsNothing, AnExhaustedReviewLedgerHoldsNoHeadBackTest, FleetCoordinatorTests.AnExhaustedPool_DispatchesEverything |
| A 429 backs off ONLY its pool, doubling, reset by a clean round | LaneSchedulerTest.Refusal_BacksOffOnlyItsPool_DoublingUntilAClearRound, AdmissionStateTest.Refusal_BacksOff_ThenEverythingStarts, AdmissionStateTest.PerPoolRefusal_Isolates |
| A flood holds nothing back | AdmissionStateTest.Flood_NothingIsHeldBackForLoad, AdmissionStateTest.HealthyPool_EveryReadyRoundStarts (all six lanes), AdmissionStateTest.HealthyPool_EveryLaneAtOnce_EverythingStarts |
| Stall, round cap and slow rounds are not refusals | AdmissionStateTest.Failures_RateLimitIsLoad_KeyAndCreditClose_OurOwnFaultIsNothing, ReviewAdmissionTests.ARound_IsARefusalOnlyOnALiveProviderRefusal |
| A closed pool refuses but probes | LaneSchedulerTest.Closed_RefusesButProbes, AdmissionStateTest.KeyLimit_ClosesTheProvider_AndEveryLaneFailsOverInsteadOfStopping |
| The day's budget is shown, never reserved | AdmissionStateTest.Budget_IsReportedOnly_NothingStopsBeforeTheProvidersRefusal, AdmissionStateTest.Spend_IsKeptPerLanePerProvider_AndHoldsNothingBack |
| Fallback + Thompson | LaneSchedulerTest.Fallback_ThompsonPrefersWhatWorksAndKeepsExploring, AdmissionStateTest.Fallback_ARefusedRoundMovesToAnAlternative |
| Hard pin | LaneSchedulerTest.Pins_NameALaneOrAnAgent, AdmissionStateTest.HardPin_WaitsVisiblyAndSignalsTheWatchdog |
| Seed-and-grow (the shown estimate) | LaneSchedulerTest.SeedAndGrow_*, LaneSchedulerTest.Spreading_ANeighbourSeedsFromAFullPool |
| Aging inside a lane | LaneSchedulerTest.Aging_AnOldLowSeverityItemComesFirstEventually, BugFixPoolTests.TheQueueOrder_IsSeverityTimesAge |
| Delegated rounds | AdmissionStateTest.Delegated_ASubThreadOfARunningRoundStartsOnAFullPool |
| Default alternatives per tier (catalog only, other provider first, explicit replaces) | DefaultAlternativesTest.* (13 tests, control's catalog plus a second provider in its compliant and non-compliant shapes) |
| A fallback keeps residency, retention and the instance requirement — explicit entries too | DefaultAlternativesTest.KeepsMarks_ResidencyRetentionAndTheInstanceRequirement, …OnlyTheCatalog_NoTierNoDefaults_ExplicitReplaces_AndIsHeldToTheMarks |
| A closed key fails over to the compliant provider, never a global one; a pin still waits | DefaultAlternativesTest.ByDefault_AClosedKeyFailsOverToTheCompliantProvider_NeverAGlobalOne_AndAPinStillWaits |
| No compliant second provider ⇒ the round waits, never leaves the route | DefaultAlternativesTest.WithoutACompliantSecondProvider_TheRoundWaits_ItNeverLeavesTheRoute |
| A new pool on a closed key is born closed | DefaultAlternativesTest.ANewPoolOnAClosedKey_IsBornClosed_SoTheRoundWaitsInsteadOfFailingAgain |
| The same model on another compliant provider wins on every draw; a Global provider is never chosen under EU residency | DefaultAlternativesTest.AClosedKey_MovesToTheSameModelOnAnotherCompliantProvider_OnEveryDraw, DefaultAlternativesTest.UnderEuResidency_AGlobalProviderIsNeverChosen_EvenForTheSameModel |
| Lanes from thread groups | LaneSchedulerTest.Lanes_FollowTheThreadGroup |
| A late signal finds its pool | AdmissionStateTest.Signals_ALateSignalFindsItsPool |
| Cost per lane | AdmissionStateTest.Cost_CountsAgainstTheLanesDay |
| The process pool on a mesh | ThreadSupervisorMeshTest |
| 🚨 A wait with a known end is Queued on EVERY replica — never relaunched or settled (negative control: a marker with no end is woken) | ThreadSupervisorMeshTest.AThreadWaitingOutAClosedKey_IsQueuedForAnotherReplica_NeverRelaunchedOrSettled, ThreadSupervisorClassifyTest.AQueuedThreadWithAFutureEnd_IsQueuedOnEveryReplica_NotParked, ThreadSupervisorClassifyTest.PastTheEnd_TheParkedBoundCountsFromTheEnd_AndTheParkNamesTheRefusal |
| Dispatch pass budget | BugFixPoolTests.ThePassBudget_IsTheSharedAdmissionsAllowance |
| Triage sweep count retired | TriageIssueSweepTests.TheCadence_IsConfiguration_AndClamped |
| The review page reads unlimited, naming a recent refusal | TriageStatusTests.TheReviewAdmission_ReadsAsTypedContent_AndShowsTheBound |
Where the code is
src/MeshWeaver.AI/Supervision/LaneScheduler.cs— the lanes, the policy, every rule (pure).src/MeshWeaver.AI/Supervision/AdmissionState.cs— the pool's state and its transitions (pure).src/MeshWeaver.AI/Supervision/ThreadDispatchPool.cs— the live pool: admission, signals, clock, census,Allowance,Plan.src/MeshWeaver.AI/ThreadSubmission.cs— the watcher asksAdmitwith the thread's lane and writes a fallback model onto the round.src/MeshWeaver.AI/ThreadExecution.cs,ProviderBudget/ProviderBudgetGuard.cs— the signals.Hosting/Deployment/Source/{BugFixPool,TriageIssueSweep,ReviewAdmission}.cs— the lanes that start work readAllowance(null unless a live refusal holds every pool of the lane).Hosting/Coordination/Source/CoordinatorPolicy.cs— the fleet coordinator dispatches everything ready; a refused provider's work moves.
Related: Review round capacity · Queues · The activity execution model · Thread groups.