The portal masks what it hashes
MeshWeaver.Plugins#2218. Two defects in the incident-to-ticket pipeline, found from the same
measurement: on 2026-09-21 Systemorph/MeshWeaver had 94 open issues opened by
systemorph-com[bot], all minted between 2026-09-19T11:44Z and 2026-09-21T09:00Z, and 42 of them
were one log site. None of the 94 carried a severity when it was opened.
1. A fold that could never see the fix
What was wrong
LogIncidentIdentityResolution exists because the identity function used to
live in mw-log-watcher, a separately shipped image, so a corrected identity reached production
only when somebody rolled it — measured five weeks and four revisions behind. #1796 moved the
function into the portal, which rolls continuously.
It moved the function and not the function's input. The identity is computed over
report.NormalizedMessage and report.NormalizedDetail, and both are produced by
LogLineParser.Normalize inside the watcher. So every masking revision since — the episode
stamp and the 8-character activation id (#2165), masking by SLOT rather than by what a token looks
like (#2167, #2168), a list's length (MeshWeaver#2157), a nodeType: term (MeshWeaver#3545) —
shipped in MeshWeaver.Observability.Contract, was present in every portal, and was dead code on
the ingest path. The same defect #1796 closed, one layer down.
The measurement
Admin/_LogIncident/54ecfdaea23fd110 on the control instance stores
[ROUTE] Routing back-pressure [e9b2{n}fb#{n} started {time}]: {n} route dispatches in flight …
The activation id is e9b212fb. It was masked only where it happened to contain digits —
e9b2{n}fb — so every grain activation normalized to a different string, hashed to a different
fingerprint, opened a different incident node and produced a different GitHub issue. Three facts,
each reproducible:
| The portal's identity function over the watcher's stored text | 54ecfdaea23fd110 — exactly the live node id |
| The same function over the same bursts' RAW lines, re-normalized here | 96c85f04cadba7f7 for 42 of the live fingerprints |
| Live fingerprints whose own samples agree here / distinct identities they resolve to | 70 → 26 (the other 16 still hold samples that disagree — pre-existing masking gaps, unchanged) |
So the identity function was live and correct; its input was five weeks stale. The agent triaging those incidents had already written the diagnosis into its own tickets — "the fingerprint appears to split per episode stamp, so every crossing opens a new ticket" — which is what the ticket flood looks like from inside it.
The fix
The report already carries its own raw evidence: LogSample.Line is the whole verbatim burst,
header and continuation lines, exactly as BurstAggregator.Aggregate joined it. The portal
re-parses that with the LogLineParser it ships and hashes its own normalization.
- The reported text stays the fallback. A report whose evidence does not parse — a watcher old enough to send no samples, a sample truncated past its header, a bodyless burst (#2222) — resolves exactly as before. The change can therefore only ever FOLD; it can never split something that was already folded.
- A vote, not sample zero. One report's samples are one reporter group, and under the portal's newer masking they can still disagree: over the 86 live incidents, sixteen reports held samples that resolve to more than one identity here, and in the clearest of them nine of ten samples agreed while a single truncated one parsed as a bodyless burst. The majority decides, and a tie breaks on the ordinal-lowest token, so the answer is a function of the report and never of enumeration order.
- The shape ledger reads the same evidence.
ShapeKeynow goes throughLogIncidentIdentityResolution.SplitIdentity, so a bucket's rows and the node they sit on can never be computed over different text.
Nothing in the watcher changes, and nothing needs to be rolled for the fix to take effect — which is the property the whole arrangement exists for.
What happens to the 42
Nothing here closes them; the corpus migration is what carries a re-addressed incident's ticket, counts and history onto its successor, keyed on the reporter's own fingerprint. The next crossing of that site lands on the folded identity and brings one of them with it; the rest are a backlog a human closes as duplicates.
2. No automated bug was ever inside the release gate
What was wrong
The issue taxonomy puts a severity on every bug and makes two of the four a release gate: sev:B
and sev:H must both be zero to cut a release, and a bug carrying neither "has not been triaged,
and working it ahead of a labelled sev:H is choosing by accident."
The pipeline had never assigned one. All 94 open bot-filed issues were opened with bug and
nothing else; the 23 that carry a severity were labelled hours or days later by a person — issue
#4824 got bug from the App at 2026-09-19T11:51:05Z and sev:M from a human at
2026-09-20T08:18:03Z. The LogTriage agent's instructions spelled the label set out literally as
["bug"], so there was nothing for the model to get wrong and nothing to fail. 71 tickets sat in a
queue no gate reads.
The fix, in two halves
- The agent judges it.
LogTriageis now given the four classes, what each means, and the warning that the two blocking ones stop a release being cut — so reaching for one to signal urgency is not available. - The filer guarantees it. Asking a model more nicely is not a guarantee; a rule applied where
the issue is actually created is.
LogIncidentFiler.WithSeverityadds the default when the draft proposed none, and the ticket's evidence footer says the severity was assigned by default, so the first person to read it knows a classification is still owed.
The default is deliberately sev:M and deliberately not a blocking class. A machine that did not
judge must not be able to hold the release line — one noisy log site would stop every cut — and it
must equally not be sev:L, which parks a real defect in a backlog nobody reads. A draft that
proposes several severities keeps all of them: choosing between two labels a human can see is a
judgement this code has no basis for, and silently dropping one would hide that triage was unsure.
One legacy incident, fifty-two successors
The corpus migration records the same finding from the other side, and it is the cleanest statement
of it. When the portal's identity re-addresses an incident, LogIncidentCorpusMigration carries the
legacy node forward and every successor's ticket says which legacy id it inherited. 52 of the 94
open bot-filed issues name the same one — e4a97855ab595beb.
That id is the REPORTER's fingerprint for the whole RoutingGrain back-pressure site: the watcher,
with all its masking flaws, still called it ONE fault. The portal then split it 52 ways. So the
split is provably not in the reporter's grouping and not in the identity function — both agree the
site is one thing — it is in the TEXT the portal hashed, which is the reporter's and stale.
It also explains a sentence those tickets carry that reads like a contradiction: "superseded and
will not fold, file or comment again". That is true, and it is about the LEGACY node, which is
correctly retired. Fold's alreadyMoved guard means only the FIRST successor inherits the ticket
and the counts; every later one gets provenance only — and then files an issue of its own. 52
tickets carrying that sentence is the defect's signature, not a second defect.
The same id, masked eleven different ways
One listing of Admin/_LogIncident is the whole finding, because the node NAME carries the
normalized text. These are all the same 8-hex activation id at the same log site, as the stale
watcher masked it:
[{n}f5{n}ea#{n} [b5{n}d3{n}e#{n} [df5d5f6{n}#{n} [{n}a3d7#{n} [{n}b2c7c6c#{n}
[aae3{n}b#{n} [{n}c7f0ee#{n} [{n}b0{n}bd1#{n} [{n}c2{n}#{n} [{n}d6fde0#{n}
[{n}b6{n}d4{n}f#{n} [{n}d8{n}e3#{n} [{n}d#{n} [df8ca0{n}b#{n} [e9b2{n}fb#{n}
The bare-number rule ate whichever runs of digits the id happened to contain and left its hex
letters standing as literals, so the masked text is a function of the id — which is exactly what
masking exists to prevent. EpisodeStamp and MixedHexId (#2165, #2167) mask the whole token; they
have been in this assembly since 2026-09-19 and had never once reached the hash.
The flood stopped for the wrong reason
Incident DETECTION on the control instance stopped around 2026-09-20T22:30Z and had not resumed 15
hours later: content.lastSeen:2026-09-21* returns 0, not truncated, over
coverage.partitions: ["admin"], while content.lastSeen:2026-09-20T2* is still truncated at 50.
POST /api/log-incidents is the ONLY way a burst enters the portal, so with the watcher silent no
report arrives, no incident is minted and no File is requested — which is why the last tickets
(#5058–#5063) were filed from incidents detected BEFORE the stop and the flood appeared to have
ended. Nothing was fixed by that outage. This fix is at ingest, so it applies to the first
report after the watcher is alive again — alive, not rolled, which is the entire point of it.
And it came back, still splitting. The same query at 2026-09-21T15:41Z returned ≥10 (truncated),
with incidents whose lastSeen is 15:40:49Z — so detection resumed some time between 13:31Z and
15:41Z. The first RoutingGrain nodes minted after it resumed are
0fc7665eaf6aae43 ([cf0fc5ed#{n} …), 18104cb0dbf5d912 ([d4cae9e3#{n} …),
0fda5918d57c6050 ([{n}f7{n}a#{n} …), 3708a213a82ec7b0 ([{n}c9{n}d#{n} …) and
4db1545b366d9574 ([a6{n}b4{n}#{n} …): five activation ids, five incidents, one log site, masked
five different ways. The defect is not historical and the flood is resuming as this is written.
🚨 And the watcher's liveness is unobservable from the portal by construction: it has no
Hosting/Deployment record (namespace:Deployments scope:descendants returns only the portal instances'
records), so nothing reports its image, samples it or can roll it — and every signal
that it has stopped is emitted BY it, including log-pipeline-behind-* and the silent-window
finding whose job is to announce exactly this. A missing self-finding is evidence FOR death, never
for health. The filer does NOT share that shape — it is a BackgroundService inside the portal, so
it rolls with the portal and is covered by the portal's own record and /health.
3. The back-pressure line's oldest-leg label: two more instance slots (MeshWeaver#5271)
Core's back-pressure line gained a new discriminator — "oldest leg in flight N ms — dispatch → <address> (delivery <id>)" — and both values in it are instances. Neither was masked:
- the delivery id (
nwzA1tz0skCOOmSyJTsSbQ) is base64url — not a guid, not hex, no slash, and it may start lower-case, soSpacedSubject's uppercase guard could not take it either. Every crossing therefore hashed to a new fingerprint. The episode stamp, which the filed tickets blamed, was already masked byEpisodeStamp. - a single-segment destination (
dispatch → Essentials,→ ThreeBody) has no slash forPath.
DeliveryId masks the id inside the line's parenthesised (delivery …) field only (so prose and a message type after a bare "delivery
failed" survive); LegDestination masks whatever follows dispatch → / stream-routed →, and
keeps the leg KIND, because the two are different routing branches.
Measured over the 146 back-pressure sample lines carried by the open core issues filed between
2026-09-19 and 2026-09-22 (43 distinct filed fingerprints): the identity function over the new
normalization yields 3 — one per leg kind on the current wording, and one for the pre-#5151
wording (deepest per-destination queue), which is a different log template and rightly its own.
Pinned in LogLineParserTest (Fingerprint_FoldsBackPressureCrossingsThatDifferOnlyInTheOldestLegsDeliveryId,
Normalize_FoldsTheOldestLegsDestination, and the two must-not-fold controls
Normalize_DoesNotMaskDeliveryOutsideItsSlot, Normalize_KeepsMissingHandlerFaultsForDifferentTypesApart and Normalize_KeepsTheLegKindApart).
This changes the fingerprint of every current back-pressure incident, so each re-addresses once
through LogIncidentCorpusMigration onto its successor; it does not close any existing issue.
4. A cut sample keyed on where the cut fell (MeshWeaver#5271, second half)
Section 3's masking was correct, reached the control instance's image (b922f070 is an ancestor of
the plugins sha the running 3.0.0-ci.9218 image was built from), and the flood went on: on
2026-09-23 the control instance minted more than forty distinct Routing back-pressure
incidents in one day. The splitter was not in the masking at all.
The portal re-derives the identity from LogSample.Line, and the watcher cuts every sample at
MaxSampleLength (2000) characters, appending …[truncated]. The back-pressure line is ~3,400
characters, so no sample of it is ever whole. The cut lands at a fixed RAW offset, but the fields in
front of it — the oldest leg's destination, the latest dispatch target — vary in length per crossing,
so it falls at a different word of the constant tail each time (AT THE CROSSIN…, AT THE CR…), and
the identity hashed a different string for every length combination. Admin/_LogIncident/78316ad8e1d82122
shows both halves at once: nine crossings from five activations folded — the ones whose cut happened
to land on the same character — while its siblings were minted beside it.
Both earlier measurements missed it for the same reason: they removed the variable. The #2286
measurement cut each line "at the constant tail" by hand before normalizing, and
PortalNormalizesWhatItHashesTest aggregates with maxSampleLength: 4000.
The fix is a statement about evidence, not a bound to widen. The identity reads the first
StructuralLogIncidentIdentity.DetailBound (1024) characters of the masked detail — a prefix every
sample of a long line keeps — and LogIncidentIdentityResolution.ReparsedBursts counts a cut sample
only when the cut (the marker is left in the line when it is parsed, so it lands in whichever field
the cut ran through) falls at least 128 characters past that bound. A cut through a stack trace
leaves the fault's text whole and still counts; a cut inside the bound cannot vouch for the prefix
and is not counted. The top frame is checked the same way: it is taken again from the sample's
COMPLETE lines (the line the cut ran through is partial) and must agree, and a cut sample carrying an
exception but no complete application frame is not counted, because the first one may lie past the
cut and the identity would switch from the frame branch to the site branch. A report left with
nothing countable resolves on its reported fields, as a report with no evidence always has. A message shorter than the bound hashes exactly as before; one
longer re-addresses once through LogIncidentCorpusMigration, like section 3.
Pinned in ATruncatedSampleKeysOnWhatItKeptTest: four dispatch crossings with the values the live
samples carry, cut at the watcher's own default MaxSampleLength (read from LogWatcherOptions, not
restated), are one incident — 3 identities before the change, one of them ec4f187fd869d0be,
which is a live node id on the control instance; the same line cut and uncut is one incident; the
clear line of the same site, and two short lines differing only at their end, still split; a cut
inside the bound is not evidence (red with the exclusion removed); a cut through a stack trace is;
a cut through or before the first application frame is not (red with the frame check removed).
Not established: the masking rule MaskSubjects collects subjects from the whole message, so a
subject that appears only past the cut could mask an earlier bare token in a whole line and not in a
cut one. No line in the measured corpus does this; it is a known edge, not a measured one.
4a. The post-roll reading, and the one split left (MeshWeaver#5271, third half)
After memex rolled onto 3.0.0-ci.9332 (Plugins 1470fbf3, which carries sections 3 and 4), the
incident listing still read [{n}f0{n}c#{n} started {time}] and four back-pressure nodes appeared
within 25 minutes of the new pod's boot. That text is not what the portal hashes. An incident's
normalizedMessage and display name are the REPORTER's — LogLineParser.Normalize as compiled into
the separately shipped mw-log-watcher image, which predates MixedHexId and still leaves an
activation id's hex letters behind. The node id is the portal's: fed the stored LogSample.Line of
Admin/_LogIncident/bab5b3e412f85ea5, LogIncidentIdentityResolution.Resolve on 1470fbf3
returns exactly bab5b3e412f85ea5, having masked that stamp to [{hex}#{n} started {time}]. So
read the fold from occurrence counts and ids, never from the name: fbd27b4e342e917c (stream-routed
leg) held 415 crossings from many activations and 981bb5aa56136a69 (dispatch leg) 162, both still
folding after the roll; 0beb991b4f0f4645 is the NACK leg label, a different branch by the same
rule that keeps dispatch and stream-routed apart; and all four were re-addressed on boot, not
minted by fresh crossings (bab5b3e4… holds one occurrence from 2026-09-24).
The one genuine split was bab5b3e4… itself: its sample is a dispatch crossing identical in shape to
981bb5aa…'s except for Latest dispatch target rsalzmann. SpacedSubject masks a subject only
when it starts upper-case or with a digit — that guard keeps prose after the noun (target was not found) out — and a user partition is lower-case. The same sample with the target capitalised
resolves to 981bb5aa…. DispatchTargetSlot masks that field by POSITION instead: the token between
dispatch target and the line's own — separator, which prose never has. It produces
SpacedSubject's own target {id}, so every spelling of the slot is one text. Pinned in
ALowerCaseDispatchTargetIsStillASlotTest over the two nodes' stored samples verbatim: the live ids
reproduce, the lower-case one folds onto 981bb5aa… (red without the rule, answering
bab5b3e412f85ea5), and prose after the noun is untouched.
Not established: whether the watcher's stale display text will be refreshed by a watcher roll or should be replaced by the portal's own normalization on ingest; neither affects the identity.
5. A resilience pipeline is a caller (MeshWeaver#4528)
Polly names the pipeline an event belongs to inside a quoted literal —
Source: '{pipeline}/{instance}/{strategy}' — and the quoted-literal rule masked it whole, so every
Polly timeout in the fleet shared ONE identity: the plugin bundle client, the plugin catalog client
(#4222), the self-update hand-over riding the portal's shared defaults pipeline, and Orleans'
placement timeouts. The incident MeshWeaver#4528 is linked to reached 4,119 occurrences with samples
from all four across its life, and could close on none of their fixes.
LogLineParser.Normalize now keeps the pipeline, the instance and the strategy's KIND as a
canonical suffix, {pipeline:<pipeline>/<instance>/<kind>}, read from the ORIGINAL message exactly
like {types:…} (#3545). Every timeout variant is the kind Timeout, so an AttemptTimeout and the
TotalRequestTimeout it rolls up into stay one incident; any other strategy (a circuit breaker, a
retry) keeps its own name. 🚨 It reads only behind Polly's own telemetry preambles
(Resilience event occurred. EventName: '…', , Execution attempt. , Resilience pipeline executing/executed. ) — a bare Source: '…' in any other message is a path and stays masked, or
every tenant's path would mint its own incident. Core's half (Doc/Architecture/EveryHttpClientNamesItsPipeline,
MeshWeaver#5664) is what gives the
shared defaults pipeline a non-empty instance — the client's own name — so the self-update hand-over
reads -standard/self-update-handover/… rather than -standard//….
This SPLITS an existing identity, which the evidence vote in section 1 never does: the next recurrence of each pipeline lands on its own fingerprint, and the corpus migration carries the old incident's history only onto the successor its reporter fingerprint resolves to.
What this does NOT fix
A second, independent duplication mechanism is still live, and it accounts for the other 8 of
the 94: one incident node, several GitHub issues. Admin/_LogIncident/54ecfdaea23fd110 alone
carries issues #5017, #5018, #5019, #5020, #5021, #5022 and #5060 — six of them inside 32 seconds —
with identical firstSeen, lastSeen and occurrences in every body, i.e. filed from
byte-identical incident state, while the node holds only four triage threads. Two other
fingerprints did it inside ONE second: #5031/#5032 (4d74fa633047b387, both 22:23:03Z) and
#5034/#5035 (2fe3118d64124621, both 22:28:53Z).
The claim that is supposed to make filing happen at most once (ClaimRequest) decides on the
content the update lambda reads from this hub's mirror and then emits an RFC 7396 merge patch.
Two merge patches carrying the same fields do not conflict, so two workers reading the same
pre-claim mirror both grant themselves the File and both open an issue. It is filed separately
rather than fixed here because the remedy is an atomic claim — a different change, in a different
layer, with its own controls.
Where it is pinned
ResiliencePipelineIsPartOfTheIdentityTest— section 5: five Polly sources are five identities; a breaker and a timeout on one pipeline are two; one request's attempt and total timeouts are one; and a non-PollySource: '<path>'stays masked.ATruncatedSampleKeysOnWhatItKeptTest— section 4: a sample cut at the production length keys on what it kept, not on where the cut fell.PortalNormalizesWhatItHashesTest— the anchor (the live split reproduced from the stored text), a fold control (three crossings of one site by a stale reporter ⇒ one incident, where hashing the reporter's text gives three), a split control (a genuinely different fault keeps its own incident), the no-evidence fallback, the header-only exclusion, and order-independence of the vote. It runs onFixtures/memex-2026-09-20-routing-backpressure.txt, four bursts copied verbatim out of the tickets' own evidence blocks.EveryAutomatedBugCarriesASeverityTest— drives the realLogIncidentFiler.Fileagainst the fake GitHub and asserts what GitHub was asked for, on both sides: a draft that proposed nothing, and a draft that judged.