The portal masks what it hashes

MeshWeaver.Plugins#2218. Two defects in the incident-to-ticket pipeline, found from the same measurement: on 2026-09-21 Systemorph/MeshWeaver had 94 open issues opened by systemorph-com[bot], all minted between 2026-09-19T11:44Z and 2026-09-21T09:00Z, and 42 of them were one log site. None of the 94 carried a severity when it was opened.

1. A fold that could never see the fix

What was wrong

LogIncidentIdentityResolution exists because the identity function used to live in mw-log-watcher, a separately shipped image, so a corrected identity reached production only when somebody rolled it — measured five weeks and four revisions behind. #1796 moved the function into the portal, which rolls continuously.

It moved the function and not the function's input. The identity is computed over report.NormalizedMessage and report.NormalizedDetail, and both are produced by LogLineParser.Normalize inside the watcher. So every masking revision since — the episode stamp and the 8-character activation id (#2165), masking by SLOT rather than by what a token looks like (#2167, #2168), a list's length (MeshWeaver#2157), a nodeType: term (MeshWeaver#3545) — shipped in MeshWeaver.Observability.Contract, was present in every portal, and was dead code on the ingest path. The same defect #1796 closed, one layer down.

The measurement

Admin/_LogIncident/54ecfdaea23fd110 on the control instance stores

[ROUTE] Routing back-pressure [e9b2{n}fb#{n} started {time}]: {n} route dispatches in flight …

The activation id is e9b212fb. It was masked only where it happened to contain digits — e9b2{n}fb — so every grain activation normalized to a different string, hashed to a different fingerprint, opened a different incident node and produced a different GitHub issue. Three facts, each reproducible:

The portal's identity function over the watcher's stored text 54ecfdaea23fd110 — exactly the live node id
The same function over the same bursts' RAW lines, re-normalized here 96c85f04cadba7f7 for 42 of the live fingerprints
Live fingerprints whose own samples agree here / distinct identities they resolve to 70 → 26 (the other 16 still hold samples that disagree — pre-existing masking gaps, unchanged)

So the identity function was live and correct; its input was five weeks stale. The agent triaging those incidents had already written the diagnosis into its own tickets — "the fingerprint appears to split per episode stamp, so every crossing opens a new ticket" — which is what the ticket flood looks like from inside it.

The fix

The report already carries its own raw evidence: LogSample.Line is the whole verbatim burst, header and continuation lines, exactly as BurstAggregator.Aggregate joined it. The portal re-parses that with the LogLineParser it ships and hashes its own normalization.

Nothing in the watcher changes, and nothing needs to be rolled for the fix to take effect — which is the property the whole arrangement exists for.

What happens to the 42

Nothing here closes them; the corpus migration is what carries a re-addressed incident's ticket, counts and history onto its successor, keyed on the reporter's own fingerprint. The next crossing of that site lands on the folded identity and brings one of them with it; the rest are a backlog a human closes as duplicates.

2. No automated bug was ever inside the release gate

What was wrong

The issue taxonomy puts a severity on every bug and makes two of the four a release gate: sev:B and sev:H must both be zero to cut a release, and a bug carrying neither "has not been triaged, and working it ahead of a labelled sev:H is choosing by accident."

The pipeline had never assigned one. All 94 open bot-filed issues were opened with bug and nothing else; the 23 that carry a severity were labelled hours or days later by a person — issue #4824 got bug from the App at 2026-09-19T11:51:05Z and sev:M from a human at 2026-09-20T08:18:03Z. The LogTriage agent's instructions spelled the label set out literally as ["bug"], so there was nothing for the model to get wrong and nothing to fail. 71 tickets sat in a queue no gate reads.

The fix, in two halves

The default is deliberately sev:M and deliberately not a blocking class. A machine that did not judge must not be able to hold the release line — one noisy log site would stop every cut — and it must equally not be sev:L, which parks a real defect in a backlog nobody reads. A draft that proposes several severities keeps all of them: choosing between two labels a human can see is a judgement this code has no basis for, and silently dropping one would hide that triage was unsure.

One legacy incident, fifty-two successors

The corpus migration records the same finding from the other side, and it is the cleanest statement of it. When the portal's identity re-addresses an incident, LogIncidentCorpusMigration carries the legacy node forward and every successor's ticket says which legacy id it inherited. 52 of the 94 open bot-filed issues name the same one — e4a97855ab595beb.

That id is the REPORTER's fingerprint for the whole RoutingGrain back-pressure site: the watcher, with all its masking flaws, still called it ONE fault. The portal then split it 52 ways. So the split is provably not in the reporter's grouping and not in the identity function — both agree the site is one thing — it is in the TEXT the portal hashed, which is the reporter's and stale.

It also explains a sentence those tickets carry that reads like a contradiction: "superseded and will not fold, file or comment again". That is true, and it is about the LEGACY node, which is correctly retired. Fold's alreadyMoved guard means only the FIRST successor inherits the ticket and the counts; every later one gets provenance only — and then files an issue of its own. 52 tickets carrying that sentence is the defect's signature, not a second defect.

The same id, masked eleven different ways

One listing of Admin/_LogIncident is the whole finding, because the node NAME carries the normalized text. These are all the same 8-hex activation id at the same log site, as the stale watcher masked it:

[{n}f5{n}ea#{n}   [b5{n}d3{n}e#{n}   [df5d5f6{n}#{n}   [{n}a3d7#{n}     [{n}b2c7c6c#{n}
[aae3{n}b#{n}     [{n}c7f0ee#{n}    [{n}b0{n}bd1#{n}  [{n}c2{n}#{n}    [{n}d6fde0#{n}
[{n}b6{n}d4{n}f#{n}  [{n}d8{n}e3#{n}   [{n}d#{n}      [df8ca0{n}b#{n}  [e9b2{n}fb#{n}

The bare-number rule ate whichever runs of digits the id happened to contain and left its hex letters standing as literals, so the masked text is a function of the id — which is exactly what masking exists to prevent. EpisodeStamp and MixedHexId (#2165, #2167) mask the whole token; they have been in this assembly since 2026-09-19 and had never once reached the hash.

The flood stopped for the wrong reason

Incident DETECTION on the control instance stopped around 2026-09-20T22:30Z and had not resumed 15 hours later: content.lastSeen:2026-09-21* returns 0, not truncated, over coverage.partitions: ["admin"], while content.lastSeen:2026-09-20T2* is still truncated at 50. POST /api/log-incidents is the ONLY way a burst enters the portal, so with the watcher silent no report arrives, no incident is minted and no File is requested — which is why the last tickets (#5058–#5063) were filed from incidents detected BEFORE the stop and the flood appeared to have ended. Nothing was fixed by that outage. This fix is at ingest, so it applies to the first report after the watcher is alive again — alive, not rolled, which is the entire point of it.

And it came back, still splitting. The same query at 2026-09-21T15:41Z returned ≥10 (truncated), with incidents whose lastSeen is 15:40:49Z — so detection resumed some time between 13:31Z and 15:41Z. The first RoutingGrain nodes minted after it resumed are 0fc7665eaf6aae43 ([cf0fc5ed#{n} …), 18104cb0dbf5d912 ([d4cae9e3#{n} …), 0fda5918d57c6050 ([{n}f7{n}a#{n} …), 3708a213a82ec7b0 ([{n}c9{n}d#{n} …) and 4db1545b366d9574 ([a6{n}b4{n}#{n} …): five activation ids, five incidents, one log site, masked five different ways. The defect is not historical and the flood is resuming as this is written.

🚨 And the watcher's liveness is unobservable from the portal by construction: it has no Hosting/Deployment record (namespace:Deployments scope:descendants returns only the portal instances' records), so nothing reports its image, samples it or can roll it — and every signal that it has stopped is emitted BY it, including log-pipeline-behind-* and the silent-window finding whose job is to announce exactly this. A missing self-finding is evidence FOR death, never for health. The filer does NOT share that shape — it is a BackgroundService inside the portal, so it rolls with the portal and is covered by the portal's own record and /health.

3. The back-pressure line's oldest-leg label: two more instance slots (MeshWeaver#5271)

Core's back-pressure line gained a new discriminator — "oldest leg in flight N ms — dispatch → <address> (delivery <id>)" — and both values in it are instances. Neither was masked:

DeliveryId masks the id inside the line's parenthesised (delivery …) field only (so prose and a message type after a bare "delivery failed" survive); LegDestination masks whatever follows dispatch → / stream-routed →, and keeps the leg KIND, because the two are different routing branches.

Measured over the 146 back-pressure sample lines carried by the open core issues filed between 2026-09-19 and 2026-09-22 (43 distinct filed fingerprints): the identity function over the new normalization yields 3 — one per leg kind on the current wording, and one for the pre-#5151 wording (deepest per-destination queue), which is a different log template and rightly its own. Pinned in LogLineParserTest (Fingerprint_FoldsBackPressureCrossingsThatDifferOnlyInTheOldestLegsDeliveryId, Normalize_FoldsTheOldestLegsDestination, and the two must-not-fold controls Normalize_DoesNotMaskDeliveryOutsideItsSlot, Normalize_KeepsMissingHandlerFaultsForDifferentTypesApart and Normalize_KeepsTheLegKindApart).

This changes the fingerprint of every current back-pressure incident, so each re-addresses once through LogIncidentCorpusMigration onto its successor; it does not close any existing issue.

4. A cut sample keyed on where the cut fell (MeshWeaver#5271, second half)

Section 3's masking was correct, reached the control instance's image (b922f070 is an ancestor of the plugins sha the running 3.0.0-ci.9218 image was built from), and the flood went on: on 2026-09-23 the control instance minted more than forty distinct Routing back-pressure incidents in one day. The splitter was not in the masking at all.

The portal re-derives the identity from LogSample.Line, and the watcher cuts every sample at MaxSampleLength (2000) characters, appending …[truncated]. The back-pressure line is ~3,400 characters, so no sample of it is ever whole. The cut lands at a fixed RAW offset, but the fields in front of it — the oldest leg's destination, the latest dispatch target — vary in length per crossing, so it falls at a different word of the constant tail each time (AT THE CROSSIN…, AT THE CR…), and the identity hashed a different string for every length combination. Admin/_LogIncident/78316ad8e1d82122 shows both halves at once: nine crossings from five activations folded — the ones whose cut happened to land on the same character — while its siblings were minted beside it.

Both earlier measurements missed it for the same reason: they removed the variable. The #2286 measurement cut each line "at the constant tail" by hand before normalizing, and PortalNormalizesWhatItHashesTest aggregates with maxSampleLength: 4000.

The fix is a statement about evidence, not a bound to widen. The identity reads the first StructuralLogIncidentIdentity.DetailBound (1024) characters of the masked detail — a prefix every sample of a long line keeps — and LogIncidentIdentityResolution.ReparsedBursts counts a cut sample only when the cut (the marker is left in the line when it is parsed, so it lands in whichever field the cut ran through) falls at least 128 characters past that bound. A cut through a stack trace leaves the fault's text whole and still counts; a cut inside the bound cannot vouch for the prefix and is not counted. The top frame is checked the same way: it is taken again from the sample's COMPLETE lines (the line the cut ran through is partial) and must agree, and a cut sample carrying an exception but no complete application frame is not counted, because the first one may lie past the cut and the identity would switch from the frame branch to the site branch. A report left with nothing countable resolves on its reported fields, as a report with no evidence always has. A message shorter than the bound hashes exactly as before; one longer re-addresses once through LogIncidentCorpusMigration, like section 3.

Pinned in ATruncatedSampleKeysOnWhatItKeptTest: four dispatch crossings with the values the live samples carry, cut at the watcher's own default MaxSampleLength (read from LogWatcherOptions, not restated), are one incident — 3 identities before the change, one of them ec4f187fd869d0be, which is a live node id on the control instance; the same line cut and uncut is one incident; the clear line of the same site, and two short lines differing only at their end, still split; a cut inside the bound is not evidence (red with the exclusion removed); a cut through a stack trace is; a cut through or before the first application frame is not (red with the frame check removed).

Not established: the masking rule MaskSubjects collects subjects from the whole message, so a subject that appears only past the cut could mask an earlier bare token in a whole line and not in a cut one. No line in the measured corpus does this; it is a known edge, not a measured one.

4a. The post-roll reading, and the one split left (MeshWeaver#5271, third half)

After memex rolled onto 3.0.0-ci.9332 (Plugins 1470fbf3, which carries sections 3 and 4), the incident listing still read [{n}f0{n}c#{n} started {time}] and four back-pressure nodes appeared within 25 minutes of the new pod's boot. That text is not what the portal hashes. An incident's normalizedMessage and display name are the REPORTER's — LogLineParser.Normalize as compiled into the separately shipped mw-log-watcher image, which predates MixedHexId and still leaves an activation id's hex letters behind. The node id is the portal's: fed the stored LogSample.Line of Admin/_LogIncident/bab5b3e412f85ea5, LogIncidentIdentityResolution.Resolve on 1470fbf3 returns exactly bab5b3e412f85ea5, having masked that stamp to [{hex}#{n} started {time}]. So read the fold from occurrence counts and ids, never from the name: fbd27b4e342e917c (stream-routed leg) held 415 crossings from many activations and 981bb5aa56136a69 (dispatch leg) 162, both still folding after the roll; 0beb991b4f0f4645 is the NACK leg label, a different branch by the same rule that keeps dispatch and stream-routed apart; and all four were re-addressed on boot, not minted by fresh crossings (bab5b3e4… holds one occurrence from 2026-09-24).

The one genuine split was bab5b3e4… itself: its sample is a dispatch crossing identical in shape to 981bb5aa…'s except for Latest dispatch target rsalzmann. SpacedSubject masks a subject only when it starts upper-case or with a digit — that guard keeps prose after the noun (target was not found) out — and a user partition is lower-case. The same sample with the target capitalised resolves to 981bb5aa…. DispatchTargetSlot masks that field by POSITION instead: the token between dispatch target and the line's own — separator, which prose never has. It produces SpacedSubject's own target {id}, so every spelling of the slot is one text. Pinned in ALowerCaseDispatchTargetIsStillASlotTest over the two nodes' stored samples verbatim: the live ids reproduce, the lower-case one folds onto 981bb5aa… (red without the rule, answering bab5b3e412f85ea5), and prose after the noun is untouched.

Not established: whether the watcher's stale display text will be refreshed by a watcher roll or should be replaced by the portal's own normalization on ingest; neither affects the identity.

5. A resilience pipeline is a caller (MeshWeaver#4528)

Polly names the pipeline an event belongs to inside a quoted literal — Source: '{pipeline}/{instance}/{strategy}' — and the quoted-literal rule masked it whole, so every Polly timeout in the fleet shared ONE identity: the plugin bundle client, the plugin catalog client (#4222), the self-update hand-over riding the portal's shared defaults pipeline, and Orleans' placement timeouts. The incident MeshWeaver#4528 is linked to reached 4,119 occurrences with samples from all four across its life, and could close on none of their fixes.

LogLineParser.Normalize now keeps the pipeline, the instance and the strategy's KIND as a canonical suffix, {pipeline:<pipeline>/<instance>/<kind>}, read from the ORIGINAL message exactly like {types:…} (#3545). Every timeout variant is the kind Timeout, so an AttemptTimeout and the TotalRequestTimeout it rolls up into stay one incident; any other strategy (a circuit breaker, a retry) keeps its own name. 🚨 It reads only behind Polly's own telemetry preambles (Resilience event occurred. EventName: '…', , Execution attempt. , Resilience pipeline executing/executed. ) — a bare Source: '…' in any other message is a path and stays masked, or every tenant's path would mint its own incident. Core's half (Doc/Architecture/EveryHttpClientNamesItsPipeline, MeshWeaver#5664) is what gives the shared defaults pipeline a non-empty instance — the client's own name — so the self-update hand-over reads -standard/self-update-handover/… rather than -standard//….

This SPLITS an existing identity, which the evidence vote in section 1 never does: the next recurrence of each pipeline lands on its own fingerprint, and the corpus migration carries the old incident's history only onto the successor its reporter fingerprint resolves to.

What this does NOT fix

A second, independent duplication mechanism is still live, and it accounts for the other 8 of the 94: one incident node, several GitHub issues. Admin/_LogIncident/54ecfdaea23fd110 alone carries issues #5017, #5018, #5019, #5020, #5021, #5022 and #5060 — six of them inside 32 seconds — with identical firstSeen, lastSeen and occurrences in every body, i.e. filed from byte-identical incident state, while the node holds only four triage threads. Two other fingerprints did it inside ONE second: #5031/#5032 (4d74fa633047b387, both 22:23:03Z) and #5034/#5035 (2fe3118d64124621, both 22:28:53Z).

The claim that is supposed to make filing happen at most once (ClaimRequest) decides on the content the update lambda reads from this hub's mirror and then emits an RFC 7396 merge patch. Two merge patches carrying the same fields do not conflict, so two workers reading the same pre-claim mirror both grant themselves the File and both open an issue. It is filed separately rather than fixed here because the remedy is an atomic claim — a different change, in a different layer, with its own controls.

Where it is pinned