A fold-in target that no longer resolves
Systemorph/MeshWeaver.Plugins#2028. Every occurrence of one fault is folded into ONE GitHub issue — that is the promise the filer prints on every ticket it opens: "Recurrences are folded into this issue rather than opening new ones." The promise has a precondition nobody had stated: the ticket it folds into has to still be there.
What was wrong
LogIncidentFiler.Comment requires IssueNumber, and Reopen opens with one cheap
GetIssue — the read that keeps the system from "reopening" an issue that was never closed. When
that number no longer names a reachable issue, the read is a 404. The error propagated, and the control plane did what it does
with any failed transition: parked the incident at Failed with the reason.
There was no way back out, and the reason is a rule that is right everywhere else:
{ IssueNumber: not null } => CommentDue(incident, options, now)
? LogIncidentRequest.Comment
: LogIncidentRequest.None,
The issue link outranks the status (LogIncidentIngestService.NextRequest) — deliberately, because
a ticketed incident that goes Failed used to be re-triaged and re-filed, and that chain opened
eight issues for ROUTER_TRAFFIC in seven minutes. So an incident that HAS a link is offered a
comment and never a file. With a link that resolves to nothing, every later recurrence re-entered
Comment, re-issued the same 404, and re-parked the incident. Comment clears Failed only when a
comment lands; File only runs for an incident with no issue yet. Nothing clears a dead
IssueNumber.
Measured
Admin/_LogIncident/9b70b639c4e77af3 on the control instance, read 2026-09-17:
"repository": "Systemorph/MeshWeaver", "issueNumber": 2950,
"status": "Failed", "error": "Not Found",
"occurrences": 2, "occurrencesAtLastComment": 1,
"firstSeen": "2026-09-01T09:04:13Z", "lastSeen": "2026-09-17T04:44:44Z"
The fault (an agent turn dying when an OpenAI-compatible gateway ends the stream with a non-standard
finish_reason: "error") had fired twice, sixteen days apart, and the second occurrence was recorded
nowhere a person looks. A complete ticket draft sat on the incident, never posted.
🚨 What actually happened to #2950, and where the 404 comes from
Systemorph/MeshWeaver#2950 was transferred, not deleted. That distinction matters, because it
is where the 404 is born — and it is not where anyone would look for it. Measured 2026-09-17:
| request | answer |
|---|---|
GET /repos/Systemorph/MeshWeaver/issues/2950 |
200 — GitHub follows the transfer redirect and returns number: 1139, repository_url: …/MeshWeaver.Plugins, closed 2026-09-01T23:38:30Z |
GET /repos/Systemorph/MeshWeaver/issues/2950/comments |
404 |
GET /repos/Systemorph/MeshWeaver/issues/2950/events |
404 |
The redirect covers the issue resource and not its sub-resources. OctokitGitHubRepoClient.GetIssue
reads the issue and then all of its comments:
Http.InvokeObservable(ct => client.Issue.Get(owner, repo, number)) // 200, redirected
.SelectMany(issue => Http.InvokeObservable(ct =>
client.Issue.Comment.GetAllForIssue(owner, repo, number)…)) // 404
so the COMMENTS leg is what throws, and the new home that the first leg resolved is discarded before
any caller sees it. gh issue view 2950 --repo Systemorph/MeshWeaver reports "Could not resolve to
an issue" for the same shape of reason — it asks over GraphQL, which does not follow the transfer —
which is why the issue that reported this concluded the ticket was gone.
What this cost while the read was not available (the state this file described until
Systemorph/MeshWeaver.Plugins#2171): a transferred ticket was treated like a deleted one, so the
fault was re-filed and the replacement named the old reference, whose link still redirects a human to
wherever the ticket went. Strictly better than losing the recurrence entirely — and for this
incident also the right outcome, since MeshWeaver.Plugins#1139 is CLOSED and the fault recurred on
2026-09-17, a new ticket by any rule.
But as a general rule it was a treadmill, and that is measured. The router sends every
MeshWeaver.* category to core (Systemorph/MeshWeaver.Plugins#2171), so a HAND-TRANSFER is the only
remedy available for a ticket filed in the repository that does not contain the code. It did not stay
done: Plugins#2132 was transferred out of core at ~05:17Z on 2026-09-19 and systemorph-com[bot]
opened a byte-identical MeshWeaver#4767 in core at 05:31:44Z — fourteen minutes. Every
hand-transfer was undone on its fault's next occurrence, which is why "move the misrouted tickets"
was never a fix.
🚨 The transfer is now FOLLOWED — the read exists, and the recurrence goes where the ticket went
Two changes, one in each repository, and the order was forced (platform first):
| half | where | what |
|---|---|---|
| the read | core MeshWeaver.GitSync, IGitHubRepoClient.FindIssueState (MeshWeaver#4629) |
one request, no comments, so it follows the redirect and emits the issue at its CURRENT home — Number and Url name the new repository. null for a number that names no issue at all. A default interface member delegating to GetIssue, so it obliged no implementer |
| the follow | this repo, LogIncidentFiler.Comment / Follow (Plugins#2171) |
the recurrence path reads through FindIssueState once, hoisted above the ReopenOnRecurrence gate, and re-points the incident's Repository / IssueNumber / IssueUrl when GitHub's answer names a different home |
Three things about the shape are load-bearing:
- The read is hoisted above the reopen policy. Following a ticket that moved is not a reopen
decision, so it must happen whether or not
ReopenOnRecurrenceis on.Reopenis handed the already-read issue instead of fetching its own, which also takes the recurrence path from two requests to one. - Nothing is superseded. A transfer is ONE ticket at a new address, not a replacement, so
SupersededIssueUrlis left alone — stamping it would put a "no longer resolves" note on a live ticket. A DELETION still sets it, still clears the link and still re-files. nullis the only answer that may clear a link, and even then nothing is concluded from it here: the ISSUES-listing probe below still has to separate "deleted" from "I cannot read this repository's issues" first. That is why the follow cannot be written as "re-point on any 404" — the two arrive as the same 404 fromGetIssue, and only the comments-free read tells them apart.
The control on each side is in LogIncidentFilingIdempotencyTest, and the negative half was run:
with the production change reverted, ATransferredIssue_IsFollowedToItsNewHome_AndNothingIsRefiled
fails (it times out waiting for the incident to reach the new number, because the old code cleared
the link and re-filed) while ADeletedIssue_IsStillRefiled_SoTheFollowDoesNotSwallowARealAbsence
passes — so the new case discriminates exactly the behaviour the change adds, and the deletion path
is untouched. The fake models the production asymmetry rather than a flag: TransferTo makes
GetIssue answer 404 for the old number (the comments sub-resource) while FindIssueState answers
the issue at its new home, which is what GitHub does.
🚨 What this does NOT fix. The ticket still arrives in the wrong repository in the first place —
routing by logger Category is Plugins#2171's remaining half, and it is blocked on something a code
change cannot establish: systemorph-com[bot] has never opened an issue in
Systemorph/MeshWeaver.Plugins, and a route into a repository the filer cannot write to is strictly
worse than a misroute (the ticket would be in no backlog and the failure a line in the watcher's
log). What the follow buys is that the manual remedy now STICKS — one transfer, by hand, and the
fingerprint learns its new home on the next occurrence.
404 does not mean "deleted"
The obvious fix — 404 ⇒ clear the link and re-file — is wrong, and wrong in the expensive direction. GitHub answers 404, never 403, for a private resource the caller may not read. An App installation that has lost access to a repository gets exactly the same answer at the issue endpoint as a deleted ticket does. Clearing the link on that reading opens a DUPLICATE of an issue that is alive and well, and duplicates are the defect this whole file exists to prevent.
That is not hypothetical, and it is not only a machine's mistake: Systemorph/MeshWeaver.Plugins#2029
was filed the same day on exactly this reading. gh repo view Systemorph/MeshWeaver.Feedback answered
"Could not resolve to a Repository" under a credential without org access — byte-for-byte what a
deleted repository returns — and gh repo list Systemorph omitted it for the same reason, so two
instruments agreed from one blind spot. The repository is private, exists, and is covered by the
systemorph-com App.
So the 404 is not the verdict, it is the question. On that rare path the filer LISTS the repository's issues once, and reads the answer off the listing rather than inferring it:
| the point read | the listing | verdict |
|---|---|---|
| 404 | comes back, and the number is NOT in it | the ISSUE is gone — clear the link and re-file |
| 404 | comes back, and the number IS in it | the ticket is there and something else refused — keep the link |
| 404 | refused, or cannot be established | nothing is established — keep the link, fail loudly |
🚨 The probe is ISSUES-scoped on purpose. The first cut of this asked
GetCanonicalRepository — GET /repos/{owner}/{repo} — and that is not sound: a GitHub App
installation token always carries repository metadata access, while issues is a separate
permission. A metadata read therefore succeeds for a token that may not read issues at all, and the
branch would clear a LIVE link and open a duplicate the moment the permission came back. Listing the
issues is refused by exactly the permission whose absence is the confounder, so a listing that comes
back has already established the access the verdict depends on — and it answers the question
directly, since a number still in it exists whatever the point read said. (Review on #2040.)
The asymmetry is deliberate and is the same one the reopen rule uses: a kept link is recoverable,
a duplicate ticket is not. "Cannot establish" therefore fails closed, the original 404 propagates,
the incident parks at Failed with a message naming which of the three rows this was, and the
platform bell rings once for the refusal episode (#4022).
The recovery is visible, not silent
Three things happen so that nobody has to find this by reading a node:
- The incident says why it lost its ticket.
Status: Failed,Errornaming the issue that no longer resolves — a row a person or a dashboard can read, and unlike before it is a state the incident can LEAVE. - The re-file is asked for on the same write AND performed in the same turn. 🚨 Asking is not
enough, and the first cut of this fix got that wrong. The incident's fingerprint is still in the
control plane's
inFlightguard while the write carryingRequestedStatus: Fileis emitted, so the live watch reachesEnqueueand skips it — and clearing the issue link also takes the incident out ofNeedsCommentReconcile, which requires one, so the comment-debt clock no longer covers it either. Nothing guarantees a later emission; the request would sit on the node, readable and unperformed.LogIncidentControlPlane.Followtherefore performs it in the current turn, through the SAME claim every other transition goes through, withmayFollow: falseinside bounding the chain to one hop by construction rather than by a counter. (Review on #2040 — whose test hid the hole by driving the second transition by hand; it now drives one entry point.) - The replacement ticket names the one it replaces.
LogIncident.SupersededIssueUrlis a field of its own — deliberately NOTError, which the claim clears before every attempt — and the new issue's body carries "This issue replaces …, which no longer exists." Without it the replacement reads as a first sighting of a fault that has been ticketed for weeks, and the only trace of the ticket that disappeared is a log line nobody is looking at.
Recovery of the incidents already parked
None is needed, and no operational action is required. An incident stuck this way still satisfies
NeedsCommentReconcile — it has a link, Occurrences > OccurrencesAtLastComment, and a comment is
due — so the comment-debt clock re-examines it every reconcile period, meets the 404, and takes the
new path. 9b70b639c4e77af3 recovers on the first tick after the portal carries this change.
Pinned by
LogIncidentFilingIdempotencyTest:
AnIssueThatNoLongerExists_ClearsTheDeadLinkAndRefilesTheFault— the issue is deleted from the fake while its issues still list. It drives one entry point (theCommentrequest) and asserts the replacement ticket, so it covers the follow-on transition too. Mutation control: with the recovery arm disabled (IsNotFoundforced false) it fails on a 60-second wait for aFilethat never comes, which is the production symptom exactly.AnIssueThatCannotBeReadAtAll_KeepsItsLinkAndFailsLoudly— the same 404 with the issue listing refused too. A guard against OVER-reach rather than a control for the fix; its own control is the opposite mutation — with the probe replaced by a constant "readable" it fails, because the incident re-files and opens the duplicate.TheIssueIsStillListedInItsRepository_KeepsTheLinkEvenThoughThePointReadSaid404— the point read refuses while the ticket is demonstrably there. This is the row that makes the listing an ANSWER rather than an inference, and it is what a metadata probe could never have decided.
The fake issue API answers Octokit.NotFoundException with HttpStatusCode.NotFound for a missing
issue, which is what the production client surfaces. The predicate that reads it
(LogIncidentFiler.IsNotFound) is structural — the HTTP status the client carries, unwrapped
through AggregateException/InnerException, never the message text, because the wording belongs
to someone else's library.