When the model's stream ends badly

An OpenAI-compatible gateway can end a streaming response in ways the wire protocol does not describe. When it does, the round ends — but it ends naming the condition, in a sentence the person can act on, with everything that was streamed before the fault kept. This page records how, and the one thing the design deliberately does not do.

The three endings, and the guard that recognises each

ProviderStreamFault (src/MeshWeaver.AI/ProviderStreamException.cs) names them:

Fault What the transport did Found by
Stalled The connection stayed OPEN and the bytes stopped StreamStallGuardChatClient — an idle bound (AiStreamingLimits)
Faulted A payload arrived that the protocol cannot represent, or an error object instead of an answer OpenAIWireStreamGuard — the SDK's own deserializer threw, or the empty update carried error
Truncated The body ended before the response did StreamStallGuardChatClient — end-of-stream where a chunk was owed

A stall is found by a bound and a truncation by an exception precisely because they are opposite transport states; conflating them would make one of the two undiagnosable.

The canonical Faulted case is OpenRouter's finish_reason: "error", which the family uses to signal a fault at the upstream model provider. That value is not in the OpenAI specification, so OpenAI.Chat.ChatFinishReasonExtensions.ToChatFinishReason throws ArgumentOutOfRangeException: Unknown ChatFinishReason value. … Actual value was error. while deserializing the SSE chunk.

The second Faulted shape throws NOTHING, and it is the dangerous one (MeshWeaver#5935). When OpenRouter has ACCEPTED a request (HTTP 200, the SSE stream open) but cannot serve it, it streams a lone {"error":{"code":…,"message":"…"}} chunk with no choices, no model and no finish_reason. The SDK has no member for error, so it deserializes the chunk to an empty update and the stream completes normally. Without the guard, the round reads "the stream completed with zero tokens" from model unknown, and the provider's reason is gone. On the control instance, 2026-09-30 22:21Z onwards, every agent ended that way: review, triage and bug-triage, on GLM and on Claude alike. Nothing anywhere said why. The SDK keeps the unmodelled error property in the update's additional raw data, so the guard serializes only an update that carries no content, no tool-call delta, no finish reason and no usage. If that update holds an error object, the guard throws Faulted with the provider's code and message, and logs both at Warning ([OpenAIWire] The provider streamed an error object …). A healthy content chunk never pays for the check. RawSdkClient_OnErrorObject_CompletesEmptyAndLosesTheReason pins the SDK behaviour. If a later SDK surfaces the error itself, that test goes red, and that is the signal to retire the check.

The sentence after the provider's words says WHOSE condition it is, and only what the code and the wording establish (MeshWeaver#5960). One canned sentence used to follow every streamed error: "a fault at the provider or the provider account …, not a problem with the request". That was false for the commonest 403 of 2026-10: the gateway's PII redaction filter answering [403] Request blocked: PII detected (redaction_context_lost), which is a verdict ON the request. The code alone cannot tell the two apart — a key limit is a 403 too — so OpenAIWireStreamGuard.Describe reads the wording:

streamed error sentence
4xx other than 408/429 (or no code) whose message says Request blocked, PII detected, flagged or moderation the gateway refused this request under a content policy; the key keeps serving other requests. The guard sets ProviderStreamException.RequestBlockedDataKey in Data (a const key, so the provider module needs no newer engine), and LocalizationKey is then chat.modelRequestBlocked
401, 402, any other 403 a condition of the provider key or account
408, 429, 5xx a temporary condition at the provider (and marked transient, as before)
anything else (404, a string code) the provider's reason is quoted, and nobody is blamed

408 and 429 are retried, so they are never described as a block whatever their text says: the sentence and the transient mark cannot disagree. The guard's own sentences carry no key-health marker ("rate limit", "key limit", "credits"), because ProviderHealthRule reads the whole exception chain for those words; only the provider's own text may say something about the key.

The user-facing cell renders LocalizationKey when the loaded platform catalog defines it, else Message. chat.modelStreamFaulted says "not a problem with your request — submit again", which is false for a block, so a block names chat.modelRequestBlocked instead. Until the platform catalog carries that key, the guard's own English stands.

ProviderHealthRule.Classify already kept the redaction block from marking the key unhealthy; this keeps the round's own message from saying the opposite. A block is not marked transient: the same content is refused again. GuardedClient_OnARedactionBlock_NamesARequestBlock_NotAnAccountFault and EachStreamedErrorClass_GetsItsOwnSentence pin it, with a 403 key limit and a 5xx that says "blocked" as the controls.

Why the payload guard lives in MeshWeaver.AI.OpenAI and not in the engine. The trigger is one SDK's specific throw, and that module owns that SDK. A general "an unexpected exception during deserialization means the provider is at fault" rule in MeshWeaver.AI would relabel our own deserialization bugs as provider faults. The match is narrow for the same reason — the "Unknown <Enum> value." message family only — so an ordinary range violation keeps reporting itself.

Why every guard sits BELOW FunctionInvokingChatClient. The scope has to be the model call and never a tool invocation, or a tool's own ArgumentOutOfRangeException would be reported as a gateway fault. The two guards compose: the engine-wide stall guard wraps the OpenAI-wire one, and each catches only what it can actually recognise.

One detail in OpenAIWireStreamGuard looks like defensiveness and is not: disposing a provider stream that faulted mid-parse makes the SDK's transport try to re-buffer content it has already consumed, and the exception coming out of that would replace the ProviderStreamException on its way up — the round would report a cleanup artefact instead of what killed it. The dispose failure is therefore logged and not allowed to become a second way for the round to fail.

What the round does with one

ThreadExecution's terminal-error path (FindStreamFault) names the condition ahead of the HTTP-status switch, because a stream fault carries no status and would otherwise fall through to ex.Message — SDK internals standing in for the thread's account of why it failed. It then:

The localized string is taken only when the loaded platform's catalog defines the key; otherwise the exception's own purpose-written English stands. That check is what lets this module ship a condition before the platform image carrying the string does, instead of rendering a raw chat.modelStream… token at a user.

The live reading (issue #2131)

#2131 was filed automatically from incident Admin/_LogIncident/9b70b639c4e77af3 and reported that the round "dies … as an unhandled ProviderStreamException" with "no response and no user-readable explanation, only a server-side error". The incident's own second occurrence falsifies that. Read on the public instance at rbuergi/_Thread/read-only-test-do-not-draft-send-reply-m-d748/15111171, the cell the 2026-09-17 04:44:44Z occurrence produced:

status      : Error
modelName   : qwen/qwen3.8-27b
text        : *Error: The language model 'qwen/qwen3.8-27b' ended its response with a provider
               error. This is a fault at the upstream model provider, not a problem with your
               request — submit again to retry, or pick a different model.*
toolCalls   : 2 × SearchMail, both isSuccess:true, results intact
timing      : modelCalls[1].outcome = "Faulted"
tokens      : 14761 in / 142 out / 14903 total

The thread root carries no status, i.e. Idle (Thread.Status's default), and its summary is that same sentence. So the turn ended as a named, localized, reportable outcome with the partial work preserved — not as a lost round. The [ThreadExec] ERROR log line the bot filed on IS the handled path; the guard and the round's branch both landed in the #2632 work on 2026-08-29, before either occurrence, and core's two catalog keys landed 2026-09-07.

The decision: the round does NOT retry automatically

#2131 asked for one — "retry the round once (the gateway already told us it is transient)". That is deliberately not done, and this is the record of why:

  1. A retry is a bound spent, not a defect fixed. The house rule is root cause only; an upstream gateway's health is not something this process can change, and a silent retry converts a visible upstream condition into an invisible one. The number of retries then becomes the knob, which is exactly the shape the no-band-aids rule refuses.
  2. The round is not idempotent. The 2026-09-17 reading above shows two SearchMail calls that had already succeeded when the stream faulted. A retried round re-runs the tool calls. For a read that is waste; for a write — a draft, a mail, a mesh node — it is a second side effect with nothing to deduplicate it, and the Executive Assistant's tools are precisely that kind.
  3. The tokens are already spent and already charged. A retry doubles the cost of a round the user may not want repeated on a model that is currently failing; "pick a different model" is often the better remedy and only a person can choose it.
  4. The remedy the message names is one click. hub.ResubmitMessage is wired to the errored cell (ThreadMessageLayoutAreas), so "submit again to retry" is an affordance and not advice.

A retry that is genuinely wanted is therefore an explicit, per-agent policy decision with a bound and a ledger — not a catch in the execution loop. Nothing in this design prevents that; it just is not what "handle the fault" means.

The one exception: a CALL that produced nothing is re-issued

The round is still never retried. What is re-issued, since 2026-10-02, is a single provider call that failed on a transient provider condition before the model produced any output (ProviderCallRetryChatClient, wrapped around the guarded provider client by ChatClientAgentFactory.RetryTransientProviderFaults).

The measurement. Control instance, 2026-10-02 09:27–09:42Z, while a burst of review rounds held the pinned upstream providers busy: thirteen rounds ended within seconds of starting on [OpenAIWire] The provider streamed an error object instead of a response (code 429): Provider returned error, and one on [ProviderStream] TRUNCATED (ResponseEnded, streamed=False). Nobody watches those rounds, so no one resubmitted; each counted against the head's three review rounds. The day before, a review round had drafted its review, and the very next call died on a 20-second provider fault — the draft was never confirmed.

Why the four reasons above do not apply to it.

  1. A bound spent, not a defect fixed. It is bounded and visible: AiStreamingLimits.ProviderCallRetries (default 2 — waits of 15 s and 45 s), none STARTING past ProviderRetryWindow (8 min from the first attempt, the wait included), each retry a Warning line [ProviderRetry] RETRY n/N …, and each attempt its own ModelCallTiming on the round's ledger with the outcome it really had (Faulted, then Completed). A condition that outlasts a minute still ends the round, naming itself.
  2. The round is not idempotent. A call that produced no output has executed no tool and written no text: the retry sits BELOW the function invoker and applies only while no update carrying output has been forwarded. After the first such update the call is never re-issued.
  3. The tokens are already spent. A request refused with a rate limit generated nothing.
  4. One click. True for a person; an unattended round has nobody to click.

What counts as transient. Only what arrives INSIDE a response the endpoint had accepted, which no SDK transport retries: a stall or a truncation before the first output, and a provider stream fault its own guard marked (ProviderStreamException.TransientDataKey in Exception.Data) — a streamed error object with code 408, 429 or 5xx, or the upstream finish_reason: "error" — that literal value only; any other finish reason the wire cannot represent is unknown, not weather, and ends the call at once. "Before the first output" is the stall guard's own test: Stalled and Truncated are raised by StreamStallGuardChatClient alone, whose HadStreamedUpdate is set from CarriesOutput, not per chunk, so a reasoning stream cut mid-thinking — content-less chunks only — is re-issued (ATruncationMidThinking_ThroughTheStallGuard_IsReissued composes the real guard to pin it). The wire guard's flag IS per chunk, but it raises only Faulted, whose arm reads the transient mark, never the flag; and the retry client's own "no output forwarded" bound applies to every arm. A streamed 402, 403 or 404 is a verdict and ends the call at once. An HTTP-status refusal of the request itself (a 429 or 5xx answered to the POST) is not retried here: the provider SDKs' transports already retry those with the provider's own Retry-After, and a second ladder would only multiply a bound already spent. The mark is a dictionary entry rather than a new member on purpose: the provider modules are separate packages from the engine, and an entry under a constant binds neither to the other's version.

The content-less updates. A failed attempt's envelope chunk must not reach the function-invoking loop — two message ids there fold into two assistant messages, one of them empty. So updates that carry no output are withheld until output arrives; the first is kept and released with it. A stream that ends with no output at all delivers its first and its last update, so a finish reason on a content-less terminal chunk is kept. Usage, text and tool calls all count as output and are never withheld.

Switch it off with AiStreamingLimits.ProviderCallRetries = 0.

A round cut while the model was still reasoning

A reasoning model streams its thinking as chunks the OpenAI adapter surfaces with no content. The stall guard correctly does not call that silence — updates keep arriving — so a long reasoning phase is ended by nothing but the round cap. On 2026-10-02 twenty review rounds were cut at 30:00 that way, and each cell asserted "The model endpoint stopped responding mid-stream", which was false and sent the diagnosis after a hung endpoint: the one call of that burst that did finish reported its first output after 27.9 minutes and 33,920 output tokens.

ModelCallTiming now records SilentUpdates (updates without output before the first one with) and LastUpdateMs, so the ledger tells a thinking endpoint from a silent one, and the cap's sentence is read off the interrupted call (ThreadExecution.RoundCapDetail): "The model was still reasoning and had not started its answer … use a lower effort or a smaller input", with the marker [ThreadExec] ROUND_CAP_WHILE_REASONING at Warning. A call that never answered, and a cap that fired between model calls (during a tool), each get their own sentence too. The first sentence — "AI streaming exceeded the maximum round duration …" — is unchanged.

Where the code is