Wrong answers — reporting the chat's mistakes
/feedback is for formal defects: a broken link, an error card, a timeout, a button that does
nothing. It hands anything that reads as technical to the Systemorph triage agent, which may file
a GitHub issue. That is the right addressee for a defect and the wrong one for this:
"The coach said it doesn't know my hourly rate. It is in the Firmenprofil."
Nothing is broken. The chat gave a poor answer — and whose problem that is depends on why it
was poor. Sent through /feedback, the words "doesn't know" and "wrong" match the technical
classifier (FeedbackHandover.LooksTechnical) and the report becomes a GitHub issue on a code
repository, where nobody can act on it. /wrong-answer (Feedback/Skill/wrong-answer) exists so
the report lands with the right person, carrying the evidence that person needs.
What the chat can know about its own answer
Every assistant message is a ThreadMessage node under the thread, and it already records what a
reviewer needs: agentName, modelName (and requestedModelName when a different model answered
than was asked for), harness, every toolCalls entry with its arguments and result, and
updatedNodes. The skill reads that cell — the user does not have to describe it.
It also runs under the user's identity. A node the skill cannot get is one the answering agent
could not read either. That single check separates "the chat is broken" from "the user lacks a
grant", and it is the reason the skill classifies by checking rather than by what the user
believes.
The kinds, and who each one belongs to
| Kind | What happened | Decided by | Addressee | What they do with it |
|---|---|---|---|---|
fabricated |
asserted a fact, number, name or path the mesh does not hold | a search for the claim finds nothing | the agent's owner — the partition hosting the agent node ({Module}/Agent/x → the module's maintainers; {viewer}/Agent/x → that user) |
tighten the instructions (a grounding rule: cite the node or say you cannot find it), reconsider the ModelTier, add the case to the agent's eval set |
no-access |
said "I don't know", the user says it is in the mesh, and the skill cannot read it either | get of the named node is refused |
the space owner of that node | grant, or explain why not. Nothing for the AI team — the chat could not have known |
retrieval |
the node is readable, a Search/Get in the answer's tool calls should have found it and did not |
get succeeds AND a missing search is on record |
the AI engine (src/MeshWeaver.AI — search, indexing, chunking), via the normal triage hand-over to GitHub |
fix, and pin it with a test — this is the one kind that is a platform defect |
not-found |
the node is readable and the answer never looked | get succeeds AND no such tool call |
the agent's owner | an instruction that says search before you say no |
wrong-source |
faithful to a node whose content is wrong or stale | the cited node says what the answer says | the content owner of that node | fix the node; the agent was right |
unasked |
created, changed or deleted something the user did not ask for | updatedNodes names nodes the request did not cover |
the agent's owner | a scope rule in the instructions; possibly a plugin removed from the agent |
other |
none of the above | — | the Inbox reviewer | classify by hand |
Only retrieval is filed as category: bug and handed over after Submit. Every other kind is
filed as category: answer, which FeedbackHandover.ShouldHandOver keeps in this instance's Inbox
regardless of how technical the message reads — the addressee is a person on this instance, not a
repository.
What the report carries
The same Feedback/Feedback node /feedback files, so preview, Submit, the Inbox and the review
lifecycle are shared. The difference is the content: the user's account verbatim in message, the
thread and message in sidePanel, and a fixed labelled block in extraContext:
Kind: not-found
Thread: sglauser/_Thread/2026-09-18-1319
Message: sglauser/_Thread/2026-09-18-1319/a7f3
Agent: AgenticOffice/Agent/OfficeCoach
Model: claude-sonnet-4-5 (requested: same; harness: native)
Tools called: 1 — Get(AgenticOffice/01-Willkommen) → ok
Answer (excerpt): "Ich weiss leider nicht, welchen Stundenansatz du verrechnest…"
Expected: "Konzept-Ansatz CHF 150 steht im Firmenprofil"
Evidence: get sglauser/AgenticOffice/Firmenprofil → readable; no Search in toolCalls
Addressee: agent owner (AgenticOffice/Agent/OfficeCoach — the course's maintainers)
Labelled lines rather than typed fields, deliberately: the volume does not yet justify a schema
change, the card renders extraContext as-is, and an aggregation can parse Kind: / Agent: /
Model: without a migration (WrongAnswerReport, Feedback/Feedback/Source, is that parser).
Agent: holds the agent's path — the user cell's agentName — because it is the key reports
are grouped by and the address an eval case runs against. When the volume justifies a schema
change, the fields to add to FeedbackContent are exactly those five.
What to do with the reports
Routing is the first use, not the only one. In order of value:
- Route to the addressee (the table above). Today the Inbox reviewer does this by reading
Kind:andAddressee:off the card and moving the item to Triaged;retrievalalone routes itself. The step to automate next: raise aNotificationsatellite on the agent node for the agent-owner kinds, so a course author sees their coach's misses without watching a shared Inbox. - Turn a report into an eval case — shipped; see the next section. A report is a labelled example — this prompt, in this context, must ground its answer in node Z / must not assert F — and replayed whenever the instructions or the model change, these cases are the only defence against fixing one hallucination and introducing another.
- Aggregate by model and by agent — shipped; see the weekly digest
below. A count of
answerreports perModel:over a window is the cheapest quality signal the platform has, and it should informProvider/AiModelTiers— a model that fabricates on theUtilitytier is doing exactly the job that tier must never fail at. PerAgent:, the same count says which instruction sets need work. - Close the loop. The reviewer's Detail view and the Triaged / Resolved / WontFix lifecycle already exist; a Resolved report whose addressee changed something should say what, so the reporter learns the chat got better because they said so.
Eval cases — a report becomes a repeatable test of the agent
A wrong answer is a sample of a rate, not a defect: the same agent with the same instructions
invents something else tomorrow on another prompt. So a category: answer report is never handed
over per report (FeedbackHandover.ShouldHandOver); it becomes an EvalCase
(Feedback/EvalCase, compiled live from Feedback/EvalCase/Source) — Plugins#2288, PR 1.
Lifting. After Submit, the Compose card of a /wrong-answer report offers Make this an eval
case. It creates a Feedback/EvalCase node under Feedback/_Evals/{agentId}/{case} (as
System; agentId is the agent path with / → _, so Agent/ExecutiveAssistant groups under
Agent_ExecutiveAssistant) carrying ONLY source — the report's path. The case's own hub lifts
the rest when it activates (EvalCaseLifting.LiftOnActivation, the same shape as the hand-over on
activation): it reads the report and the thread the report names, and fills
| Field | From the report |
|---|---|
agentPath |
Agent: |
model |
Model: — the model that answered, before the parenthesis; auto (a router) pins nothing |
prompt |
the user message that precedes Message: in the thread — verbatim |
contextPath |
the report's mainNodePath |
mustGround |
the node paths in Evidence: |
mustNotAssert |
the telephone-shaped numbers in Answer (excerpt) that are not the expected value (v1 lifts numbers; any other invented value is edited onto the case by hand) |
mustAssert |
Expected: — the one number in it, or the line whole |
kind |
Kind: |
confirmed |
false — the classifier is usually the accused agent; a person confirms in the editor |
Any client can do the same with one create of that node type with only source set — the lift
is the hub's, not the button's. Not under the agent node: module agents live in system-synced
partitions where a live write is reverted at the next sync, and built-ins are served from the
engine's content pack. Not in the reporter's space: a case outlives the reporter.
Running. The Eval area (the case's default view) shows the definition, the last run and a
Run button. A run starts runs (default 5) real conversations with the agent, one after the
other, under the case ({case}/_Thread/…), on the case's model when one is pinned — the
composer's explicit selection, which the engine honours over the agent's tier and the deployment
default; a headless round decides its model before the thread hub has loaded its agents, so the
tier alone does not reach it — else on whatever the engine chooses for any thread, waits for each completed reply
(ThreadFlow.ObserveResponses) and scores it by three deterministic checks on the reply's own
ThreadMessage record — no model-as-judge (EvalScoring). Whose identity: the thread node is
created under a System scope (the case lives in the shared space, which the viewer cannot write),
but each round runs as the person who pressed Run — the AI engine never runs a round as an
infrastructure identity (HubThreadExtensions.CaptureSubmitter rejects the ambient System context
and takes the real circuit user; ThreadSubmission dispatches under that submitter) — so the agent
reads the mesh with that person's access, and lastRun.ranBy records who, because two runs are
comparable only when it agrees. The checks:
- Looked —
toolCallsholds aGet/Searchwhose arguments or result name amustGroundnode; - Did not invent — the text does not contain a
mustNotAssertvalue: numbers compare by their last nine digits (the subscriber part survives+41,0041,(0)), text after case and whitespace folding; - Correct — if
mustAssertis set, the text contains it.
The run is summarised as rates: check 2 must be N/N; checks 1 and 3 report a rate against a
threshold (0.8 and 0.6 until thresholds are configured per agent). The last run is written back
onto the case (lastRun) with every thread's score and a link to it, so a reviewer opens the answer
a score came from. With no usable language model on the instance — or none for the pinned
model — the run starts nothing and says so; it never fabricates a result. A reply that is the
model-configuration error, or one the engine answered on another model than the pinned one, is a
non-answer, not an invention.
What is next (Plugins#2288, PR 2): a nightly runner on an instance that holds provider credentials (models drift with no commit). Whether the CI gate mesh can hold provider keys, so cases run on a PR that changes an agent's instructions, is open.
The weekly digest — one standing issue per agent
The second consumer of answer reports (Plugins#2288, PR 3) is the triage agent, which drains a
queue on a cadence — but not one item per report. Once a week the Feedback module folds every
report of the last completed ISO week into a digest (Feedback/Digest, compiled live from
Feedback/Digest/Source) and posts one signed digest event per agent; the triage agent
keeps one standing issue per agent current from it.
When and where. The schedule (DigestSchedule.RunWeeklyDigest) runs on the hubs of
Feedback/Digest — the shipped Feedback/Digests node and every digest — checking hourly whether
the last completed week is due (Monday 06:00 UTC) and not yet recorded. It then makes the ONE
declared cross-partition read this module allows itself (nodeType:Feedback/Feedback partitions:all,
once a week — a report lives in its reporter's space or in Feedback/_Submissions), reads every
eval case, builds the digest and CREATES Feedback/_Digests/{week} as System, then reads the new
node once so its own hub comes up and posts; a second hub loses the create and does nothing. Both
reads are asked as system BY VALUE and fail closed (a timer tick has no ambient identity), and both
carry an explicit page bound (200, newest first — the surface's upper end): a full page is recorded
on the digest as truncated, a lower bound said out loud rather than a silent short count.
A hub ticks only while it is activated, and a per-node hub is activated only by something reaching
its address — so every Feedback/Feedback hub wakes the schedule node when it activates
(FeedbackDigestWake): the reports themselves guarantee that a week WITH reports has a running
schedule, and a missed quiet week is caught up at the next activation with nothing to post.
What it counts (WrongAnswerDigest.Build, pure): per agent (Agent:) the SUBMITTED reports
(a draft — the preview the skill files first — is not a report until Submit) by Kind:, the
models (Model:), the count of the week before as the trend, the eval cases lifted from this
week's reports and, over all the agent's cases with a last run, how many passed — the health line;
per model, across agents, the counts by kind. A report without a labelled block is counted under
the unknown agent, never lost.
What it posts (DigestSchedule.PostOnActivation). The digest's own hub, seeing itself
unstamped, sends one event per agent through the same route the feedback hand-over uses
(FeedbackHandover.Deliver — the configured control inbox, or the local target on the control
instance) and stamps postedAt, route and the accepted identities; that stamp is what makes the
post run once. The event is feedback-shaped so the same intake records it: event: digest,
feedbackPath: Feedback/_Digests/{week}/{agentId} (the identity — a redelivery is one item),
message the agent's digest, extraContext the same as JSON, url the digest. A week with zero
reports posts nothing — the digest is still written, as the record that the week was looked at.
With no route configured the digest stands and says so.
What triage does with it (Hosting/Skill/triage §5). The intake carries the standing issue
forward: an item's id is instance + feedbackPath and the path carries the ISO week, so the same
agent's items of the previous twelve weeks are spelled — never searched, so no page limit or
ordering can hide one — and the first issueUrl found lands on the new item. The agent then reads
one field: issueUrl set → CommentIssue the week's numbers onto it; empty → FileIssue it once
— title Wrong answers — , labels feedback and wrong-answer-digest, in the repository
that owns the agent (never the fleet's inbox Systemorph/MeshWeaver.Feedback). retrieval counts are informational:
those reports were filed as bugs when they were submitted. The ledger counts a digest on its own
column.
Reading it. Feedback/Digests lists every week; a digest's Digest area shows the tables the
events carried and what was posted. The area is a template that declares both shapes at once and binds
the node's projection into /data/digestView (DigestLayoutAreas), so neither waits on the node
or the index query before it renders.
What this is not
- Not a place to fix things. The skill never edits the node, never grants access, never re-runs the search and presents that as the answer. It records and classifies; a person acts.
- Not a verdict. The agent filing the report is usually the agent being reported. Its classification is a hypothesis with its evidence attached; the reviewer decides.
- Not a substitute for
/feedback. A broken preview, an error card, a timeout:/feedback.
Tests
Feedback/Feedback/Test/FeedbackTests.cs—HandOver_BugAndIdeaAlways_PraiseNever_QuestionWhenTechnicalpins thatcategory: answeris never handed over, however technical the message reads, and that aretrievalreport filed asbugis;EvalCaseStub_IsUnderTheAgentsEvalsWithOnlyItsSourcepins that the Compose card creates the case typedFeedback/EvalCase, under the agent's_Evalsnamespace, with only its source.Feedback/Digest/Test/DigestTests.cs(run by the type'sTestsarea) — the window (the last completed ISO week, due from Monday 06:00 UTC), the counts by agent, kind and model, the trend, the cases lifted and their last run, a quiet week posting nothing, the rendering, and one feedback-shaped event per agent;Hosting/Deployment/Test/TriageIntakeTests.cspins that adigestparses, maps to its agent's item and carries the standing-issue rule.Feedback/EvalCase/Test/EvalCaseTests.cs(run by the type'sTestsarea) — the lift from the labelled block (every field; a/feedbackbug is refused; an invented value is never the expected one), the case path, the three checks (looked only on aGet/Search, digit-insensitive invention, correct by digits or n/a) and the summary (not-invented unanimous, the rest rates, zero runs never a pass).- Follow-up: an e2e beside
e2e/feedback-skill.spec.tsthat drives/wrong-answerin a thread and asserts the draft'sextraContextcarries every labelled line.