Public Web Presence
The portal serves two audiences from one process: people who sign in and work, and everybody else, including search engines. This page is the design for the second audience. It replaces a separate static marketing site, and it is the reference for every question of the form "why is this page not on Google".
What was measured on 2026-09-11
The brand hosts had already collapsed into one: systemorph.com, www.systemorph.com,
meshweaver.cloud and portal.meshweaver.cloud all answered 301 to www.meshweaver.cloud, a
static one-page nginx site with no meta description, no Open Graph card, no robots.txt and no
sitemap. It was the only MeshWeaver page Google showed. Google also still listed the old WordPress
addresses on systemorph.com (/team/, /contact/, /imprint/, /privacy-policy/), every one of
which landed on the homepage, which a search engine files as a soft 404.
The public portal instance (here portal.example.com) had good crawler plumbing in its <head>: a per-page title,
description, canonical link, Open Graph card, Course and Product JSON-LD for store items, a real
robots.txt and a sitemap of 103 public roots. It had zero pages in Google's index, and the
reason was in the <body>:
$ curl -sA Googlebot/2.1 https://portal.example.com/Doc/Architecture/Localization
<title>Localization · Memex</title>
<link rel="canonical" href="https://portal.example.com/Doc/Architecture/Localization" />
…
<body>
<div id="components-reconnect-modal"> Reconnecting… </div>
<script src="_framework/blazor.web.js" autostart="false">
</body>
A Blazor Server page ships an empty body and fills it over the SignalR circuit. Googlebot runs JavaScript, but it does not hold a circuit, so it indexed nothing. Three separate defects stacked up behind that one symptom:
- The only body source was the node-level mirror.
SeoNoScriptBodyrenderedMeshNode.PreRenderedHtmlinside<noscript>. That mirror is set by the file-system and partition storage readers, and only when the content deserialises as a typedMarkdownContent. The page resolver issues a mesh-widepath:query, which on the partitioned Postgres deployment runs through the cross-schema reader, and that reader never sets the mirror. A plugin cover (PluginContent.body) never had one to begin with. So even the noscript fallback was empty. - The interactive page skipped prerender for strangers.
ApplicationPagefetched cached HTML during prerender only for authenticated visitors, on the theory that serving it to a stranger could leak a page the anonymous gate would later refuse. The gate already decides the head; the body had simply never been wired to the same decision. - The sitemap stopped at the roots. It listed partition roots of three node types, plus each
store plugin's declared
publicSegments, read through the raw storage adapter from an anonymous HTTP entry. On the partitioned deployment that read returned nothing, so no course chapter was ever listed; and nothing belowDocwas ever a candidate at all.
Smaller findings in the same sweep: a path that does not exist renders its nearest ancestor with
HTTP 200 (the canonical link limits the damage, the status does not); HEAD answered 405 from
every page; the default site name was "Memex Portal" because Portal:SiteName was unset on the
public instance; the signed-out landing (/welcome) was a generic sign-in page with half of its
copy hard-coded in English.
The shape
One process, two hosts, one address per page.
| Host | Role | Indexed |
|---|---|---|
www.meshweaver.cloud |
Public. The landing, the documentation, the store, every course cover and free chapter, every public space. Canonical for all of it. | yes |
the portal host (portal.example.com) |
App. Sign-in, the workspace, /api, /mcp, gRPC, the plugin registry. A stranger asking it for a public page is sent to the public host with 301. |
no |
| the brand hosts | 301 to the public host, per path, so the old WordPress addresses land on real pages. |
— |
Why not rename the portal host to www: the Entra, GitHub, Google, Apple and LinkedIn redirect
URIs, every other deployment's plugin-registry URL, every MCP client configuration and every shared
link name the app host, and each of those is its own outage class. The two-host shape gets the same
public result and leaves that decision open.
Configuration
Two keys, both read through PublicSite and both off by default so a single-host deployment is
unchanged:
Portal:PublicHost— the public host name (www.meshweaver.cloud). When set, the canonical link,og:url, every sitemap<loc>and theSitemap:line ofrobots.txtare built onhttps://{PublicHost}whichever host served the request; a request on any other host is on the app host, whoserobots.txtreadsDisallow: /and whose pages carrynoindex.Portal:LandingPath— the node a signed-out visitor sees at/(the landing Space). Unset, the root stays the portal's own welcome route.Portal:AuthHost— the host that owns SIGN-IN (the app host, e.g.portal.example.com). An OAuth challenge buildsredirect_urifrom the host the request arrived on, so a sign-in started on a brand host asks every identity provider to redirect to a host none of them has registered — measured onwww.meshweaver.cloud, where Microsoft, Google and LinkedIn were each handedhttps://www.meshweaver.cloud/signin-*while the registered value is the app host. With this set,UseAuthHostRedirectsends a sign-in that STARTS anywhere else to this host with a temporary redirect, query intact, soredirect_uriis always the registered one and each provider needs one registration, not one per brand host. Only/auth/loginmoves; a/signin-*callback never does, because it carries the correlation and nonce cookies of the host that issued the challenge. Unset, nothing is redirected — a single-host deployment is untouched. The portal registers the middleware besideUsePublicHostRedirect.
Publicness is decided by one gate, per node
Nothing here has its own notion of "public". A page is public because its node carries an
Anonymous Read grant, and AnonymousGate is the single instrument that reads it. The head, the
body, the sitemap, the Open Graph card and the app-host redirect all ask that gate for the exact
node in question and withhold everything on a refusal or on an undetermined answer. That is what
makes descending the tree safe: the installer expresses a course's free and paid chapters as a root
grant plus a deny on every chapter that is not free (see PackageInstaller.EnsureDeclaredAccess),
and a commercial plugin's cover can be public while its content is not. Asking per node is both the
fix for the empty sitemap and the reason the body can be served.
The app-host redirect asks as the stranger. UsePublicHostRedirect sits before
UserContextMiddleware, so a request that only bounces never mints a guest identity. That also
means no AccessContext exists yet when it decides. SeoResolver.Resolve takes the ambient identity
for its owner read, so the decision is run explicitly as UserContextMiddleware.ResolveHttpCaller.
Every request that reaches the decision is unauthenticated, so that identity is the well-known
Anonymous one. This widens nothing: Anonymous reads exactly what the gate admits.
Without the explicit identity, the owner read left portal/reads-{meshId} with none. The never-null
guard refused it (message=GetDataRequest, target=Store was posted with no AccessContext). The
resolver's fail-open then read the refusal as "not public", and the app host served every public
page with 200 instead of redirecting it (MeshWeaver#5227, measured on the public instance:
every synthetic-probe run logged app host: 200 -> (no redirect) within a second of a
target=Store refusal). PublicHostRedirectAsksAsTheStrangerTest runs the default decision over a
real mesh with no identity anywhere. The older PublicSiteTest stubs the decision, so it could not
see this.
The body is in the first response
SeoPageData.Body is the page text a crawler reads: the node's current markdown (content for a
markdown node, body for a Space or plugin cover, or a bare string) rendered by
MarkdownBody.Render, which calls MarkdownViewLogic.Render, the same renderer the interactive
markdown view uses. The signed-in prerender read (IMeshService.GetPreRenderedHtml) reads the node
from its owner and uses this same source-first helper, so an eventually consistent query snapshot
cannot keep showing the source from before an edit. Both pass the
node path, so relative links and embeds resolve against the same page. An edit therefore reaches
both views from the same source, even when the node still carries HTML generated before the edit
or before a renderer update. An explicitly empty source renders empty; it never revives the old
page. The node's mirrored HTML, then the content's prerenderedHtml, are fallbacks only for nodes
that carry no markdown source. Neither cache records which source or renderer produced it, so
neither can establish freshness. SEO resolution refreshes the admitted node from its owner, then
rechecks anonymous access before returning its current title and body. A signed-in caller's own
read grant cannot turn a newly private page into public HTML. The owner read carries the original
HTTP or circuit viewer across the asynchronous permission result; it does not rely on that result's
thread retaining an ambient identity. It is computed only for a node the
gate admitted, and it is rendered visibly in the server response, not inside <noscript>:
Googlebot renders with JavaScript on and may ignore noscript content. The document, head and body
await the same per-request resolution, so they use one current source and access decision.
Legacy Space content stored as a serialized JSON object is recovered before rendering, just as
in the live Space view: body wins over content, including an explicitly empty body. A bare
JSON document on a Markdown node remains authored text.
Plain public documents finish on the server
An anonymous visitor requesting an exact, authored Markdown page receives that HTML as the final
page, without starting Blazor, Monaco or the reconnect UI. The same applies to authored Spaces in
the configured landing subtree, and to Spaces that explicitly exclude their live contents
catalog. PublicPageResponse makes the request decision and PublicPageRendering identifies
supported document bodies in MeshWeaver.Plugins. There is no separate static build: each
request reads the current node and renders its current source with the shared markdown renderer.
An edit therefore reaches the next request without refreshing a second copy of the page.
Public pages can still be interactive. Framework embeds, executable cells, Mermaid and math keep Blazor, as do custom node types, applications, layout-area routes, satellites, empty bodies and signed-in sessions. Query parameters that select another view also keep the interactive route; ordinary tracking parameters do not change presentation. Gated pages still go through their existing access and sign-in flow. Navigating from a circuit to an eligible static page performs a normal document load. Where interaction is needed, the visible server-rendered article remains the initial response and the circuit replaces it on hydration.
The interactive Space view follows the same source-first priority and preserves an explicitly
empty body. The generic Overview/Data markdown body uses the shared renderer on each live node
emission too, keeping its source, HTML and node path together when embedded content needs Blazor.
Navigation derives HTML for its current snapshot without persisting that result: a
delayed cache write must not overwrite the HTML after another author has edited the source.
The landing page and its descendants request the existing showHeader=false presentation, so
the authored hero and headings survive hydration without an additional Space title. The SEO head
and interactive public page titles use the same PublicPageTitle formatter; the landing title
also follows the node stream when its name changes. These interactive changes live in
MeshWeaver.Plugins, alongside the Space view and the portal pages.
The sitemap descends
For every root the gate admits, SeoEndpoints.EnumeratePublished lists the root's descendants
whose node type is a page — Markdown, Space, Store/Plugin, Store/Catalog, Edu/Module,
Edu/Page — skipping any path with a satellite segment (_Thread, _Access, _GitSync, Source,
Test, Release), and asks the gate about each. A reinsurance plugin's partition carries hundreds
of amount types, cashflows, source files and release markers; none of them is a page, and listing
them would bury the twenty that are. The listing runs as System and unbounded (limit:all — a
stated cap would silently clip the one list that claims to be complete; a stale negative here costs
a URL, never a leak, because every candidate still passes the gate), the gate checks run eight at a
time, and the same enumeration feeds Administration → Published to the web, so the list a
person reads and the list a crawler gets cannot drift. The sitemap protocol bounds one file at
50,000 URLs; a deployment that publishes that many pages needs a sitemap index, which is a
follow-up, not a reason to cap the query.
The rest of the crawl surface
- A path with a remainder beyond the resolved node (a layout-area route, or a missing page that
fell back to its ancestor) carries
noindexand a canonical link to the node it actually rendered, so a fallback never competes with the real page. HEADis answered as aGETwith the body discarded (PublicSite.UseHeadAsGet).- The public host's
robots.txtkeeps/login,/welcome,/api/,/_blazorand/dev/out. - Course covers keep their
CourseJSON-LD with Systemorph as the provider and the price as anOffer; the English and German twins of a course point at each other withhreflang.
Verifying a deployment
Each check is one command, and each has a negative control: run it on the app host too, where the answer must differ.
# the body: the article text, not the reconnect modal
curl -sA Googlebot/2.1 https://www.meshweaver.cloud/Doc/Architecture/Localization | grep -c '<article'
# the sitemap descends: a chapter and a doc page are listed
curl -s https://www.meshweaver.cloud/sitemap.xml | grep -c 'AgenticPrimer/01-TheMagicWish'
curl -s https://www.meshweaver.cloud/sitemap.xml | grep -c 'Doc/Architecture/'
# the app host sends strangers away and keeps crawlers out
curl -sI https://portal.example.com/Doc | grep -i '^location: https://www.meshweaver.cloud/Doc'
curl -s https://portal.example.com/robots.txt | grep -c 'Disallow: /$'
# HEAD is answered
curl -sI https://www.meshweaver.cloud/Doc | head -1
It is watched continuously, and the 301 is the assertion
prod-synthetic-probe.yml runs those two app-host checks against every portal every 15 minutes: the
public page must answer 200 on the public host, and the app host must answer 301 to exactly
that URL — never curl -L, because following the redirect makes a real outage and a correct
canonicalisation produce the identical green.
🚨 Which host is public is a per-deployment decision, so the probe READS it rather than pinning
it, from the Sitemap: line of the portal's own robots.txt — the one line that names
CanonicalBaseUrl on either host. That makes the two-host split machine-readable at probe time and
a single-host portal (which declares itself) assert 200 with no redirect, with nothing to keep in
step in CI. Two consequences for anyone editing SeoEndpoints.MapSeo: that line is a contract, and
so is Disallow: / on an app host that differs from the public one. Before this, the probe asserted
200 for a public page on the app host and went red for three runs on a healthy portal the day
Portal:PublicHost was set — the shape of a check that is true about a broken portal and false
about a working one.
Search Console's URL inspection, "rendered HTML" view, is the acceptance test for the whole design: it must show the article text for a documentation page, a course cover, a course lesson and a plugin cover. Indexing before that view shows text only fills the report with "crawled, currently not indexed".
Related
- Access Control — the anonymous grant and the gate.
- Localization — the landing renders as authored, in English and German.
- Operating from the portal — how the ingress change that adds the public host is rolled.