Commit Graph

1270 Commits

Author SHA1 Message Date
sencho-quartermaster[bot] 8fd02ef39b chore(main): release 0.84.2 (#1127)
Co-authored-by: sencho-quartermaster[bot] <275163604+sencho-quartermaster[bot]@users.noreply.github.com>
v0.84.2
2026-05-20 13:06:01 -04:00
Anso 982c7830b1 fix(mesh): address audit follow-ups (early-data, recompose, admiral gate, error codes, docs) (#1126)
* fix(mesh): buffer early TcpData on reverse-relay path

The reverse-relay code dropped the local-shaped reservation and any
TcpData frames buffered in it before acceptReverseRelay had wired up
the target stream. A peer that sent a request body immediately after
tcp_open_reverse lost those bytes when the target lived on a third
node, since reverse_local buffers them but reverse_relay did not.

Carry the reservation through acceptReverseRelay: transplant its
pendingData and pendingBytes into the new reverse_relay state, gate
TcpData writes on targetOpen, and flush buffered frames into the
target stream before sending tcp_open_ack. Mirrors the reverse_local
pattern. Regression test fires tcp_open_reverse plus an immediate
TcpData while ensureBridge is in flight and verifies the bytes
arrive intact, in order, before the ack.

* fix(mesh): recompose affected stacks when node-level mesh is disabled

disableForNode used to clear DB rows and override files but leave the
running containers attached to sencho_mesh with stale /etc/hosts
alias entries until an operator redeployed every stack by hand.

Mirror optOutStack: after the existing alias/forwarder cleanup, call
regenerateOverridesAcrossFleet, cascadeRecomposeAcrossFleet, and
triggerRedeploy for each previously meshed stack on the disabled
node so containers detach from sencho_mesh and shed the alias
entries they owned. The disabled node's mesh_stacks rows are deleted
before the cascade so listMeshStacks returns the right set with no
skip tuple required. Route threads the actor through actorFor(req)
for parity with optInStack/optOutStack. Tests cover the redeploy
fan-out, the cascade no-skip-tuple invariant, and the default actor
fallback for non-route callers.

* fix(mesh): require Admiral on the WS proxy-tunnel upgrade

HTTP mesh routes in routes/mesh.ts all enforce requireAdmiral, but
the /api/mesh/proxy-tunnel WS upgrade accepted any node_proxy or
full-admin api_token regardless of the receiver's license. A node
downgraded from Admiral kept serving mesh data-plane traffic to a
sibling central while refusing every mesh management call.

Read the receiver's local LicenseService at the upgrade and 403 when
the tier is not paid+admiral. The check sits after the existing
credential gate and uses LicenseService directly rather than
effectiveTier (which trusts forwarded proxy headers); a remote peer
dialing in cannot be trusted to assert our entitlement. Dialer and
node_proxy token format are unchanged. Three regression tests cover
community-tier node_proxy, skipper-tier node_proxy, and
community-tier full-admin api_token all rejected with 403.

* fix(mesh): handle no_target alongside push_failed in inspectStackServices

proxyFetch throws MeshError('no_target') when getProxyTarget returns
null (pilot tunnel offline, proxy bridge unreachable), but the
inspectStackServices catch branch only matched push_failed. Offline
remotes fell through to the generic 'remote unreachable' error log,
which the Routing tab surfaces as an unexpected fault.

Match both error codes and emit the operator-friendly warn message
that names the unreachable node and the error code. Regression test
spies on console.warn/console.error to pin the branch.

* docs(mesh): align env defaults and forwarder comments with current architecture

SENCHO_MESH_PROXY_TUNNEL_IDLE_MS in .env.example carried the old
five-minute idle-close value (=300000), but the code default is
DEFAULT_IDLE_TTL_MS=0 (persistent tunnel). Copying the example
silently reintroduced the idle-close behavior the dialer removed.

MeshForwarder.ts's leading docblock and inline listen comment still
described host-network mode plus extra_hosts: host-gateway as
required for forwarder reachability. Sencho runs in standard bridge
mode and attaches to the shared sencho_mesh network at a stable IP;
meshed user containers reach the forwarder by that IP directly.

Flip the env default to 0, rewrite the env comment to describe the
persistent behavior and the opt-in for idle teardown, and rewrite
both forwarder comments to match the bridge-network reality.

* fix(mesh): trust forwarded tier on proxy-tunnel WS and remove remote overrides on disable

Admiral entitlement on the WS data plane now follows the same trust model
as the HTTP mesh routes: the central asserts its tier via x-sencho-tier
and x-sencho-variant on the WS handshake, and the receiver trusts those
headers only when the upgrade carries a node_proxy credential. When no
headers are present or the credential is a full-admin api_token, the
receiver falls back to its own local license. Without this an Admiral
central could be rejected by a Community remote and a Community central
could dial a locally-Admiral remote.

disableForNode now routes through removeOverrideFromNode so override
files pushed earlier via applyLocalOverride are removed on remote nodes
via DELETE /api/mesh/local-override/:stack. Sequential awaits are wrapped
in Promise.allSettled to match regenerateOverridesForNode's parallel
push pattern.

Other changes:
- Cover the post-state-swap buffering window in the reverse-relay test
  (TcpData arriving after openTcpStream returns but before target open).
- Refresh stale MeshService comments that still referenced the removed
  sidecar layer and host-network listener model.

* test(mesh): pin removeOverrideFromNode remote HTTP shape

The disableForNode regression test mocks removeOverrideFromNode itself,
so a regression inside the helper would not be caught. Add a narrow
contract test that spies on global fetch and asserts the request shape:
DELETE /api/mesh/local-override/:stack against the resolved proxy
target, with Authorization Bearer plus the x-sencho-tier and
x-sencho-variant headers. Also covers the encodeURIComponent path and
the swallowed-network-error behavior the disable cascade relies on.

* test(mesh): use createTestApiToken helper in proxy-tunnel api_token gate test

The full-admin api_token branch of the WS Admiral gate test inlined the
canonical generateApiToken + sha256 + addApiToken triple that already
lives in the createTestApiToken helper. Switching to the helper removes
the duplicated insertion logic and aligns this test with the helper used
by the other api_token call sites.
2026-05-20 12:53:29 -04:00
Anso 08caa914ce docs: v1 docs refresh (batch 2) (#988)
* docs(atomic-deployments): refresh page around current UI and behavior

Rewrites the page to match the v1 docs refresh template. Corrects
several factual errors against the current code, fills in missing
detail, and adds a screenshot of the rollback overflow menu.

Notable corrections:
- Scheduled tasks do not run atomically; only stack editor Deploy and
  Update, App Store installs, webhook triggers, and image auto-updates
  pass the atomic flag through to ComposeService.
- Rollback lives in the stack editor's More actions overflow menu, not
  on the action bar directly. The backup timestamp renders as a
  sub-line of the menu item.
- Health probe is a 3-second window with an exit-code check on every
  container labelled with the compose project name; describe this
  exactly rather than as 'waits briefly'.
- Document where backups live (DATA_DIR/backups/<stack>/), why they
  are kept outside the compose folder, and that the slot is one per
  stack with overwrite semantics.
- Document the four streamed log markers users see in the deploy
  progress modal during the atomic flow.
- Add a troubleshooting accordion group covering missing menu entry,
  late crashes outside the probe window, manual-intervention message,
  and the single-slot retention edge case.

* docs(deploy-enforcement): refresh page for v1 and align with current enforcement paths

Update the page to match the current pre-flight gate behavior, the v1 modal chrome on the
block dialog, and the AccordionGroup troubleshooting pattern used across the v1 docs.

Drift items corrected:
- Replace the broken vulnerability-scanning/deploy-blocked-dialog.png reference with three
  fresh captures under docs/images/deploy-enforcement/ (policy list, policy editor, block
  dialog).
- Drop "Recreate from the stack actions menu" and the git-source apply pre-flight claim;
  neither path runs the gate.
- Add bulk label deploy and the auto-update scheduler to the enforced code paths, with a
  dedicated subsection for the auto-update interaction (alert-and-skip, not 409).
- Drop the false claim that severity chips in the block dialog are clickable; the dialog
  is informational.
- Document the compose-parse-fails-closed branch with its synthetic violation label.
- Refresh dialog copy to reflect the v1 ModalDestructiveHeader (kicker, title, button
  variants).
- Convert the troubleshooting Q&A into AccordionGroup blocks and add accordions for the
  compose-parse-error case and the auto-update-skipped case.
- Quote the verbatim audit-log summary format.

* docs(blueprints): refresh against v1 UI and add federation/state-review coverage

* docs(git-sources): refresh page against v1 UI and current behavior

Rewrites the page against the v1 docs refresh template (Note tier-gate,
sectioned anatomy, AccordionGroup troubleshooting), aligning prose with
the live UI labels and the current code paths.

Corrections:
- Authentication toggle reads "Public (no auth)" / "Personal Access
  Token" (not "None"), and apply mode "Auto-write files" (not
  "Auto-write").
- Diff dialog kicker is GIT . PULL PREVIEW; local-edits state opens an
  Overwrite local edits? confirmation modal whose primary button is
  Overwrite and apply.
- Sidebar pending indicator is a small GitBranch icon, not a brand-color
  dot, and the image-update dot takes priority over it on the same row.
- Pending update banner appears in the panel; Review re-fetches the
  commit and opens the diff (no client-side payload caching).

Adds coverage for:
- Anatomy of the panel (pending banner, form, last-applied stat strip,
  footer actions).
- 10-second webhook debounce window.
- Pending compose/env content is encrypted at rest in the database, not
  just the token.
- Auth/host failures map to HTTP 400, never 401, so they do not sign
  the user out.
- Per-stack lock serializes pull, apply, and create-from-git so a
  webhook firing during a manual apply waits rather than racing.
- Compose validation has a 10-second budget; clone fetches have a
  30-second timeout.
- New troubleshooting accordion for Pending commit has changed since
  this pull was fetched.

Recaptures all five screenshots from the v0.74.x production node,
signed in as admin: panel, create-from-git tab, pull-preview diff
dialog, sidebar GitBranch pending icon, webhook Action select with
Git source sync highlighted.

* docs(stack-labels): refresh page for v1 sidebar grouping and fleet-action surface

- Lead with the v1 behavior the previous page did not cover: the sidebar
  groups stacks under collapsible label headers (PINNED first, label
  buckets sorted by stack count desc then name asc, UNLABELED last)
  with a count chip per group. Trailing colored dots on each row
  (max 3 + N overflow, paid-only) supplement the headers.
- Drop the stale claim that a label-pill filter bar lives between
  search and the stack list; that UI no longer exists.
- Drop the right-click-on-pill bulk actions table (Deploy all / Stop
  all / Restart all). The legacy per-node action endpoint stays in
  the backend but no longer has a UI binding, so the page documents
  only what users can click today.
- Document the two Skipper+ Fleet Action cards: Stop fleet by label
  (name match across nodes, autocomplete, per-node breakdown,
  HTTP 429 on per-node concurrency) and Bulk label assign (per-node,
  replace semantics, clear on empty selection).
- Document the inline 'New label' form inside the stack right-click /
  three-dot Labels submenu, the Settings - Advanced - Labels masthead
  N/50 stat, the LABELS - NEW / EDIT modal kickers, and the
  LABELS - DELETE - IRREVERSIBLE confirmation copy verbatim.
- Document the Fleet Overview Tags multi-select filter (filters by
  stack labels aggregated across nodes), with cross-link to fleet-view.
- Capture every screenshot fresh from production signed in as admin:
  sidebar-grouping, context-menu-labels, inline-create-form,
  settings-labels, create-label-dialog, fleet-tags-filter,
  fleet-actions. Drop the now-stale sidebar-with-labels,
  sidebar-filtered, and bulk-actions-menu captures.

* docs(dashboard): refresh page for v1 layout (status masthead, gauges, fleet heartbeat, restart map)

Aligns docs/features/dashboard.mdx with the redesigned Home tab. Replaces the obsolete
Recent Activity feed coverage with the actual DashboardActivityCard split (Fleet Heartbeat
when remote nodes are registered, Stack Restarts (7d) otherwise) and recaptures every
screenshot from the v0.74.x production node.

* docs(global-search): refresh page for v1 palette

- Note tier and role gating on the Pages list (Auto-Update, Console,
  Schedules, Audit) so the prose matches what the top bar exposes.
- Document the ACTIVE chip on the currently active node row.
- Document the 50-result cap counter and the Searching... loading state.
- Mention the ~250 ms debounce and clarify that filename matching
  includes the file extension.
- Replace stack screenshot with a redesigned capture and add empty-state
  Pages and Nodes captures showing the ACTIVE chip.

* docs(global-observability): refresh page for v1 layout (masthead, signal rail, filter strip, paused-resume chip)

Full rewrite against the current Logs tab and the v1 docs refresh template
(hero Frame, sectioned anatomy, AccordionGroup troubleshooting, refresh-cadence table).

Replaces the single overview screenshot with seven captures under
docs/images/global-observability/ (overview, masthead, signal-rail,
filter-strip, feed-bands, paused-resume-chip, error-only-filter), all from
the v0.75.x production node signed in as admin with PII scrubbed
(profile chip patched to AD, in-feed LAN IPs and third-party hostnames
substituted via DOM injection while the stream was paused).

Aligns prose with the actual UI labels and code:

- Masthead kicker reads LIVE LOGS · NODE · <NAME> with LOCAL for the
  local node; state word toggles Streaming / Idle / Offline; SESSION
  uses uppercase letter suffixes (1H 43M / 0M 12S) per formatUptime.
- Signal rail tile counts are scoped to the 2000-entry buffer and reset
  with Clear; CONTAINERS is buffer-bound, not a monotonic accumulator.
- Filter strip controls quoted verbatim (Stacks · All / Stacks · n,
  segmented controls All / Out / Err and All / Info / Warn / Error).
- Feed row anatomy: severity dot, timestamp, brand-cyan container name
  with stack/container tooltip, message tinted by source. Row tint
  follows detected level, which is regex-based, so an STDOUT line
  containing ERROR: still classifies as ERROR.
- Day bands: NOW, Nm AGO, Nh AGO, calendar date.
- Empty states: two-tier kicker over caption (Awaiting events / No matches).
- Pause keeps the SSE buffer filling up to the 2000-entry cap; resume pill
  reads <n> NEW · RESUME and counts the queue, not total arrivals during
  the pause.
- Download filename and row format quoted: sencho-logs-<ISO8601>.txt and
  [<ISO>] [<stack>/<container>] <LEVEL>: <message>.

Documents behavior the previous page never covered:

- Active-node scoping; node switch resets the stream and the buffer.
- SSE primary transport with 30-second server heartbeat and a 5-second
  polling fallback against /api/logs/global (server-capped at 500 lines
  per snapshot).
- Initial replay of the last 500 lines per container when the SSE
  connection opens, so the feed has context immediately.
- Display limits (2000 client buffer, 300 rendered rows, Showing last
  300 of N overflow notice).
- Refresh cadence table covering UI tick, flush cadence, polling
  cadence, SSE heartbeat, sparkline window, and the Idle threshold.

Adds a seven-accordion troubleshooting block (Offline state, gray Idle
dot, ERROR-without-tint, growing Resume pill, Clear-cutoff lag,
node-switch buffer drop, fleet-wide aggregation expectations).

Tightens the closing Note so it makes clear that Notification Log
Retention does not govern this live container stream.

* docs(alerts-notifications): refresh page for v1 and absorb notification-routing

Full v1 template rewrite of /features/alerts-notifications. Bundles in
the entire Notification Routing page so a reader sees channels, routing,
per-stack rules, and retention in one place; deletes the standalone
notification-routing.mdx and points all five cross-link sites at the new
in-page anchor.

* docs(alerts-notifications): drop "What's not in scope" section

The page should describe what Sencho does, not enumerate what it does
not ship. Users find missing integrations through the Webhook section
and the routing matcher reference; the explicit disclaimer added noise
without adding guidance.

* docs(audit-log): refresh page for v1 layout, expanded action list, troubleshooting accordion

- Clarify that the search/method/date filter strip lives in Table view only.
  Stream view always shows the unfiltered chronological feed.
- Fold the total-entries readout into the card subtitle wording where it
  actually renders, instead of describing it as a separate header element.
- Sharpen the Peak hour off-hours window to the literal 08:00 to 17:59
  working window the tile keys off, plus the 5% / 20% failure-rate tints.
- Note that the Actors tile names a sample actor alongside the new-IP count.
- Expand the example actions list to cover surfaces that have shipped since
  the last edit: per-service stack lifecycle, node cordon/uncordon, fleet
  replica role changes, Sencho Cloud Backup operations, Fleet Secrets, and
  blueprint federation pin updates.
- Correct the Settings path: Settings · Developer · Data retention card,
  Audit log input, Save settings button.
- Add a Troubleshooting AccordionGroup matching the rest of the v1-refresh
  pages: missing tab, filter scope, anomaly thresholds, export cap, and
  retention pruning.
- Replace all four screenshots with fresh captures of the current UI.

* docs(multi-node): refresh page for v1 layout, pilot agent mode, refreshed table columns

Rewrites docs/features/multi-node.mdx against the current product. The previous page predated the v1 Settings hub redesign and the Pilot Agent enrollment model, so it documented only the Distributed API Proxy add-node flow and missed the new Mode, Endpoint, and Labels columns on the Nodes table.

Restructures the page into 13 sections: intro, How it works, the local node, Choose a remote mode (decision table comparing Pilot Agent vs Distributed API Proxy), Add a remote node: Pilot Agent (three steps plus re-enrollment), Add a remote node: Distributed API Proxy (three steps), Switching between nodes, the Nodes table (full column reference), What Settings apply per node (verified against settings/registry.ts), License enforcement across nodes, Editing and deleting nodes, Security (token security, transport encryption, why no application-layer TLS), and Troubleshooting (AccordionGroup matching the v1 template used on audit-log, atomic-deployments, and deploy-progress pages).

Refreshes seven screenshots against the production node signed in as admin, scrubbing IPs and usernames before capture: full Nodes panel overview, Generate Node Token card with a placeholder token, Add node modal in Pilot Agent mode, Add node modal in Distributed API Proxy mode (with the inline plain-HTTP warning visible), Edit modal showing the Regenerate enrollment token card for a pilot agent, Pilot enrollment modal with the docker run command, refreshed node switcher popover, and a close-up of the table columns. Drops the obsolete add-node-form.png, http-warning.png, and per-node-scheduling/ folder.

* docs(fleet-view): refresh page for v1 layout, expanded tabs, cordon, sheet-based updates

- Aligns the Overview, Status, and Node Updates content with today's UI:
  the masthead's `The fleet` headline plus CPU / MEM / CONTAINERS stat tiles,
  the eight-tab strip (Overview, Snapshots, Status, Deployments, Traffic,
  Federation, Fleet Actions, Secrets) with per-tier visibility, and the
  Check Updates surface that is now a system sheet rather than a modal.
- Documents the toolbar (search, sort, filter popover with Status / Type /
  Severity / Tags sections) and the Grid / Topology segmented control
  including the topology graph's status pill (Online / Critical / Offline),
  connector colouring, ReactFlow controls and minimap.
- Documents the per-card surfaces that were missing from the prior page:
  Cordoned badge with cross-reference to Fleet Federation, fleet stack
  label dots in the drill-down, container drill-down rows (state dot,
  badge, image, status, open-in-editor hover button), and the Admiral
  three-dot Node actions menu for cordon / uncordon.
- Documents the Node Updates sheet anatomy (Recheck and Update all (n)
  header actions, four summary cards, node table columns, Update flow,
  reconnecting overlay timing, admin enforcement) and the GitHub Releases
  with Docker Hub fallback resolution path with its 30-minute cache.
- Replaces every stale screenshot with a fresh capture (overview,
  topology, drill-down, status tab, node updates sheet) and removes the
  obsolete files plus the empty docs/images/fleet/ folder.
- Reformats troubleshooting as an AccordionGroup matching dashboard,
  multi-node, and audit-log refreshes.

* docs(fleet-backups): refresh page for redesigned fleet and settings UI

Replace all six screenshots with current production captures. Update
content to reflect the new fleet header card, eight-tab layout, full-
page Cloud Backup settings with header stats, and corrected navigation
paths. Add cloud backup rows to the access control table.

* docs(fleet-backups): convert troubleshooting to AccordionGroup pattern

Match the foldable-accordion pattern used across the v1 docs refresh
batch. Merges the standalone Cloud Backup troubleshooting subsection
into a single Troubleshooting section at the bottom of the page with
seven accordions covering skipped nodes, two restore failure modes,
three cloud-upload failure modes, and a diagnostic logging entry.

* docs(remote-updates): refresh page for v1 sheet, accordion troubleshooting, factual fixes

Rewrites the page against the v1 docs refresh template (Note tier gate,
sectioned mechanism deep-dive, Frame screenshots with detailed alt text,
inline AccordionGroup troubleshooting), bringing it in line with the
recently-refreshed fleet-view, fleet-backups, dashboard, and audit-log
pages.

The page is repositioned as the mechanism deep-dive (prerequisites, what
runs on a node during an update, completion and failure detection,
recovery actions). The full UI tour for the Node updates sheet remains in
fleet-view so the two pages stop overlapping; remote-updates now links
into fleet-view#node-updates instead of restating the table anatomy.

Captures three screenshots from the production node, signed in as admin:
fleet-node-updates.png shows the Node updates sheet with eight nodes and
seven remote updates available; local-update-confirm.png shows the
LOCAL · UPDATE alert dialog with the Cancel and Update & restart buttons;
node-card-update-available.png shows the Opsix card with the Update
available pill and the Update to v0.76.7 outline button.

Corrects several factual claims that no longer matched the current code:

- The remote early-fail threshold is about 3 minutes, matching
  EARLY_FAIL_MS in backend/src/routes/fleet.ts, not 90 seconds.
- The Recheck button sits in the sheet header, not the footer.
- The component is a SystemSheet, so the page now consistently calls it
  the Node updates sheet instead of a dialog, with lowercase "Node
  updates" and lowercase "Update all (n)" matching the live UI.
- Reconnecting overlay polls /api/health every 3 seconds, not "every few
  seconds".
- The local Failed badge surfaces as soon as the helper writes its error
  file, by the 3-minute mark at the latest.

Documents the LocalUpdateConfirmDialog kicker, title, body, and CTA
verbatim, the Triggering... loading state on the Update buttons, the
four completion signals the gateway accepts (version change, process
startedAt change, offline-then-online transition, version at or above
the comparison target after 15 seconds), and the 60-second auto-clear
of the Updated badge.

Drops references to two screenshots that never existed
(fleet-node-updating.png, fleet-node-failed.png); the in-flight and
failed states are described in prose instead, the same way fleet-view
handles them.

* docs(scheduled-operations): refresh page for v1 timeline, fleet-wide update action, sheet-based run history

Rewrites the Scheduled Operations page against the v1 template
(Note tier gate, sectioned anatomy, Frame screenshots, AccordionGroup
troubleshooting) applied to sibling pages in this batch. Captures
seven fresh screenshots against the production node signed in as
admin (timeline, all-tasks, action-picker, create-restart,
create-prune, create-scan, run-history) and removes every legacy
PNG.

Documents the new "Auto-update All Stacks" action that was absent
from the page, extends the Skipper allow-list to all four Skipper
actions (Auto-update Stack, Auto-update All Stacks, Fleet Snapshot,
Vulnerability Scan) and clarifies that the action picker hides
operations the active tier cannot run.

Corrects several factual claims that no longer matched the code:

- Scheduled scan completion is `info`/`scan_finding` on a clean run
  and `warning`/`scan_finding` when findings are present (not
  `info`/`system` as previously stated). Cross-link now points at
  `alerts-notifications#vulnerability-scanning`.
- Lifecycle actions (auto_backup, auto_stop, auto_down, auto_start)
  execute against the local Sencho instance only; only Auto-update
  Stack / All Stacks have a remote-proxy code path. The page
  reinstates the guidance to schedule remote lifecycle operations
  from that node's own UI.
- Run history lives in a right-side sheet with a "Schedules ›
  <task> › Runs" breadcrumb and a Download CSV secondary action.
- Timeline masthead is described in terms of the v1 visual
  (`NEXT 24 HOURS` kicker, italic display heading, monospace date
  range, right-anchored Next pill with countdown, glowing cyan now
  rail, six-tick bottom axis).

* docs(rbac): refresh RBAC & user management page against v1 template

Bring /features/rbac onto the v1 docs refresh template (Note tier gate,
sectioned anatomy, Frame screenshots, AccordionGroup troubleshooting).
Recapture five screenshots from the production node signed in as admin
and remove the three stale captures under docs/images/rbac/.

Corrections vs. the prior page:
- Deployer no longer claims node:read in the permission matrix; the
  backend grants only stack:read and stack:deploy.
- Add the system:registries row (container registry management).
- Document the form as inline below the Add user button (not a modal).
- Note the (you) marker on the signed-in admin's row and the disabled
  delete icon on that row.

Additions:
- Settings nav location and hub-only visibility.
- 2FA reset row action with verbatim modal kicker, title, and body.
- Five-failure / 15-minute MFA lockout behavior and admin reset recovery.
- Token-version session-security table covering deletion, role change,
  password change, and admin 2FA reset.
- SSO password-fields-hidden line quoted verbatim and the per-provider
  Require MFA toggle.
- Audit-log emissions list for every user-management mutation.
- API tokens cross-link explaining the user-vs-machine boundary.
- Scoped permissions section retightened: scoped role picker is
  Deployer / Node Admin / Admin only; resource type is Stack or Node.

AccordionGroup with eight troubleshooting entries covering missing nav,
greyed role options, seat-limit errors, unexpected sign-outs, scoped
deployer mismatches, missing shield icon, re-locking MFA accounts, and
SSO role drift at provisioning.

* docs(2fa): refresh two-factor authentication and admin guide against v1 template

Bring /features/two-factor-authentication and /operations/two-factor-admin
onto the v1 docs refresh template (Note tier gate, sectioned anatomy, Frame
screenshots with descriptive alt text, AccordionGroup troubleshooting,
verbatim modal copy with kicker callouts). Recapture every screenshot under
docs/images/two-factor-auth/ from a fresh session and add six new captures
for surfaces the prior page did not document.

Corrections vs the prior pages:

- Panel rename: Settings -> Account & Security is now Settings -> Account,
  under the Identity group of the settings sidebar. Replaced every
  occurrence on both pages.
- Enrol dialog titles match the current modal: Pair your authenticator,
  Confirm the pairing, Save your recovery codes (was: Set up 2FA, Confirm,
  Save your backup codes). Step rail 01 PAIR / 02 CONFIRM / 03 ARCHIVE
  documented.
- Manual-entry affordance is the always-visible Secret manual entry row
  with a copy icon, not the toggleable Can't scan Show secret key link.
- Confirm step auto-submits on the sixth digit; no submit button. Verified
  in MfaChallenge.tsx and MfaEnrollDialog.tsx and called out explicitly.
- Authenticator-app list trimmed to match in-app copy (1Password, Bitwarden,
  Google Authenticator, or any TOTP app). Authy and Microsoft Authenticator
  dropped because the dialog does not mention them.
- Disable dialog: kicker SECURITY MFA DISABLE, title Turn off two-factor,
  destructive header, Disable button. Replaces the prior Disable 2FA
  paragraph that did not describe the dialog chrome.
- Regenerate dialog: two-step flow with kicker SECURITY BACKUP CODES, Confirm
  identity then New recovery codes, with the verbatim PREVIOUS CODES HAVE
  BEEN INVALIDATED warn rail on the show step. Documented that the dialog
  only accepts a TOTP, not a backup code.
- Per-user SSO toggle label corrected: Require 2FA on SSO sign-in (was:
  Require 2FA even when signing in via SSO). Added the per-provider vs
  per-user distinction on both pages (admins can also enable Require MFA
  on the SSO provider config, which is independent of the per-user toggle).
- Admin reset modal: verbatim USERS RESET 2FA kicker, Reset 2FA for
  <username> title, full-body copy reproduced. Documented that the reset
  bumps the target's token version and invalidates active sessions.

Additions:

- Sign-in throttle: five failed verifications lock the account for 15
  minutes, server returns 423 with Retry-After, UI shows the Retry in MM:SS
  countdown plus Rate limited label. Lockout recovery section explains
  that the counter only clears on a successful sign-in, so retries after
  the window expires re-lock immediately.
- Account panel anatomy section enumerates the three rows (Authenticator
  app, Backup codes, Require 2FA on SSO sign-in) plus the destructive
  Disable 2FA link, and the masthead 2FA on / BACKUP N left chips.
- Recovery codes section now covers all three count states (3 plus, 1 to 2,
  0) with verbatim helper text, tone, and the standalone No backup codes
  left callout that renders at zero. New screenshots for the 2-remaining
  and 0-remaining states.
- Cross-references to the admin operations page (CLI fallback, token version
  rotation, what a reset changes in the DB), the SSO page, and the RBAC
  page (per-provider Require MFA toggle, SSO auto-provisioning).

Troubleshooting on the feature page rewritten as an AccordionGroup with
nine entries: clock drift, wrong account selected, QR will not scan, lost
phone with no codes, lost codes with authenticator, ran out of codes,
unexpected SSO prompt (with both toggle causes), repeated lockout after
the window expires, missing shield icon on Users panel.

The admin operations page also gains the SSO + 2FA two-toggles table so
administrators can answer the per-user vs per-provider question without
context-switching between pages.

Six new images added; six existing images replaced. Total 14 captures.

* docs(rbac,host-console): drop enforcement-boundary detail from tier-gate notes

Operator-facing docs should state tier or role requirements once, in plain
customer-facing language, and leave the enforcement chain to the source.
Two surfaces on the v1-refreshed pages over-specified the gate:

- `features/rbac.mdx::Scoped permissions`: the Note enumerated both the UI
  hide on Skipper and the `/api/users/:id/roles` write rejection. The first
  half ("Scoped permissions require Admiral.") is the operator-relevant
  fact; the rest reads as a fence specification, which is awkward for an
  open-core product where the gate is readable in source anyway. Trimmed
  to just the tier claim.

- `features/host-console.mdx::Availability`: the paragraph already says
  who can use the console and that the Console tab is hidden on Community
  or Skipper. The trailing "Attempting to access the console endpoint
  directly without the correct license or role is rejected" is the same
  bypass-prevention coda. Dropped.

No functional behavior change; the gates themselves are untouched.

* docs(sso): refresh SSO & LDAP authentication page against v1 template

Rewrites docs/features/sso.mdx against the v1 docs refresh template (intro
+ tier callout, sectioned Configuration anatomy, Frame screenshots,
AccordionGroup troubleshooting), bringing it in line with the previously
refreshed two-factor-authentication and rbac pages on this branch.

Recaptures all four screenshots from the production node signed in as
admin: sso-settings (overview with the five collapsible provider cards),
sso-settings-ldap (LDAP form expanded), sso-settings-oidc (Google form
expanded), sso-settings-custom-oidc (Custom OIDC form expanded with all
eleven fields).

Refreshes the Settings UI section to match the redesigned panel: each
provider is a collapsible card with an Active badge on the header, an
enable / disable toggle pill, and a footer with Save, Test Connection
(green check or red X next to the button), and Remove (only after a
config has been saved). Documents the static callback-URL helper that
sits below all five cards.

Clarifies that the per-OIDC claim mapping environment variables
(SSO_OIDC_*_ID_CLAIM, *_USERNAME_CLAIM, *_EMAIL_CLAIM) are accepted for
Google, GitHub, and Okta, not just Custom OIDC. The Settings UI hides
those fields on the presets because the defaults match.

Converts the troubleshooting section to an AccordionGroup with five
entries (Test Connection discovery failure, issuer validation error,
wrong username or missing email after sign-in, invalid redirect URI,
SSO buttons missing on the login page). Cross-links the operations
troubleshooting page for setup-time errors.

Tightens the LDAP TLS env var note to spell out the literal string
'false' requirement. Syncs the Combining SSO with 2FA section to use
the live toggle label 'Require 2FA on SSO sign-in'.

* docs(sso): drop the Community-tier Custom OIDC workaround tip

The Tip walked through how a Community-tier operator could integrate
Google, GitHub, or Okta by pointing Custom OIDC at the provider's
discovery URL, bypassing the Skipper preset gate. Operator docs should
state the tier rule once and stop; they should not describe how to
circumvent it.

The tier matrix above the removed block already names which providers
are paid; the Custom OIDC row already lists "any spec-compliant OIDC
provider" as its scope. That is enough.

* docs(vulnerability-scanning): refresh page for v1 UI and corrected tier mapping

The page was last revised before the v1 visual redesign and before the
tier-mapping changes shipped in v0.81.2 (open Community access to
secret scanning, compose misconfig scanning, scan history, and scan
comparison). This refresh:

- Rewrites the tier matrix to match the shipped Community / Skipper /
  Admiral split. Secret detection, compose misconfig scanning, scan
  history, scan comparison, and misconfig acknowledgements are now
  correctly marked as Community. Scheduled fleet scans, scan policies
  with block_on_deploy, SBOM, SARIF, and Trivy auto-update stay paid.
- Drops two stale Notes that said secret detection and compose
  misconfig scanning required Skipper or Admiral. The page now states
  each tier requirement once, in plain language.
- Refreshes all six existing screenshots from the production node:
  resources-badges, scan-details-sheet, scan-history-sheet,
  scan-compare-sheet, security-settings, app-store-toggle.
- Adds a new scan-config-button screenshot showing the stack-page
  overflow menu where Scan config now lives.
- Describes the scan drawer header accurately: Re-scan + Compare + CSV
  + SARIF as top-level buttons, with SBOM as a separate button below
  the summary.
- Updates the compose misconfig flow to point at the stack overflow
  menu (not the Deploy controls).
- Converts the troubleshooting section to a single AccordionGroup per
  the v1 template, and audits each entry for legacy phrasing and the
  removed tier claims.
- Adds a TRIVY_BIN reference to the How it works section so operators
  know about the host-binary override.

* docs(cve-suppressions): refresh page for v1 UI and corrected suppression specifics

- Recapture all three screenshots from the production node signed in
  as admin under `docs/images/cve-suppressions/` (`settings-panel`,
  `create-dialog`, `suppressed-row`). The previous file referenced
  three image paths that did not exist in the repo.
- Align prose with the actual UI labels:
  - Dialog kicker `SUPPRESSIONS . NEW`, title `New suppression`.
  - Field labels match the form: `CVE or advisory ID`, `Package
    (optional)`, `Image pattern (optional)`, `Reason`, `Expires in
    (days, optional)`.
  - Remove confirmation reads `Remove suppression` with kicker
    `SUPPRESSIONS . REMOVE . IRREVERSIBLE`.
- Factual corrections:
  - Fleet sync truncation cap is 5,000 rows (not 10,000).
  - State the admin-role requirement once in the lead Note.
  - Drop references to a `Fleet . Sync status` page and a `Reanchor`
    button; neither exists in the UI. The reanchor flow is an admin
    API call and is documented in /features/fleet-sync.
  - Sharpen the specificity scoring section (package + image scores
    3, package only 2, image only 1, neither 0) so the order matches
    the read-time filter logic.
  - Note that the image-pattern glob is case-sensitive.
- New coverage:
  - Suppressing directly from a scan result, including which fields
    are read-only in that inline flow and when to fall back to
    Settings to broaden scope.
  - The `replicated` and `expired` row badges in the panel.
  - Hovering the package column on a suppressed row to surface the
    Reason.
  - Two distinct read-only modes: viewing a remote node from the hub
    (panel hidden, banner shown) versus signing into a replica
    instance (panel visible, read-only).
  - SARIF export carries suppressions through as
    `kind: external, status: accepted`, cross-linked to the
    Vulnerability Scanning page.
- Convert troubleshooting to AccordionGroup with six entries; update
  the truncation entry to reflect the 5,000-row cap.

* docs(private-registries): refresh page for v1 UI and fleet-wide credential model

Rewrites the page against the v1 docs template (Note tier gate, opening Frame,
sectioned anatomy, AccordionGroup troubleshooting) and replaces every
screenshot with a fresh capture taken against the current product.

Corrects several factual claims that no longer matched the current code:

- Registries are stored once on the control instance and applied fleet-wide,
  not configured per node. The old Multi-node behavior section and the
  matching troubleshooting entry described a per-node model that the product
  no longer has.
- The Registries section is hidden on remote nodes (global scope) and on
  Sencho versions that do not surface the feature. New troubleshooting
  entries explain both visibility states.
- The feature is admin-only on Admiral. Non-admin operators do not see the
  section even on Admiral; previous copy implied any Admiral license user
  could manage credentials.
- Registry endpoints are not reachable from API tokens; only an admin
  browser session can manage credentials. The Security section now states
  this without naming internal route paths.

Documents UI behavior the previous page omitted: the inline form (not modal),
the four type-specific form variants, the Docker Hub read-only URL field, the
destructive delete confirmation with its stack-pull warning, the masthead
REGISTRIES count, and the empty-state callout copy.

Screenshots replaced:
- registries-overview.png: section with one configured GHCR card and the
  masthead stat at one.
- registries-empty.png: empty state with the Add registry button and callout.
- registries-add-form.png: inline form with the Docker Hub default and the
  read-only URL field.
- registries-ecr-form.png: form switched to ECR, showing the AWS Region
  field and the relabelled AWS credential inputs.
- registries-card-detail.png: card close-up with the three action icons and
  the metadata row.
- registries-delete-confirm.png: destructive ConfirmModal with the kicker,
  title, and stack-pull warning body.
- registries-with-entry.png removed (superseded by registries-overview.png
  and registries-card-detail.png).

* docs(auto-update): refresh readiness page for v1 redesign

Bring the Auto-Update Readiness doc in line with the shipped UI:

- Replace the hero screenshot with a fresh capture of the redesigned
  board (italic-display hero, brand-cyan accent, per-node groups with
  local/remote pills, dashed-border changelog separator).
- Rewrite the card-anatomy list. Drop the rollback-target bullet (the
  field exists in the backend payload but is not rendered). Add the
  "Rebuild available" inline label and the primary-image / multi-service
  count line.
- Rewrite the risk-tags table as a risk-badges table using the actual
  badge labels and colors emitted by the UI (Safe / Review / Blocked
  with the corresponding icons; Digest rebuild for non-semver tags).
- Add an Empty state section and document the per-node group header.
- Tighten the hero subtitle paragraph to match the actual UI string
  (only major-bump count is surfaced separately; preview failures are
  not).
- Fix workflow step 4: major-bump apply path is the stack lifecycle
  Update action, not the Schedules editor (a scheduled task hits the
  same block).
- Add the 2-minute manual-refresh cooldown to the Recheck workflow.
- Remove the broken cross-link to the non-existent
  /features/image-update-detection page and inline the 6-hour cadence
  fact from ImageUpdateService.INTERVAL_MS.
- Convert troubleshooting to AccordionGroup format per the troubleshoot
  ing convention used on /features/deploy-progress.
- Sync the Auto-Update entry in /features/overview.mdx to the new
  badge labels and the corrected hero-counter description.

* docs(auto-update): fix Auto-Update entry point in Workflow step 1

Workflow step 1 said "Open the Auto-Update view from the sidebar." The
Auto-Update view is opened from the top nav strip (alongside Home,
Fleet, Resources, App Store, Logs, Schedules, Console, Audit). The
sidebar carries the stack list and the per-stack right-click / kebab
context menu that toggles auto-updates on or off; it does not house
the Auto-Update top-level view.

* docs(auto-update): trim enforcement detail from per-stack control note

State the tier requirement once and stop, per Directive 27. The
"The toggle does not appear on Community" sentence enumerates the
enforcement effect of the gate, which the source already reflects;
operator docs do not need to narrate it.

* docs(auto-heal): refresh page for v1 UI and policy hardening

Rewrite Auto-Heal Policies docs against the current Stack Monitor
sheet: corrects the Max restarts / hr field label, documents the
per-policy enable toggle, the consecutive-failures pill, the full
Recent activity action set (including Docker unavailable), the 30s
evaluation cadence, multi-node behavior, notification dispatches,
and the dashboard Configuration status counter.

Replaces the broken /images/auto-heal-policies/policy-sheet.png
reference with three fresh screenshots captured against a live
node: the sheet on the Auto-heal tab, a single policy row, and
the expanded Recent activity panel.

* docs(webhooks): refresh page for v1 UI, correct tier and add Git source sync

- Fix tier note: gate is Skipper or Admiral, management is admin-only.
- Update Settings path to Settings -> Alerts -> Webhooks; document the
  read-only Node field and the green secret-reveal callout.
- Add the missing Git source sync action and the git-pull override value.
- Refresh the configured-webhooks card description: action/stack/node
  badges, On/Off toggle, copy URL, and the Recent executions disclosure.
- Tighten the trigger section with a constant-time signature check note
  and a status/body/meaning response table.
- Add an Accordion troubleshooting block covering common signature
  failures, the 404 case, no-op actions on 202, and git-pull prereqs.
- Re-capture all three screenshots from the v1 UI.

* docs(webhooks): wrap troubleshooting accordions in AccordionGroup

* docs(sidebar): refresh page for v1 redesign with filter chips, bulk mode, row anatomy, and troubleshooting

Rewrites the Stack Sidebar page against the live v1 sidebar and the v1
docs refresh template (Frame screenshots, Note tier callouts,
AccordionGroup troubleshooting). Recaptures all four existing
screenshots and adds three new captures: filter chips, row anatomy,
and bulk mode.

Adds coverage for features the previous page omitted entirely: the
ALL / UP / DOWN / UPDATES filter chips with their counts cap and
collapse toggle; bulk mode (B key, sticky toolbar with Start / Stop /
Restart, and Update gated on Skipper or Admiral); stack-row anatomy
(status pill, label dots with +N overflow, image-update dot vs Git
source icon priority, hover kebab); the Auto-update toggle, Schedule
task, and Open App entries in the context menu; the B shortcut for
bulk mode.

Corrects three claims that no longer matched the code or UI:
Auto-Heal is gated on Skipper or Admiral, not universal; the global
Ctrl+K opens the command palette, not the sidebar search box; the
activity footer kicker reads LIVE / IDLE with the verbatim copy from
SidebarActivityTicker. Documents the in-menu ↗ and L › glyphs as
visual hints rather than global keybindings to match
useStackKeyboardShortcuts.ts.

* docs(sidebar): trim enforcement-effect sentence from context-menu tier note

State the Skipper / Admiral requirement once and stop, per Directive 27.
The "They do not appear in the menu on Community" clause described the
enforcement effect alongside the gate, which the directive bans in
operator-facing docs.

* docs(host-console): refresh page for v1 UI and clarify shell metadata

Rewrite the Host Console page to match the current Cockpit layout
(masthead + terminal well + chip strip), replace the legacy PowerShell
screenshot with a fresh bash capture, and document the masthead tone
states, kicker, metadata pills, and session/heartbeat behavior. Trim
the security section to state the tier and role rule once.

* docs(licensing): refresh page for v1 UI, corrected pricing, and trial flow

Rewrites the Licensing & Billing page to match the redesigned v1
Settings layout. The previous draft still described the legacy
Settings Hub: in-app "Upgrade your plan" Skipper/Admiral cards,
"Start monthly trial" / "Start annual trial" buttons, the
"Have a license key?" field, "Manage Subscription" with a capital S,
"Deactivate License" as the button label, and the license-active.png
asset rendering the literal "Sencho Pro" string in the card title.
None of that exists in the current product.

- Refreshes the Plans table to the live pricing on sencho.io/pricing
  and adds an Enterprise mention with the floor price ($3,500/year).
  Skipper now $11.99 annual / $14.99 monthly / $449 lifetime, Admiral
  now $69.99 annual / $89.99 monthly / $2,499 lifetime.
- Rebuilds the Feature breakdown from a code-level audit of every
  requirePaid, requireAdmiral, requireScheduledTaskTier, and
  requireTierForSsoProvider call site in backend/src/routes, not
  from the marketing page. Notable code-grounded items: CVE
  suppressions on Community (no requirePaid guard), manual fleet
  snapshots on Community (scheduled snapshots on Skipper),
  Sencho Mesh under Admiral (entire mesh.ts router is requireAdmiral),
  and scheduled-task tiering names update/scan/snapshot as the
  Skipper subset with everything else under Admiral.
- Rewrites the Free trial flow end to end. The previous steps told
  operators to click in-app "Start monthly trial" or "Start annual
  trial" buttons; no such buttons exist. The new flow starts on
  sencho.io/pricing, switches to the Annual or Monthly tab, clicks
  "Start 14-day trial" on the Admiral card, completes the Lemon
  Squeezy checkout (card-required, no charge before day 14), and
  pastes the issued key into Settings -> License -> License key.
- Adds a new "The Plan section" anatomy block describing the masthead
  SCOPE / PLAN / DURATION (or RENEWS, TRIAL, STATUS) stat pills and
  the Plan card fields (Customer, Product, masked License key, status
  helper).
- Adds a new "License states" reference table covering
  Community / Trial / Active subscription / Active lifetime /
  Expired / Disabled, what each surface renders, and which of the
  Plan / Activate / Pricing sections is visible in each state.
- Corrects every UI label that drifted: section heading is Activate,
  field label is License key (not "Have a license key?"), buttons are
  Manage subscription (lowercase s) and Deactivate (not "Deactivate
  License"), and the action-row hint reads "Lemon Squeezy manages
  billing".
- Documents the redesigned profile dropdown: identity header with
  initials chip, role badge, and tier badge, then Settings,
  conditional Billing, Documentation, Feedback, an Appearance
  segmented control, and Log Out. Billing only appears when the
  license is an active non-lifetime subscription.
- Replaces all four screenshots under docs/images/licensing/:
  license-admiral-active.png (production Admiral lifetime view),
  profile-menu.png (redesigned popover), and two new captures for
  the Community-tier surfaces (license-activate-section.png,
  license-community.png). Removes the stale license-active.png
  (legacy "Sencho Pro" card) and profile-billing.png (legacy
  dropdown).

* docs(settings-reference): refresh page for v1 UI with new sections and masthead

Rewrites docs/reference/settings.mdx against the current Settings Hub so a reader
encounters an accurate map of every section. Adds the previously missing **Cloud
Backup** and **Security** sections, restructures **System Limits** into Host
thresholds and Docker hygiene subsections (GiB units, "Global crash capture"
toggle), fixes the Account password minimum to 8 chars and documents the
two-factor subsection, refreshes License/Routing/Webhooks/App Store with the
field labels actually rendered today, and documents the masthead pills
(SCOPE/NODE, EDITED, plus the per-section stats like 2FA, PLAN, CHANNELS, ROUTES,
WEBHOOKS, LABELS, TRIVY, POLICIES, PROVIDER, USED, SNAPSHOTS, DEV MODE).

Replaces five existing screenshots that predated the v1 redesign and adds five
new captures: Account with the 2FA card, License panel, System Limits with both
subsections, Security with the Trivy installer, and Cloud Backup with Sencho
Cloud Backup provisioned. All shots taken against the production node.

* docs(licensing): drop billing-provider name from operator-facing copy

The previous draft named the third-party billing provider in nine
places (checkout, receipt email, error toast verbatim, Customer /
Product field descriptions, the action-row hint, the billing portal,
and the validation API). Operator docs don't need to advertise which
vendor sits behind the checkout, billing portal, and validation
calls. Rewrite each instance to describe what the operator sees and
does without naming the upstream service.

* docs(node-compatibility): refresh page for v1 UI with lock card visuals and current capability list

- Replaces the legacy "dim + blur + pill overlay" description with the
  current CapabilityGate behavior: a centered lock card with an Unplug
  icon, title "<feature> is not available on this node", and a body line
  that names the node and its running version.
- Corrects the tier-interaction section: on the wrong tier the entry
  point is hidden entirely, so the lock card only appears for users who
  already cleared the license gate.
- Documents the public /api/meta endpoint, the 5-minute success cache,
  the 30-second failure cache, and the lazy-fetch behavior visible in
  the switcher (the version pill appears once a node has been visited).
- Refreshes the capability table against the current CapabilityRegistry
  list, adding container-exec and vulnerability-scanning, with a note
  that vulnerability-scanning is only advertised when Trivy is installed.
- Adds three production screenshots captured on the live fleet:
  switcher popover with mixed-version pills (one node on v0.76.9, rest
  on v0.81.11), a real lock card on an older pilot agent, and the
  Connection Details panel from Settings · Nodes.

* docs(security): refresh security architecture page for v1 UI

Add Fleet Secrets and Webhook signatures cards plus tier-matrix rows for
shipped-but-undocumented features. Rename SSO presets from "one-click" to
"preset providers" (presets still require OAuth-app provisioning on the
upstream IdP). Update settings paths to the v1 middle-dot convention:
Settings · Users, Settings · Account, Settings · Developer · Data retention.

Extend the encryption-at-rest list with registry credentials and Fleet
Secrets bundle payloads (both sealed with the same AES-256-GCM data key)
and clarify the password section with bcrypt cost factor 10.

Add a Webhook signature authentication subsection covering the per-webhook
HMAC-SHA256 secret, one-shot display, masked preview thereafter, and
constant-time comparison on inbound triggers.

Replace the API Tokens screenshot with a fresh capture against the v1
Settings · Identity · API Tokens panel.

* docs(security-advisories): retire reference page

The reference/security-advisories page does not survive the v1 docs
refresh:

- Misuses the term "Security Advisories", which industry-wide refers to
  published notices for confirmed product CVEs (ID, severity, affected
  versions, fix version, remediation). The retired page was a narrative
  changelog of internal hardening work between v0.19 and v0.25.2.
- The narrative is also frozen at v0.25.2 (April 2026) while current
  release is v0.81.11. Refreshing it would require backfilling ~56
  release entries' worth of hardening copy.
- The framing is uniformly "improved from prior behavior" (minimum 8
  characters up from 6, 1-year token expiry previously without expiry,
  CORS previously allowed all origins, users should upgrade promptly).
  Sencho has not shipped publicly; there are no users to address as
  upgraders.

All operationally relevant content already lives elsewhere: the
security architecture page covers the current posture, verifying-images
covers the supply-chain attestations, cve-suppressions covers operator
acknowledgment, vulnerability-scanning covers the in-app scanner, and
contact + the security architecture page both surface the disclosure
path. Published Sencho-product advisories, when any exist, will appear
on the GitHub Security tab, which is already linked from those pages.

Inbound-link audit returned a single hit on the nav entry itself; no
other doc, README, or operator artifact deep-links the slug.

* docs: rewrite Pilot Agent page with deep architecture and operations reference

Reframes docs/features/pilot-agent.mdx as the architecture-and-operations
companion to the operator walkthrough in Multi-Node Management. Adds a
mental model section, an explicit security and trust model, a full agent
env-var reference, an honest limitations list, and a 5-item FAQ. Verifies
every constant and label against the current backend source. Refreshes
four production screenshots (admin login, scrubbed) and resolves the
previously-broken /images/pilot-agent/enrollment-dialog.png reference.

Adjacent edits keep the cross-linking coherent:
- multi-node.mdx adds a one-line forward link to the rewritten page
- security.mdx adds a Pilot Agent tunnel credentials subsection

* docs(fleet-federation): deep rewrite with production screenshots

Rewrites the Fleet Federation page against the v1 docs refresh template
following the recent fleet-view, pilot-agent, and multi-node refreshes.
Doubles the page length (92 to 204 lines) while keeping the cut-line v1
MVP scope: operator-driven placement controls (cordon + pin) for
Blueprints, no expansion into mesh/sync/pilot territory.

What changed:

- Adds four production-captured screenshots under docs/images/fleet-federation/:
  the Federation tab with a cordoned node populated, the node-card kebab
  menu showing the Cordon node entry, the cordon confirmation dialog
  with a reason filled in, and a node card displaying the Cordoned pill.
- Expands the page to eleven sections: opening summary, philosophy
  (kept), key capabilities, prerequisites, step-by-step usage with
  embedded screenshots, behaviour and lifecycle table, security and
  audit, limitations and non-goals (expanded), practical workflows (new:
  OS patching, host-to-host migration, gateway pinning), troubleshooting
  (eight accordions, up from five), and a Where Federation fits
  cross-link table.
- Documents the exact production UI strings observed: the cordon
  dialog description, the uncordon confirmation copy, the reason field
  cap (256 chars), and the audit log action names (node.cordon,
  node.uncordon, blueprint.pin).
- Documents the audit visibility surface so operators know how to
  filter the Audit view for cordon and pin history.
- Adds eight cross-links to sibling pages (Fleet View, Multi-Node,
  Pilot Agent, Mesh, Fleet Actions, Fleet Sync, Blueprints, Licensing)
  with one-line scope contrasts so newcomers can place Federation in
  the broader fleet picture.
- Tightens lifecycle table to operator-relevant terms (no DB column
  names) and audit section to operator-facing wording (no middleware
  names), keeping the page operator-focused rather than
  implementation-focused.

Validation:

- Captured screenshots against the production node logged in as admin,
  using Playwright MCP. Cordoned and pinned actions reverted; audit log
  confirmed the matched cordon/uncordon pair.
- Verified every cross-link target exists in the v1-refresh worktree
  (/features/fleet-view, /features/multi-node, /features/pilot-agent,
  /features/sencho-mesh, /features/fleet-actions, /features/fleet-sync,
  /features/blueprint-model, /features/licensing).
- Compliance: no em dashes, no PII, no "previously"/"used to" framing,
  no fence-spec language, tier rule stated once in plain language.

* docs(fleet-federation): drop fence-spec phrasing in the open-core context

Sencho is open-core: anyone can clone the repo and read the tier gate.
Operator docs that name exactly where the UI gate sits ("hidden at the
Community and Skipper tiers", "lower-tier users do not see the toggle",
"only the Federation tab is gated") work as a dig-target for a
tech-savvy reader and undercut the open-core posture. Directive 27
already bans enforcement-chain spelling; the open-core threat model
makes the same phrasings risky even when they describe UI surfaces
rather than route guards.

Removes three instances of the pattern on this page:

- Top Note callout: drops "The tab is hidden at the Community and
  Skipper tiers." Keeps the one-line requirement: "Federation is an
  Admiral feature. Cordon and pin actions require an admin user role."
- Security and audit section: drops the sentence enumerating which UI
  affordances are hidden from which tiers. Keeps the customer-visible
  behavior (the Cordoned pill stays visible at every tier as a
  read-only signal).
- Troubleshooting "Federation tab is not visible" accordion: rewrites
  to lead with the requirement and the role check, drops the
  "Federation is hidden by design" and "only the toggle and the
  Federation tab are gated" phrasings.

Other claims on the page unchanged; rule is still stated once in plain
language at the top of the page.

* docs(fleet-sync): deep rewrite with production screenshots

Replace fleet-sync.mdx with a verified end-to-end reference. The previous
page named two replicated resources but the code syncs three, described a
sync-status panel and a fleet-vs-node scope picker that do not exist in
the shipped UI, and was missing prerequisites and several edge cases.

Highlights of the rewrite:

- Names all three replicated resources (scan policies, CVE suppressions,
  misconfig acknowledgements) and treats them uniformly.
- Drops the sync-status-panel and node-scope-picker UI claims; both move
  to the Limitations section as honest caveats.
- Adds prerequisites covering the paid-tier requirement on the control,
  admin-role requirement, proxy-mode remotes, and reachability.
- Expands lifecycle coverage: per-node serialised pushes, add-node
  backfill, monotonic pushedAt, per-resource watermarks, identity-drift
  notifications, the 5000-row truncation cap, stale-target warnings,
  audit-log entries on the replica.
- New "Where Fleet Sync fits" closing table cross-linking to Fleet View,
  Multi-Node Management, Pilot Agent, Vulnerability Scanning, CVE
  Suppressions, Fleet Federation, Fleet Actions, and Licensing.
- Two fresh production screenshots: control Security panel and the
  "Scanner is per-node" callout shown when proxying to a remote.

* docs(fleet-actions): deep rewrite with production screenshots

Three cards are documented end to end: Stop fleet by label, Bulk label
assign, and Prune Docker resources fleet-wide. Adds the execution-path
distinction (control-orchestrated fan-out vs single-node proxy), per-card
behaviour and partial-failure semantics, prerequisites, limitations,
practical workflows, an Accordion troubleshooting section, and a Where
Fleet Actions fits comparison table linking the surrounding Fleet view
features.

Corrects the prior page's tab-neighborhood claim, confirm-dialog wording,
autocomplete-vs-request scope, and missing batch ceiling. Replaces the
ten-day-old single screenshot with five fresh production captures under
docs/images/fleet-actions/.

* docs(fleet-secrets): deep rewrite with production screenshots

Full rewrite of /features/fleet-secrets matching the fleet-actions
structure. Replaces the sparse v1 page (no Frames, inline Q&A) with a
gold-standard layout: opening Frame, single Note for the tier gate,
'What it covers' table, mental model, prerequisites, create + edit +
versions + push (Target / Preview / Results) sections each with a
production Frame, Import from stack section, behaviour and lifecycle
table, audit-trail mapping with the six exact audit strings,
limitations and non-goals, practical workflows, AccordionGroup
troubleshooting, and a Where-it-fits cross-link table.

Adds six fresh production screenshots under
docs/images/fleet-secrets/ : overview, create, versions, target,
preview, and results.

Documents the Import-from-stack flow (depends on the bundle editor's
new Import action) and uses the post-rename 'Send' wording on the
bundle-row action (depends on the aria-label fix).

Corrects three factual drifts vs the code: env-key regex described as
'letter or underscore, then letters/digits/underscores; case-
sensitive' to match ^[A-Za-z_][A-Za-z0-9_]*$ ; documents only the
'ok' and 'failed' status pills (the 'skipped' enum value is unused);
replaces the bogus 'stack not found' troubleshooting entry with the
real 'env file not declared' cause.

Drops the fence-spec phrasing 'The tab is hidden on Community.' per
Directive 31; the tier requirement is now stated once in plain
language.

* docs(sencho-mesh): deep rewrite with mental model, lifecycle, security, screenshots

Replace the feature-reference page with a deep product + technical guide.

Adds:
- Opening hook framing audience and problem (cross-node service-to-service
  without a separate VPN or service-mesh sidecar).
- Mental model: three moving parts (sencho_mesh bridge, alias registry,
  cross-node transport) with direction-of-flow described in prose.
- Key capabilities, prerequisites, step-by-step usage with inline screenshots.
- Full lifecycle section covering opt-in, opt-out, sticky stack-stopped state,
  peer reconnect, and the proxy-mode bridge with its real default (persistent,
  env-override for idle).
- Security and trust boundaries split into authentication, inbound exposure,
  encryption, audit, and app-layer caveats.
- Limitations and non-goals: one-alias-per-port, port 1852 reserved,
  central-relay for remote-to-remote, shared 1024-stream pool with the Pilot
  tunnel, no L7, host-network unsupported, in-memory activity log.
- Three concrete workflow examples and a complete troubleshooting accordion
  (every data-plane reason, every probe stage, every unreachable cause) plus
  a Common questions FAQ.
- Where Mesh fits CardGroup linking Pilot Agent, Multi-Node, Federation,
  Licensing.

Corrections vs prior text:
- Tab is labelled Traffic in the UI (not Routing); all navigation references
  updated.
- Proxy-mode bridge default is no idle close (env-overridable to opt into idle
  teardown); prior 5-minute-teardown claim removed.
- Audit trail scope tightened: only opt-in / opt-out write durable rows;
  tunnel-state and probe events live in the in-memory activity log.

Adds seven production screenshots under docs/images/sencho-mesh covering
Table view, opt-in sheet, graph (Tunnels and Aliases edge modes), Diagnostics,
activity log, and per-stack topology.

* docs(blueprints): add missing detail-state-review screenshot

Captures the Blueprint detail sheet with a deployment row in the
"Awaiting confirmation" status (stateful first-deploy gate), to fix the
broken image referenced at blueprint-model.mdx:132. mint broken-links
now reports zero broken references.

* docs(blueprints): deep rewrite with mental model, lifecycle, security, prerequisites

Restructures the Blueprints page against the v1-refresh template used by the
recently-refreshed mesh, secrets, and atomic-deployments pages. Adds a mental
model, prerequisites table, lifecycle and status-transition map, security and
trust boundaries section, practical workflows, common questions accordion,
and a Where Blueprints fits CardGroup. Removes the internal-style rollout
and watch-plan section. Replaces all nine production screenshots with fresh
captures against the production node signed in as admin, and adds two new
captures (federation pin policy table, stateless eviction dialog). Rewrites
the tier-gate Note to drop the fence-spec phrasing that violated Directive
31. Every retained claim is anchored to current backend or frontend code.

* docs(pilot-agent): recapture enrollment dialog with compose payload

Replaces the pre-0.84 docker-run capture with the current dialog (Compose
file, two-step instructions, "Copy compose file" button) and refines the
alt text to describe the captured content. URL and token redacted to
placeholder values during capture.
2026-05-20 08:43:18 -04:00
sencho-quartermaster[bot] 6bd4645ea9 chore(main): release 0.84.1 (#1125)
Co-authored-by: sencho-quartermaster[bot] <275163604+sencho-quartermaster[bot]@users.noreply.github.com>
v0.84.1
2026-05-20 03:37:34 -04:00
Anso ee891b093b fix(self-update): resolve /app/data through bind or named volume (#1124)
SelfUpdateService scanned the container's Mounts list for Type='bind'
only when picking the host path to forward to the helper container.
A pilot enrolled with the recommended Docker Compose snippet uses a
named volume (sencho-agent-data:/app/data, Type='volume'), so the
lookup returned null and the boot log read "/app/data mount not found
- update error recovery will be unavailable". The helper still ran
on update but could not persist UPDATE_ERROR_FILE on a failed pull,
so the next gateway process had nothing to surface.

Extract findDataDirHost(mounts) and accept Type='bind' or
Type='volume'. Docker populates Source with the on-disk volume
directory for either type, so the existing :rw bind in the helper
spawn works unchanged. The independent hostBindMounts filter stays
strict to Type='bind' (forwarded operator-declared compose paths
only; named volumes are not in scope there).

Tests: pure unit coverage for the helper across 6 cases (bind hit,
volume hit, no /app/data, mixed list with siblings, empty Source,
tmpfs ignored).
2026-05-20 03:36:08 -04:00
Anso 0b50c88eb3 fix(fleet): route remote update trigger through getProxyTarget (#1123)
POST /api/fleet/nodes/:id/update and POST /api/fleet/update-all read
node.api_url + node.api_token directly, so a pilot-agent row (which
carries neither) returned "Remote node not configured." and was
filtered out of bulk update. Both routes now resolve the target via
NodeRegistry.getProxyTarget(), which returns the loopback URL for an
active pilot tunnel and the configured api_url for proxy-mode remotes.
fetchMetaForNode replaces the direct fetchRemoteMeta call for the
self-update capability check.

self-update is removed from PILOT_DISABLED_CAPABILITIES so a
Compose-deployed pilot can advertise it; the host-console filter
stays. SelfUpdateService still gates the local-end capability on the
container actually carrying docker-compose labels, so a docker-run
pilot self-disables and the Fleet UI sees a clean 503 instead of an
ambiguous failure.

Tests: new fleet-pilot-update covers the four single-node branches
(success via loopback, null target, no self-update capability, meta
offline) and two update-all cases (mixed-fleet dispatch, all-targets-
null skip). capability-registry-pilot rewired to assert the new
filter set.
2026-05-20 03:35:39 -04:00
Anso 282ab8d844 fix(pilot): let SENCHO_PUBLIC_URL override the request Host in enrollment (#1122)
The enrollment minter inferred SENCHO_PRIMARY_URL from the request Host
header, which baked loopback or LAN addresses into the compose YAML when
the admin opened Add Node on the central's own machine. Pilots on a
different network (a public cloud VPS, for example) cannot dial that.

SENCHO_PUBLIC_URL on the primary now wins when set and well-formed
(http(s)://, no loopback). Trailing slashes are stripped. Falls back to
the request Host when unset or invalid.
2026-05-20 03:35:25 -04:00
sencho-quartermaster[bot] b4d92621d4 chore(main): release 0.84.0 (#1120)
Co-authored-by: sencho-quartermaster[bot] <275163604+sencho-quartermaster[bot]@users.noreply.github.com>
v0.84.0
2026-05-20 02:06:02 -04:00
Anso 3ad6ea9c5d feat(pilot): make Docker Compose the canonical pilot enrollment payload (#1121)
* feat(pilot): make Docker Compose the canonical pilot enrollment payload

The Add Node dialog for a pilot-agent now returns a Compose snippet
instead of a single-line docker run command, and the enrollment dialog
walks the operator through a save-and-up flow. The compose project name
and container name align with what SelfUpdateService looks up at boot,
so a Compose-deployed pilot can be updated remotely through the Fleet
view without intervention on the remote host.

Docs (pilot-agent, remote-updates) were rewritten to match.

* test(e2e): align pilot enrollment spec with Compose payload

The spec was written against the docker-run payload; it now asserts the
Compose YAML the dialog renders.
2026-05-20 02:05:21 -04:00
Anso 0117556bea fix(fleet): rename Traffic tab label to Routing for consistency (#1119)
The Fleet view sub-tab that renders RoutingTab.tsx was labeled "Traffic"
while every adjacent identifier already used "Routing": the backend
route file (backend/src/routes/mesh.ts), the component path, the
localStorage key (sencho-routing-view-mode), the SegmentedControl aria
label ("Routing view mode"), and the engineering vocabulary across the
codebase. Operators looking for the "Routing tab" could not find it
because the visible label said something else.

Align the user-visible label with the rest of the implementation and
update the two doc references that named the tab by its old label
(docs/features/fleet-actions.mdx, .env.example).
2026-05-19 23:25:30 -04:00
sencho-quartermaster[bot] 9362c7215e chore(main): release 0.83.1 (#1118)
Co-authored-by: sencho-quartermaster[bot] <275163604+sencho-quartermaster[bot]@users.noreply.github.com>
v0.83.1
2026-05-19 19:50:34 -04:00
Anso e05099f2a1 fix(fleet-sync): make control-identity-mismatch sticky and surface in UI (#1117)
Treat 409 CONTROL_IDENTITY_MISMATCH from a replica as a non-retriable
failure instead of looping the same 409 through the 5-minute retry
service forever and silently writing identical failure rows.

Backend
- DatabaseService: add `sticky_error_code`, `sticky_error_expected`,
  `sticky_error_got` columns to `fleet_sync_status` via an idempotent
  migration. New methods setFleetSyncSticky, getFleetSyncStickyCode,
  clearFleetSyncStickyForNode. recordFleetSyncSuccess clears the sticky
  flag on a clean push. getFailedSyncTargets SQL adds
  `AND sticky_error_code IS NULL` so the retry loop skips sticky rows.
- FleetSyncService.executePushToNode: short-circuits at the top when
  sticky is set (covers event-driven pushResourceAsync calls). On a 409
  with code CONTROL_IDENTITY_MISMATCH, records the failure once and
  pins sticky with the expected/got fingerprints carried in the 409 body.
- routes/nodes.ts: new POST /api/nodes/:id/fleet-sync/reset-anchor.
  Admin + paid + node:manage. Proxies POST /api/fleet/role/reanchor to
  the peer with `{override:true}` using the stored Bearer node_proxy
  token. On peer 200, clears every sticky row for the node so the next
  push re-anchors and resumes replication. Distinct 502 / 504 responses
  for peer-rejected / peer-unreachable so the UI can show a useful toast.

Frontend
- New lib/fleetSyncApi.ts + hooks/useFleetSyncStatus.ts. Polling hook
  (30s visibilityInterval) skips fetch when !isPaid.
- NodeManager.tsx: destructive banner per affected node listing both
  fingerprints, with `Reset anchor on peer` and `Remove node` buttons.
  Hidden for community-tier users via empty hook data.
- FleetConfiguration.tsx (Fleet -> Status): read-only `Policy sync`
  SummaryRow per remote node card. In sync / degraded / paused with
  a tooltip; no action buttons (the action lives in NodeManager).

Tests
- fleet-sync-service.test.ts: 4 new cases for sticky-set on first
  mismatch, short-circuit on subsequent pushes, null fingerprints,
  and non-mismatch failures not setting sticky.
- database-fleet-sync-sticky.test.ts (new): 6 cases pinning the DB
  contract incl. retry-loop SQL filter and migration idempotency.
- nodes-fleet-sync-reset-anchor.test.ts (new): 6 cases covering
  happy path, peer 401 -> 502, peer unreachable -> 504, local-node
  rejection, unknown node id, and community-tier 403.

Gate parity (Directive 30): the new POST .../reset-anchor enforces
requireAdmin + requirePaid + node:manage (matches the existing read at
GET /api/fleet/sync-status). UI banner + SummaryRow only render when the
hook returns data, which it only does for paid-tier authed users. No
existing tier-gate file moved; this is greenfield parity.

Auth audit: the peer's POST /api/fleet/role/reanchor route already uses
requireAdmin, which accepts the central's stored node_proxy Bearer
token because authMiddleware maps `scope === 'node_proxy'` to
`req.user = { username: 'node-proxy', role: 'admin', userId: 0 }`.
No widening required.

Backend tsc clean. Frontend tsc -b clean. 59 fleet-sync tests pass; full
backend suite green minus the pre-existing Windows-only file-lock flake
on filesystem-backup.test.ts that reproduces unchanged on main.
2026-05-19 19:49:16 -04:00
sencho-quartermaster[bot] 69bc955c3b chore(main): release 0.83.0 (#1114)
Co-authored-by: sencho-quartermaster[bot] <275163604+sencho-quartermaster[bot]@users.noreply.github.com>
v0.83.0
2026-05-19 19:13:57 -04:00
Anso 37413c1020 feat(cloud-backup): paginate cloud snapshots list (#1110)
Splits the Cloud Snapshots panel in Settings > Cloud Backup into pages of
10 items with prev/next chevrons and a page counter, matching the existing
pattern shipped in Fleet Snapshots. Long-running deployments with many
snapshots no longer overflow the settings panel.

The empty state, refresh, download, and delete flows are unchanged. The
safePage clamp handles the last-page-delete case without extra reset
logic. Chevron buttons carry aria-label values for screen reader users.
2026-05-19 18:19:57 -04:00
Anso 96d9bb9f76 fix(fleet-secrets): align bundle-row action aria-label with icon ("Send") (#1116)
The bundle-row action uses the lucide Send icon but the button's
aria-label was "Push", which surfaced inconsistently to screen-reader
users vs the visible/iconic intent. The same action is referred to
as "Send" in the feature docs, so standardise on Send everywhere.
2026-05-19 18:19:44 -04:00
Anso e3e5943b57 feat(fleet-secrets): add Import from stack action to bundle editor (#1115)
Adds an in-sheet "Import from stack" affordance to the Secret bundle
editor (edit mode only). Picks a node + stack + env file basename,
calls the existing /secrets/:id/import-from-stack endpoint, and
overlays the returned key/value pairs onto the editor: existing keys
have their values updated in place (preserving row order and any
duplicate-keyed rows for save-time validation), new keys append at
the bottom, empty placeholder rows are kept.

The endpoint and audit log row already shipped; only the UI trigger
was missing.
2026-05-19 18:19:28 -04:00
Anso 6722335a79 fix(stack-update): refresh frontend state automatically after a stack update (#1113)
After applying a stack update the sidebar's "update available" dot stayed
visible and the stack's status indicator was stuck on the optimistic value
until the page was manually refreshed. Two root causes:

1. Image-updates state refresh was a fire-and-forget call in some paths and
   entirely missing from the bulk-update, auto-update, and state-invalidate
   WebSocket-handler paths.
2. stackActionsRef.current was resynced only at render time, so the post-
   update refreshStacks(true) running in the action's finally block read a
   stale "busy" map and preserved the optimistic mask via prev[file] ?? status.

Backend now broadcasts a state-invalidate event with scope='image-updates'
and action='stack-updated' after every successful update (single-stack route
and auto-update loop). The frontend useNotifications hook routes this to a
new onImageUpdatesChange callback wired to fetchImageUpdates in EditorLayout,
so every connected client refreshes the dot through the same code path.

Bulk update also calls fetchImageUpdates directly for fast local feedback,
and setStackAction/clearStackAction now keep stackActionsRef synchronously
in sync with state so the busy-stack check inside refreshStacks observes
the cleared map immediately.

Adds 3 unit tests covering the new WS branch (positive, scope-mismatch
negative, auto-update-settings-changed negative).
2026-05-19 17:56:41 -04:00
sencho-quartermaster[bot] 7f81f06bbd chore(main): release 0.82.1 (#1112)
Co-authored-by: sencho-quartermaster[bot] <275163604+sencho-quartermaster[bot]@users.noreply.github.com>
2026-05-19 07:29:56 -04:00
Anso 9dbce9c3c7 fix(spawn): attribute ENOMEM and ENOENT-under-memory-pressure spawn failures to host OOM (#1111)
Operators previously saw "spawn docker ENOENT" or "spawn /bin/sh ENOENT" when
the host was under memory pressure, which sent them down a missing-binary
debugging path. Linux libuv's posix_spawn can fail to allocate its argv /
path-search arena under low free memory and surface the underlying ENOMEM
as ENOENT.

Centralizes spawn-error mapping in a new utils/spawnErrors.ts helper:
- Explicit ENOMEM is rewritten to "Out of memory while launching <command>
  (host free memory: X MiB of Y MiB)".
- ENOENT under the 128 MiB free-memory floor is rewritten with the same
  wording plus a "reported as ENOENT under memory pressure" hint.
- ENOENT for docker on a healthy host preserves the existing
  "Docker CLI unavailable on this node" mapping.
- Other errors pass through unchanged.

Applied at the four named offenders: ComposeService.execute(),
ComposeService.captureCompose(), DockerController.getContainersByStack(),
and FileSystemService.getStacks() (which gets an ENOMEM-aware log line
for the scandir failure).

Startup also logs host free/total MiB once and warns when free memory is
below the 128 MiB floor, so the diagnostic surfaces before the first
spawn attempt rather than after it fails.

37 tests cover the mapping function directly and the ComposeService /
FileSystemService integration paths.
2026-05-19 07:27:12 -04:00
sencho-quartermaster[bot] 5a2aed22fd chore(main): release 0.82.0 (#1109)
Co-authored-by: sencho-quartermaster[bot] <275163604+sencho-quartermaster[bot]@users.noreply.github.com>
2026-05-19 00:14:46 -04:00
Anso 523ba5854c fix(stacks): return 404 for nonexistent stacks on deploy/down/update (F-7) (#1108)
POST /api/stacks/:name/{deploy,down,update} previously returned HTTP 500
with body {"error":"spawn docker ENOENT"} when invoked against a stack
whose compose directory was missing. The status code was wrong (the
named resource did not exist, so 404 is the right answer) and the
message misled operators into thinking the docker CLI was unavailable.

Add a small requireStackExists(nodeId, stackName, res) helper in
routes/stacks.ts that validates the stack name and confirms a compose
file is present via FileSystemService.hasComposeFile before any of the
three handlers spawn docker compose. The helper is called immediately
after requirePermission and before runPolicyGate so unauthorized
callers still get 403 first and the policy gate never runs against a
phantom stack.

In ComposeService.execute(), narrow the child.on('error') handler so
the genuine docker-binary-missing case (ENOENT on the spawn itself)
rejects with "Docker CLI unavailable on this node" instead of the raw
"spawn docker ENOENT". This is defense in depth for the rare case the
pre-check cannot cover, and it fixes the misleading-message half of
the bug as well.

Cover the new contract with stack-actions-missing-stack.test.ts (four
cases: deploy/down/update return 404, invalid name returns 400). Mock
ComposeService as a tripwire so a future code path that bypasses the
guard would fail loudly. Fix stacks-failure-notifications.test.ts by
adding hasComposeFile to its FileSystemService partial mock so the
existing happy-path-error-handling cases continue to flow into
ComposeService.
2026-05-19 00:13:57 -04:00
Anso 3a839b781b fix(stack-editor): reset tab to compose.yaml when clicking edit (#1107)
The Edit affordance on the stack anatomy panel previously only flipped
the editor visibility flag and left activeTab whatever it was. After a
user clicked Files (which set activeTab to 'files'), closed the editor,
then clicked Edit, the editor reopened still on the Files tab instead of
showing the compose.yaml editor.

Make onEditCompose mirror the sibling onOpenFiles handler by also
setting activeTab to 'compose', so the Edit button always lands on the
compose editor regardless of which tab was last viewed.
2026-05-19 00:13:44 -04:00
Anso 1f673073ca feat(fleet): add fleet-wide Docker prune to Fleet Actions (#1104)
Adds a third card to the Fleet Actions tab that fans out Docker prune
(images, volumes, networks) across every node in one submit. Local nodes
call DockerController under a bulk-prune lock; remote nodes receive one
POST /api/system/prune/system per target. Per-node + per-target results
with reclaimed bytes are surfaced inline via ResultsList.

Tier: Skipper / Admiral (requirePaid + requireAdmin), matching the rest
of Fleet Actions. The frontend card is mounted inside the existing
isPaid branch at FleetActionsTab; no new frontend gate is required.

The card uses an amber accent rail and the Eraser icon so it reads as
'cleanup' rather than 'destructive stop'. Scope toggle defaults to
Managed only (Sencho-tagged resources) with an All unused option that
escalates the destructive-confirm copy.
2026-05-19 00:13:32 -04:00
sencho-quartermaster[bot] 9f22f73cfb chore(main): release 0.81.15 (#1106)
Co-authored-by: sencho-quartermaster[bot] <275163604+sencho-quartermaster[bot]@users.noreply.github.com>
2026-05-18 22:10:34 -04:00
Anso 0cffd60173 fix(mesh): clean up activeStreams on src socket close/error (F-10) (#1105)
The cross-node MeshService.openCrossNode handler attached src.on('close')
and src.on('error') listeners that only called tcpStream.destroy() and
never deleted the activeStreams Map entry. MeshTcpStreamLike.destroy()
sends a tcp_close frame to the remote but does not synchronously emit
the local close event; that fires only when the remote sends back a
tcp_close ack via _dispatchClose (or when the entire tunnel tears down
and the bridge force-emits close on every stream).

Failed dials whose remote ack was lost (peer gone, network drop, or the
F-9-era timeouts) therefore leaked records into activeStreams until the
tunnel itself idle-closed, producing the monotonically-climbing
activeStreamCount documented in the mesh E2E audit (F-10).

Add a cleanupRecord() closure that clears the open timer and deletes
the record idempotently, then route all four lifecycle handlers
(tcpStream.on('error'), tcpStream.on('close'), src.on('close'),
src.on('error')) through it. Map.delete is naturally idempotent so the
double-delete that fires when both sides close normally is a free
no-op. The same-node openSameNode path already had this shape via its
existing teardown() closure; this brings the cross-node path to parity.

Three new Vitest cases in mesh-service.test.ts use an EventEmitter-
backed fake src so emit('close') actually delivers to the listener
(the existing makeFakeSocket uses vi.fn for .on which records calls
but never fires them). They cover src.close, src.error, and the
idempotency case where both src and tcpStream emit close in sequence.
2026-05-18 22:08:06 -04:00
sencho-quartermaster[bot] 1b33f2d064 chore(main): release 0.81.14 (#1103)
Co-authored-by: sencho-quartermaster[bot] <275163604+sencho-quartermaster[bot]@users.noreply.github.com>
2026-05-18 17:39:24 -04:00
Anso 596cce3507 fix(security): suppress CVE-2026-41567 and CVE-2026-42306 against vendored docker-compose moby library (VEX) (#1102)
The v0.81.13 Docker publish workflow failed on the post-build Trivy
re-scan because the moby/Docker vulnerability database picked up two new
HIGH CVEs against `github.com/docker/docker v28.5.2+incompatible`, which
docker-compose v5.1.3 vendors statically as a client-side API/codec
library. The CVEs were not present in the PR-side Trivy DB three hours
earlier, so PR CI passed but the release re-scan failed.

Both CVEs are daemon-side: CVE-2026-41567 targets the daemon's
PUT /containers/{id}/archive handler, CVE-2026-42306 is a race in the
daemon-side bind-mount resolution that backs `docker cp`. docker-compose
never acts as the daemon, never serves these routes, and Sencho only
invokes compose for up/down/ps. The vulnerable code paths are
unreachable. The fix lives on the github.com/moby/moby/v2 module path;
until upstream compose migrates, v28.5.2+incompatible remains the only
Go-module-resolvable version compose can reference. Same triage shape as
the existing CVE-2026-34040 entry, so the new statements mirror its top-
level products purl form.

Bumps OpenVEX document version 4 -> 5 and updates last_updated /
timestamp to 2026-05-18.
2026-05-18 17:36:52 -04:00
sencho-quartermaster[bot] 7dff6127e7 chore(main): release 0.81.13 (#1101)
Co-authored-by: sencho-quartermaster[bot] <275163604+sencho-quartermaster[bot]@users.noreply.github.com>
2026-05-18 16:51:05 -04:00
Anso c460bb87a8 fix(mesh): probe upstream synchronously in route diagnostic (F-11) (#1100)
GET /api/mesh/aliases/:alias/diagnostic returned a cached state derived
from the last latency/error maps. Those maps only mutated when someone
called POST .../test or when a cross-node connect logged an event, so
once an upstream stopped the diagnostic kept reporting "healthy" until
the 60 s alias-cache refresh pruned the alias entirely.

The GET now calls testUpstream synchronously after the alias-resolved,
opted-in, tunnel-up short-circuits. Probe failures land in routeErrorMap
via logActivity (cross-node already did this; same-node timeout/error
paths now log probe.fail in the same shape), so the state computation
flips to "unreachable" on the same request that exposed the stopped
upstream. A new routeProbeAtMap stamps freshness and surfaces in the
response as lastProbeAt; the route detail sheet renders "Last probe
<age> · <ms>" using the existing formatTimeAgo helper.
2026-05-18 16:48:09 -04:00
sencho-quartermaster[bot] 7d2b8bee7a chore(main): release 0.81.12 (#1099)
Co-authored-by: sencho-quartermaster[bot] <275163604+sencho-quartermaster[bot]@users.noreply.github.com>
2026-05-18 01:53:10 -04:00
Anso 554f662563 fix(mesh): surface stopped-stack opt-ins on routing node cards (#1098)
Add `currentlyResolvable: boolean` per entry in `MeshNodeStatus.optedInStacks`,
derived from the existing alias cache so the new field stays consistent with
`/api/mesh/aliases` without any extra Dockerode or cross-node inspect on the
status path. The Routing tab renders an amber `suspended` pill for entries
whose stack is opted in but currently has no running services, plus a single
explanatory caption below the suspended list.

Resolves the contradictory `Mesh stacks: 1 / Aliases: 0 / No mesh services
on this node yet` copy on the node card when a meshed stack's container has
been stopped; the misleading line is now only shown when the node truly has
no opt-ins. The opt-in itself remains sticky: when the stack starts again,
its aliases reappear automatically on the next refresh.

Defensive de-dup in the UI filters suspended entries against the live alias
snapshot to handle the transient gap where `/mesh/status` and `/mesh/aliases`
return slightly inconsistent views from their separate fetches.

Tests:
- new `mesh-status-resolvability.test.ts` (6 cases) locks the resolvable /
  suspended / mixed / empty / per-node-scoping / stale-alias-no-phantom
  invariants for `getStatus`.
- `mesh-topology-layout.test.ts` gains a `stacksKey` resolvability-flip case
  and a `meshNodeStateEqual` case asserting a resolvability flip on an
  otherwise identical stack registers as a state change so the topology
  layout re-runs.

Operator docs gain one new troubleshooting accordion in
`/docs/features/sencho-mesh.mdx` explaining the suspended state.
2026-05-18 01:52:18 -04:00
sencho-quartermaster[bot] 77782ce1ff chore(main): release 0.81.11 (#1097)
Co-authored-by: sencho-quartermaster[bot] <275163604+sencho-quartermaster[bot]@users.noreply.github.com>
2026-05-18 00:16:31 -04:00
Anso 9f238e187c fix(mesh): cascade opt-out when a stack is deleted (F-1 / F-14) (#1096)
DELETE /api/stacks/:name now calls MeshService.optOutStack after the
DB cleanup so the mesh_stacks row, the override file under
<DATA_DIR>/mesh/overrides/<nodeId>/, and any derived aliases do not
outlive the deleted stack. Pre-fix, those artifacts leaked and the
reconcile loop logged "No compose file found for stack" every tick.

optOutStack is idempotent (early return when the stack was never
opted in) and already cascades override-regen plus recompose across
the rest of the fleet, so peers' /etc/hosts drop the dropped alias.
The new cascade call is best-effort relative to the delete itself:
mesh cleanup failures warn but never regress the delete contract.

New regression test asserts three contracts: cascade on delete,
no-op on never-meshed, and 200 + warn when the cascade rejects.
2026-05-18 00:14:10 -04:00
sencho-quartermaster[bot] d9b4f71e0d chore(main): release 0.81.10 (#1095)
Co-authored-by: sencho-quartermaster[bot] <275163604+sencho-quartermaster[bot]@users.noreply.github.com>
2026-05-17 22:05:24 -04:00
Anso f6e42535c8 fix(mesh): route peer→central traffic over the existing forward WS (#1094)
* fix(mesh): route peer→central traffic over the existing forward WS

The reverse mesh callback path (`/api/mesh/proxy-tunnel-from-peer`) needed
SENCHO_PRIMARY_URL on central plus a publicly reachable origin from the
peer's perspective. In a typical homelab where central sits behind NAT,
peer→central dispatch silently failed at the dialer's short-circuit and
the headline "call any service on any node by hostname" worked one way
only.

The forward WS at `/api/mesh/proxy-tunnel` is already bidirectional end
to end. Make the bridge a persistent control-plane primitive: dial every
mesh-enabled proxy peer at startup, reconcile every 60 s, never idle-close.
Peer→central traffic multiplexes over the same WS via `tcp_open_reverse`.

Removed:
- `meshProxyTunnelFromPeer.ts` WS handler and dispatch
- `MeshCentralRegistry`, `PeerToCentralMeshSessionDialer`
- `mesh_handshake` first-frame state machine in `meshProxyTunnel.ts`
- `maybeSendBootstrap`, `buildHandshakeFrame` in the dialer
- `mesh_proxy_callback_bootstrap` capability and `maybeWarnUnsetPrimaryUrl`
- `mesh_centrals` table (drop migration; greenfield, no users)
- `PilotTunnelManager.replaceOrRegisterProxyBridge` (dead after handler removal)
- twelve associated unit/integration tests plus the peer-recovery branch
  in `MeshService.openCrossNode`

Added:
- `MeshService.proactiveBridgeFanout` selects every mesh-enabled proxy
  peer (no longer gated on `mesh_stacks` rows)
- `startBridgeReconcileLoop` runs the fanout every 60 s (override via
  `SENCHO_MESH_RECONCILE_INTERVAL_MS`)
- `MeshProxyTunnelDialer` default idle TTL is now `0` and exposes
  `isDialing(nodeId)` for the status surface
- `MeshNodeStatus.reverseCallbackStatus` discriminator
  (`connected | connecting | unavailable | not_applicable`) surfaced via
  `/api/mesh/status` and rendered as a pill in the Routing tab
- `openCrossNode` error message distinguishes "no proxy target" from
  "waiting for central to dial the reverse bridge"
- New tests: `mesh-service-proxy-tunnel-reconcile`,
  `mesh-status-reverse-callback`, `mesh-proxy-tunnel-dialer-no-idle-close`

SENCHO_PRIMARY_URL is no longer required for any mesh function.

* fix(mesh): rewrite proxy-tunnel reconcile test contents

The previous commit renamed the file but the rewritten test bodies stayed
unstaged on top of the rename. This commit lands the actual rewrite: the
fanout assertion now requires every mesh-enabled proxy peer to be dialed,
not just those with `mesh_stacks` rows, and adds a reconcile-tick
repeated-call test.
2026-05-17 22:00:31 -04:00
sencho-quartermaster[bot] 54c07d4930 chore(main): release 0.81.9 (#1093)
Co-authored-by: sencho-quartermaster[bot] <275163604+sencho-quartermaster[bot]@users.noreply.github.com>
2026-05-17 20:15:09 -04:00
Anso ed49ed6165 fix(http): strip zstd from Accept-Encoding so compression sets Content-Encoding (#1092)
compression@1.8.1's negotiator silently fails to set the Content-Encoding
response header when Accept-Encoding carries an unknown token like zstd,
even though the body is still compressed with brotli. Chromium 123+ sends
"Accept-Encoding: zstd, gzip, br" by default, including Playwright's
bundled Chromium over plain HTTP, so any homelab Sencho viewed without a
TLS terminator triggered ERR_CONTENT_DECODING_FAILED in the browser and
rendered as a blank page.

Insert a tiny middleware (step 4 of the canonical pipeline) that drops
zstd tokens from Accept-Encoding before compression's negotiator runs.
The negotiator then picks br or gzip and writes the matching header. The
fix is no-op when zstd is absent. If only zstd was offered, the header
falls back to identity so the response is served uncompressed.

Renumber the canonical-order JSDoc in app.ts from 17 to 18 steps and
update the step-number references in hub-only-guard and proxy-mount-order
test docstrings to match.

11 new tests (8 unit + 3 supertest integration against real compression
with a 3000-byte body) lock in the behavior: zstd-stripping across casing
and q-values, surviving tokens preserved verbatim, identity fallback when
only zstd was offered, Content-Encoding header always written when at
least one supported codec remains.
2026-05-17 20:11:44 -04:00
sencho-quartermaster[bot] 77ccce41fe chore(main): release 0.81.8 (#1091)
Co-authored-by: sencho-quartermaster[bot] <275163604+sencho-quartermaster[bot]@users.noreply.github.com>
2026-05-17 19:08:19 -04:00
Anso 1d7418a233 fix(mesh): recompose previously-meshed containers when alias set changes (#1090)
When a stack is opted into the mesh, every previously-meshed stack's
override.yml on disk gains the new alias, but the running containers
keep their stale /etc/hosts because extra_hosts is read at container
creation time. Cross-stack DNS for the new alias silently fails until
an operator manually redeploys each prior stack.

Complete the cascade: after regenerateOverridesAcrossFleet finishes,
fan out triggerRedeploy across every (node_id, stack_name) tuple in
mesh_stacks, skipping the just-opted-in pair (the explicit
triggerRedeploy at the end of optInStack covers it). Mirror the same
cascade in optOutStack so the dropped alias exits every container's
/etc/hosts.

triggerRedeploy already dispatches local vs remote, already wraps
runRedeploy in fire-and-forget error logging to the mesh activity
ring and audit log, and already enforces compose policy gates. The
cascade just reuses it.

regenerateAllOverrides (boot and manual /regen-overrides) is
intentionally NOT covered: a Sencho restart must not force a
fleet-wide recompose; override files alone are sufficient there.

Adds four unit tests in mesh-service.test.ts: cascade fires for every
prior meshed stack on opt-in, no double-redeploy of the just-opted-in
tuple, cascade is a no-op when no peers exist, cascade fires for
every survivor on opt-out.
2026-05-17 19:02:02 -04:00
sencho-quartermaster[bot] c2a23378aa chore(main): release 0.81.7 (#1089)
Co-authored-by: sencho-quartermaster[bot] <275163604+sencho-quartermaster[bot]@users.noreply.github.com>
2026-05-17 17:32:31 -04:00
Anso 578cac89da fix(mesh): surface data-plane failures in health, meta, and Routing tab (#1088)
The three previously-silent console.warn paths in MeshService.setupMeshNetwork
now route through a typed recordSetupFailure helper that classifies the
failure (subnet_invalid, subnet_overlap, subnet_mismatch, ip_in_use,
attach_failed, not_in_docker), emits a mesh.disable activity entry at the
matching level (error for real failures, warn for the expected dev-mode
not_in_docker case), and strips mesh_proxy_callback_bootstrap from
advertised capabilities via CapabilityRegistry.

/api/health gains a mesh.dataPlane block carrying the typed status.
/api/mesh/status carries localDataPlane at the top level so the Routing tab
renders a red banner with an operator-actionable recovery hint (set
SENCHO_MESH_SUBNET to a free /24 and recreate the container) when the data
plane is down. The success path re-enables the capability and flips the
status to ok.

Generalizes the previously-documented F-0 failure mode (IP-in-subnet
collision) to also cover the subnet-pool-overlap case where another Docker
bridge on the host already owns the requested CIDR.
2026-05-17 17:31:57 -04:00
sencho-quartermaster[bot] f11e5ef58c chore(main): release 0.81.6 (#1087)
Co-authored-by: sencho-quartermaster[bot] <275163604+sencho-quartermaster[bot]@users.noreply.github.com>
2026-05-17 14:12:13 -04:00
Anso a318e6b3c1 fix(mesh): close data-plane race that dropped early TcpData frames (#1086)
The tcp_open and tcp_open_reverse receivers registered their stream
entry only after awaiting resolveTarget / resolveContainerIp. The
peer, which is free to send TcpData immediately after the open frame,
hit the lookup-miss path in handleBinaryFrame and the bytes were
silently dropped. Probes passed because they only exercise the
handshake; real HTTP hung with 0 bytes received.

Reserve the stream entry synchronously, buffer early TcpData in a
per-stream pendingData queue capped at 1 MiB, and flush onto the
local socket inside the connect handler before sending tcp_open_ack.
The payload is copied because decodeBinaryFrame returns a subarray
view of the WS receive buffer; holding the view would pin the parent
buffer past its lifecycle.

Both sides of the data plane are patched:
- tcpStreamSwitchboard.onTcpOpen (forward, agent + proxy-mode peer)
- PilotTunnelBridge.handleTcpOpenReverse / acceptReverseLocal
  (reverse, primary acting as the local dial target)

Adds unit coverage for the race window and for the overflow path.
The cap constant lives in pilot/protocol.ts alongside the other
per-stream limits.
2026-05-17 14:01:39 -04:00
Anso 50b89db3b8 chore(mesh): drop mesh_stacks table on pilot-mode DBs (#1085)
Per the C-3 design, mesh state lives on central. Pilots learn aliases
via the D-1 override push and hold them in `pilotAliasOverlay`. The
local `mesh_stacks` table on pilots has been dead since C-3 shipped.

This is a greenfield cleanup (Directive 20):

- `migrateMeshTables()` runs `DROP TABLE IF EXISTS mesh_stacks` when
  `SENCHO_MODE === 'pilot'` and skips the CREATE. Central-mode
  behavior is unchanged.
- The four mesh CRUD methods (`listMeshStacks`, `isMeshStackEnabled`,
  `insertMeshStack`, `deleteMeshStack`) short-circuit on the same
  check. Reads early-return empty/false; writes early-return;
  `insertMeshStack` also warns so an accidental pilot-side write is
  loud.
- New `isPilotMode()` helper mirrors the existing one in
  `bootstrap/startup.ts`. Layering rules forbid the shared import; a
  third call-site would justify extraction to `helpers/`.

Operator-visible behavior is unchanged on either side. On central, the
table and its rows are untouched. On pilots, the table is removed (it
held no rows in steady state) and any caller of the CRUD methods
continues to see empty list / false / no-op without touching SQLite.

Test plan

- `npx tsc --noEmit`: clean.
- New `database-service-pilot-mesh-stacks.test.ts`: 5/5 green (table
  absence + four short-circuit assertions).
- `mesh-service.test.ts`, `mesh-diagnostic-local.test.ts`, and
  `mesh-service-proactive-bootstrap-fanout.test.ts`: 58/58 green.
- Full backend suite: 2280/2281; the single failure is a pre-existing
  Windows EBUSY flake in `filesystem-backup.test.ts` that reproduces
  on a clean tree.
- Code reviewed; findings applied.
2026-05-17 05:25:00 -04:00
Anso 3f620591c8 chore(ci): retire release-blog-scaffold workflow (#1084)
Removes the auto-publish workflow that fired on every v* tag and
generated a release post in the sencho-website repo on every 5th
release. Release posts are no longer the desired blog content for
the project, so the workflow, its helper script, and its template
are all removed together.

Deletes:
- .github/workflows/release-blog-scaffold.yml
- .github/scripts/scaffold-release-post.mjs
- .github/release-blog-template.tsx.tmpl

The 7 historical release posts are removed from the sencho-website
repo in a separate PR.
2026-05-17 05:24:40 -04:00
Anso 5bc4add099 refactor(mesh): retarget proxyFetch null-target throw to no_target (#1083)
`MeshService.proxyFetch` previously threw `MeshError('push_failed', ...)`
when `NodeRegistry.getProxyTarget(nodeId)` returned null, but semantically
no push was attempted: the target simply did not exist. Callers had to
special-case the code and explain the overload in a comment.

Retarget the throw to the existing `'no_target'` variant in
`MeshErrorCode`. `listStacksOnNode` matches the new code directly; the
explanatory comment is gone since the code reads for itself.

`MeshError('push_failed', ...)` continues to signal real outbound push
failures: `applyLocalOverride` (mesh data plane not yet ready on the
receiving node) and `pushOverrideToNode` (HTTP 404 from an outdated
remote, non-OK HTTP response). Those sites are unchanged.

Operator-visible behavior is unchanged in the steady state. The rare
"remote disappeared between opt-in start and override push" case now
returns HTTP 400 with `code: 'no_target'` instead of 503; the response
body carries a clear message ("no proxy target for node N") so clients
can disambiguate.

Test plan
- `npx tsc --noEmit`: clean.
- `npx vitest run src/__tests__/mesh-list-stacks-remote.test.ts`: 8/8 green
  (6 existing + 2 new cases pinning the catch-arm semantics).
- `npx vitest run src/__tests__/mesh-service.test.ts src/__tests__/mesh-inspect-remote.test.ts`:
  56/56 green.
- Code reviewed; findings applied.

Notes
- PR #1082 (refactor: route inspectStackServices through proxyFetch) needs
  a rebase + one-line update to its catch arm before merging on top of
  this change.
2026-05-17 05:06:58 -04:00
Anso dcc43205e0 refactor(mesh): route inspectStackServices through proxyFetch (#1082)
inspectStackServices was constructing its Authorization, x-sencho-tier,
and x-sencho-variant headers inline. proxyFetch is the centralized
helper used by listStacksOnNode, pushOverrideToNode, triggerRemeshRedeploy,
and removeOverrideFromNode. Equivalent today for a GET-no-body call, but
any future header added to proxyFetch (audit context, license fingerprint,
request id) would silently miss inspectStackServices.

Wire output is byte-identical: same URL, same method, same three headers,
same 5 s timeout. The new catch arm converts MeshError('push_failed') back
to the existing warn + return [] contract callers (optInStack,
refreshAliasCache) rely on. Pattern mirrors listStacksOnNode.
2026-05-17 05:06:45 -04:00
sencho-quartermaster[bot] 662036bcb7 chore(main): release 0.81.5 (#1080)
Co-authored-by: sencho-quartermaster[bot] <275163604+sencho-quartermaster[bot]@users.noreply.github.com>
2026-05-17 04:16:49 -04:00
Anso 24155cad66 fix(stacks): prevent right-side clipping in Create Stack modal (#1081)
The "From Docker Run" tab's converted compose.yaml preview uses
<pre whitespace-pre> inside a ScrollArea that defaulted to
overflow-x: hidden. Long YAML lines (long env values, long volume
paths) were silently clipped on the right with no horizontal scroll
affordance, and Radix's display: table viewport child propagated the
overflow back up the chain, pushing the form's ModalBody past the
dialog's 576px max-width.

Opt into the ScrollArea component's existing `block` prop on the
YAML preview and on both tabpanel form ScrollAreas. `block` switches
the viewport child to display: block; min-w-0 and renders a
horizontal ScrollBar (Radix mounts it only when content actually
overflows). DOM trace confirms the outer `block` is load-bearing
when the inner pre is wider than the viewport.

Matches the established pattern used by VulnerabilityScanSheet,
ScanComparisonSheet, and SecurityHistoryView.
2026-05-17 04:13:02 -04:00
Anso a32183d198 fix(mesh): cascade override regen across all meshed nodes on opt-in/opt-out (#1079)
Opting a stack into mesh on node A now propagates the new alias to every
meshed node's override file, not just stacks on node A. Same on opt-out.

Pre-fix, optInStack and optOutStack called regenerateOverridesForNode
which only walks db.listMeshStacks(nodeId). Other meshed stacks on other
nodes did not learn about the new alias until something else re-pushed
their overrides; operators worked around it with a manual
POST /api/mesh/regen-overrides plus a redeploy of each affected stack.

New regenerateOverridesAcrossFleet helper iterates db.listMeshStacks()
(no arg, fleet-wide) under Promise.allSettled, skipping the
(nodeId, stackName) tuple opt-in just pushed loudly. Per-stack failures
surface as forwarder.error activity events, matching the pattern used by
regenerateAllOverrides. Offline remote nodes leave stale overrides until
the next opt-in/opt-out, the next tunnel reconnect, or a manual rerun of
the regen-overrides endpoint.

regenerateOverridesForNode stays for the tunnel-up retry listener: a
reconnecting pilot only needs its own stacks re-pushed, not the fleet.

Tests: redirect five existing spies in mesh-service.test.ts to the new
method; add two new cases asserting cross-node cascade for opt-in and
opt-out. 52 tests in mesh-service.test.ts pass.
2026-05-17 03:48:58 -04:00