Files
sencho/docs/features/fleet-view.mdx
T
Anso 275c654407 feat(fleet): detect and update from a new sencho-dev:dev build (#1871)
* feat(fleet): add self dev-build detection primitives

Split compareLocalToRemoteTag into compareLocalToRemoteTagDetailed (returns
the probe's primary digest alongside the match/update/error verdict) with
compareLocalToRemoteTag now a thin wrapper, so a caller that needs both the
verdict and the digest no longer has to probe the same mutable tag twice.

Add detectSelfDevBuildUpdate, which compares the running container's own
image against the rolling ghcr.io/studio-saelix/sencho-dev:dev tag using the
new detailed comparison, laying the groundwork for surfacing dev-build
updates in Fleet.

* feat: add isSenchoDevRepository and isSenchoDevFloatingTag predicates

Add two pure predicate functions to helpers/selfUpdateCompose.ts for
identifying Sencho dev repository references and floating tag variants:

- isSenchoDevRepository: checks if a reference is to the ghcr.io/studio-saelix/sencho-dev
  repository, including digest-pinned and dev-<sha> tag variants
- isSenchoDevFloatingTag: checks if a reference is specifically the floating :dev tag
  on the Sencho dev repository (not digest-pinned, not immutable dev-<sha>)

Both functions reuse existing parsing patterns (normalizeImageRepository for repository
extraction, classifyImagePin idiom for digest and tag detection) to maintain consistency.

Add comprehensive test coverage in self-update-compose.test.ts covering all specified
test cases including edge cases (malformed refs, unrelated repos, digest pins, etc.).

* feat(gitops): wire dev-build detection into MonitorService

Adds a dev_build_update_available notification category and a new
checkSenchoDevBuild() cycle in MonitorService that detects when the
running container has fallen behind the rolling
ghcr.io/studio-saelix/sencho-dev:dev build it is pinned to, using
detectSelfDevBuildUpdate() and isSenchoDevFloatingTag(). Availability
state is written unconditionally so the Fleet update affordance never
depends on notification delivery succeeding, while a separate dedup
key prevents re-notifying for a digest already announced. Also guards
checkSenchoVersion() so a dev-repo pin no longer produces a false
positive stable-release update notification.

* feat(fleet): surface dev-image status and build availability

Fleet's GET /update-status now reports isDevImage (any reference to the
sencho-dev repository, including digest pins) and devBuildUpdateAvailable
(the exact floating :dev tag with a newer build observed, read from the
system-state key MonitorService already maintains). A dev-pinned local
node forces updateAvailable to false and clears any stale stable-release
skip, since that skip was computed before image-pin classification and
would otherwise leak a bogus "Skipped" state onto a dev row.

Made MonitorService's SENCHO_DEV_BUILD_AVAILABLE_KEY constant public so
both call sites share one string instead of duplicating it.

* fix(fleet): omit targetVersion for a dev-image update trigger

updateRequestInit() always forwarded latestVersion (the latest stable
release) as targetVersion whenever it was valid semver, even for a
dev-pinned node. The backend already ignores targetVersion safely for a
floating pin, so this never caused an actual repin, but it produced a
misleading "Update to X.Y.Z" button label and confirm-dialog copy for an
update that installs the dev image, not that stable release.

* feat(fleet): add integration-image badge and dev build update button

NodeCard now shows a persistent "Integration image" badge whenever a
node's compose image is any sencho-dev reference, independent of update
availability, visible to every role. When a newer dev build is available,
a solid brand-colored "Update dev build" button appears alongside it,
admin-only, reusing the existing update trigger and requireAdmin route.
Styled distinctly from the neutral stable "Update to X.Y.Z" button so an
operator always knows which channel they're acting on.

* feat(fleet): add dev-image copy to the local update confirm dialog

LocalUpdateConfirmDialog now recognizes isDevImage and shows a distinct
LOCAL - DEV UPDATE kicker plus copy stating the sencho-dev:dev reference
will be pulled without rewriting the compose image, and that integration
images are unsigned and carry no release attestations. Without this, a
dev-pinned node's update confirmation fell through to the generic "Pulls
Sencho the latest release" copy. FleetView.tsx threads isDevImage from
the node's update status through to the dialog, same source as its other
pin fields.

* feat(fleet): separate dev and stable availability in the Node Updates sheet

The sheet counted stable and dev availability together via the same
updateAvailable field, so a dev-pinned node with a build available fell
into neither the summary counts nor any row action, and would have
misleadingly rendered as "Up to date" once devBuildUpdateAvailable
existed. stableAvailable and devAvailable are now tracked separately: the
changelog dot lights only from stableAvailable (a dev build has no
release changelog), the summary and meta text report the combined total,
a dev row shows "Integration build" instead of a stable version in the
Latest column, and the existing Update button/badge now also fires for
devBuildUpdateAvailable. Update all and Skip stay stable-only, since both
already gate on fields a dev row never satisfies.

* feat(fleet): bring dev-build detection and update to Mobile Fleet

Mobile Fleet previously had no update capability at all: it only polled
/fleet/overview and never called useFleetUpdateStatus, so it could not
show the stable update flow either. It now fetches update status
alongside the overview poll, shows the same "integration" marker as
desktop on any dev-pinned node's card (visible to every role), and gives
admins a dev-build update action.

The action renders as a sibling of the card's own button rather than
nested inside it, since the card is itself a <button> and a nested
button is invalid HTML with broken touch semantics. It reuses the exact
same triggerNodeUpdate/confirmLocalUpdate flow and LocalUpdateConfirmDialog
/ReconnectingOverlay components desktop already renders, so there is no
parallel API implementation to keep in sync.

* feat(notifications): wire dev_build_update_available through the frontend

Adds the category to the frontend NotificationCategory union, its bell
label, the per-node "mute update notifications" bundle, and the bell's
friendly dot-color memo. The changelog navigation and "View changelog"
button stay scoped to node_update_available only: a dev build has no
release changelog entry to navigate to.

* docs: document dev-build detection and update on Fleet

Adds the dev_build_update_available notification category, the
persistent Integration image marker, and the dev-build update action
(desktop and mobile) to the alerts-notifications, verifying-images,
fleet-view, remote-updates, and upgrade pages. States the detection
cadence explicitly: it polls on a fixed interval and reflects the newest
build observed, not necessarily every individual build.

* fix(gitops): sanitize the inconclusive-reason debug log for log injection

CodeQL flagged the dev-build check's debug log as depending on a
user-influenced value (a registry probe failure reason can trace back to
external input). Wraps it with sanitizeForLog(), the existing repo-wide
remediation for this class of finding, matching how registry-api.ts
already handles the same pattern.

* test(gitops): cover the no-repin invariant on a dev-build self-update

Proves triggerUpdate(), called with neither targetVersion nor
targetImageRef (the exact dev-build update call), pulls the current
compose-declared ref unchanged and never stages a compose rewrite.

* fix(fleet): use the shared busy-button pattern on Mobile Fleet's dev update action

Replaces the local Loader2 plus boolean pending logic with BusyButton
so busy behavior and interaction locking stay in sync with the rest
of the app's async click surfaces.

* test(gitops): exercise the production call shape in the no-repin regression

Fleet substitutes the stable compare target when the request body omits
one, so SelfUpdateService receives a targetVersion even for a dev-build
update. The guard that protects a :dev install is therefore the semver
check inside the repin branch, not the absence of a target.

Drives triggerUpdate with a forwarded target against a floating :dev pin
and asserts the reference is pulled unchanged with no staged patch, and
pairs it with a semver case so the negative assertions cannot pass
vacuously.
2026-08-30 15:54:19 -04:00

340 lines
33 KiB
Plaintext

---
title: Fleet View
description: Monitor every node in your Sencho deployment from one screen, drill into stacks and containers, and orchestrate Sencho version updates across the fleet.
---
The **Fleet** tab is the single page you open when you want to see the whole estate at once: every node's health, every stack, every container, and which boxes need a Sencho update. It is the home for fleet-wide tabs that go beyond monitoring (Snapshots, Docker Labels, Deployments, Routing, Federation, Actions, Secrets), each linking out to its own dedicated page.
<Note>
Fleet is hub-only and is hidden from the nav strip when a remote node is the active selection. See [Multi-Node Management](/features/multi-node#what-top-level-views-show-when-a-remote-node-is-active).
</Note>
<Frame>
<img src="/images/fleet-view/fleet-overview.png" alt="Fleet page showing the masthead with the editorial 'The fleet' word, a meta line reading '4 nodes · 4 online · last sync 5s', and CPU / MEM / CONTAINERS stat tiles next to a 0-alert counter. Below, the tab strip runs Overview, Snapshots, Status, Map, Docker Labels, Deployments, Federation, Actions. Under that, the toolbar (collapsed search icon, sort, filters, Grid/Topology toggle, Node Update, Manage nodes) sits above a grid of four node cards: a pinned Local card with a cyan ★ Local rail, an Online badge, a version chip, and a 'Networking · exposed' pill, followed by Opsix, Pitt-Moba, and SLX-Mars." />
</Frame>
## Page layout
The page has three rails stacked top to bottom: a status masthead, a tab strip plus action buttons, and the active tab's content.
### Fleet masthead
A single rail summarises the state of every registered node so you can read the whole estate without scrolling.
- A **state word** that always reads `The fleet`, set in the editorial display face. The word changes color with the fleet's overall state: foreground when **Healthy** (every node online, none critical), warning when **Degraded** (any node offline), destructive when **Critical** (any online node above the CPU or disk threshold).
- A **rail tint** down the left edge of the masthead that mirrors the state colour, plus a subtle gradient wash across the bar. When Healthy, the rail carries a slow ambient shimmer; when Degraded or Critical, it switches to a steady glow instead, so an alert state reads as calm, sustained emphasis rather than a distracting pulse.
- A **meta line** in uppercase mono tracking that reads `<n> nodes · <m> online · last sync Xs`, where the sync delta uses compact units (`s` / `m` / `h`). While data is loading the line reads `syncing…` instead.
- A **reasons line** below the meta line that names exactly what is wrong, e.g. `1 offline · 1 critical`. The reasons line is hidden when the fleet is Healthy.
- Three **stat tiles** on the right edge, in mono tabular numerals:
| Tile | What it shows |
|------|---------------|
| **CPU** | Average CPU across online nodes, with a sub-line that names the peak node (`peak <name> <percent>%`). The value tints amber once average CPU is at or above 80%. |
| **MEM** | Total RAM used across online nodes, with `of <total> · <percent>%` underneath. |
| **CONTAINERS** | Active running container count across online nodes, with `of <total> total` underneath. |
- An **alerts indicator** pinned to the far right: a bell icon next to the current critical-node count, with an `alert` / `alerts` mono label. The icon and number tint destructive while the count is above zero.
### Tabs
The Fleet view is a tab strip. Every tier sees Overview, Status, Map, Docker Labels, Deployments, Federation, and Actions. Snapshots appears for admins. Secrets appears for admins. Routing is a limited-availability fleet surface and is not part of the default tab strip. A vertical separator after **Docker Labels** (or after **Map** when Docker Labels is not present) divides the per-node monitoring tabs from the fleet-wide orchestration tabs.
| Tab | Tier | What it does |
|-----|------|--------------|
| **Overview** | Community | The grid or topology view of every node and its health. Covered in the next section. |
| **Snapshots** | Community (admin role) | Snapshot every compose file across the fleet. See [Fleet-Wide Backups](/features/fleet-backups). |
| **Status** | Community | One card per node summarising which automations and security features are configured. Covered below. |
| **Map** | Community | A read-only map of how stacks, services, networks, volumes, and ports relate across the fleet, with anomaly flags. Covered below. |
| **Docker Labels** | Community | Estate-wide audit of Docker and Compose labels across every node. See [Docker Label Audit](/features/docker-label-audit). |
| **Deployments** | Community | Blueprint deployments and reconciler state. See [Blueprints](/features/blueprint-model). |
| **Routing** | Limited availability | Cross-node service routing via Sencho Mesh when that surface is enabled on the instance. See [Sencho Mesh](/features/sencho-mesh). |
| **Federation** | Community | Cordon nodes and pin blueprints to specific hosts. See [Fleet Federation](/features/fleet-federation). |
| **Actions** | Community (admin role) | Fleet-wide bulk operations: stop stacks by label, bulk-assign labels, prune Docker resources. See [Fleet Actions](/features/fleet-actions). |
| **Secrets** | Community (admin role) | Encrypted env-var bundles you push to labeled nodes across the fleet. See [Fleet Secrets](/features/fleet-secrets). |
### Action buttons
Two buttons sit next to the tab strip in the top-right corner. They are visible from every tab, not just Overview.
| Button | What it does |
|--------|--------------|
| **Refresh** | Forces an immediate re-fetch of fleet data. The icon spins while the refresh is in flight; the button is disabled until it completes. |
| **Export Dossier** | Admin-only. Walks every node and stack and downloads a `homelab-dossier.zip` Markdown archive of the whole fleet. See [Fleet Dossier](/features/fleet-dossier). |
The **Node Update** and **Manage nodes** buttons live in the Overview tab's own toolbar rather than this shared row; see [Toolbar](#toolbar) below.
## The Overview tab
Overview is the default tab and is where most operators spend their time. It offers two layouts of the same data: a card grid and an interactive topology graph.
### Toolbar
In **Grid** mode the toolbar exposes search, sort, filters, and the view toggle. In **Topology** mode only the view toggle remains (the grid-only controls collapse). **Node Update** and **Manage nodes** sit at the end of the row in both modes.
| Control | Behaviour |
|---------|-----------|
| **Search** | Collapsed to a single icon button by default; click it to expand the input (it stays expanded while a query is entered and collapses again on blur once cleared). Filters in real time against node names and the names of stacks deployed on each node. Typing `plex` keeps only nodes that have a `plex` stack; typing `dev` keeps only nodes whose name contains `dev`. |
| **Sort** | Combobox with five orderings: Name, CPU Usage, Memory Usage, Containers, Status. The selection persists in the browser. |
| **Direction toggle** | Arrow button next to the sort combobox. The icon rotates 180° and the title flips between *Switch to ascending* / *Switch to descending* to confirm the current direction. |
| **Filters** | Popover with five sections: **Status** (All / Online / Offline), **Type** (All / Local / Remote), **Severity** (Critical Only toggle), **Networking** (All / Exposed / Unknown / Drift, matching the [networking badge](#grid-view) on each card), and **Tags** (multi-select of fleet label palette, only visible if labels exist). The button shows a count badge for active filters; a **Clear all filters** action appears at the bottom of the popover when at least one filter is set. All filter and sort selections persist in the browser. |
| **Grid / Topology** | Segmented control, with a Grid icon and a Topology icon. The control resets to Grid when the page is reloaded. |
| **Node Update** | Opens the [Node Updates sheet](#node-updates) and immediately re-checks every node's version, same as **Recheck** inside the sheet. |
| **Manage nodes** | Admin-only. Navigates to **Settings → Infrastructure → Nodes**, where you register, edit, or remove nodes. |
### Grid view
Every node renders as a card. The local node is pinned at the top of the grid with a cyan accent rail, a brand-tinted gradient wash, and a `★ Local` mono label in the top-right so it can never be confused with a remote.
| Element | Description |
|---------|-------------|
| **Server icon tile** | Green when online, muted when offline. Sits next to the node name. |
| **Online / Offline badge** | Online (success palette, Wi-Fi icon) or Offline (secondary palette, Wi-Fi-off icon). |
| **Type badge** | Outline pill reading `local` or `remote`. |
| **Version badge** | The node's Sencho version in mono tabular numerals (e.g. `v0.76.3`). Hidden if the node cannot report a version. |
| **Update available** badge | Warning pill shown when a newer Sencho release is published for this node. |
| **Pinned** badge | Shown instead of **Update available** when the node's own Sencho image is pinned to a digest (or an unrecognised pin) and cannot be updated automatically from here. A tooltip explains why. |
| **Integration image** badge | Warning pill shown, for every role, whenever the local node's Sencho image is pinned to the `sencho-dev` integration repository (see [Verifying images](/operations/verifying-images)). Persistent: it shows whether or not a newer build is currently available. |
| **Critical** badge | Destructive pill with a triangle icon, surfaced when the online node is above 90% CPU or 90% disk. |
| **Cordoned** badge | Warning pill with a Ban icon, surfaced when the node is cordoned. The badge tooltip carries the cordon reason or the default *Unschedulable: new blueprint deployments skip this node*. See [Fleet Federation](/features/fleet-federation) for the full cordon and pin flow. |
| **Networking** badge | Warning pill reading `Networking · exposed`, `Networking · drift`, or `Networking · unknown exposure`, shown when the node has a stack with published ports, detected network drift, or unresolved exposure. Click it to jump to that node's Networking page. |
| **Updating / Updated / Failed / Timed out** badge | Update progress indicator, shown only while or just after an update flows through. Failed and timed-out states surface inline retry and dismiss buttons and a cursor-following error tooltip. |
| **Skipped** badge | Shown on the card when an available update has been deferred from the [Node Updates sheet](#node-updates); the same skip state is reflected there. |
| **Container stats grid** | Three cells: **Running** (active containers), **Stopped** (exited containers), **Stacks** (count, or `-` if the node has not reported). Hidden on offline nodes. |
| **CPU / RAM / Disk bars** | Each row shows the metric icon, the percent (CPU) or `used / total` (RAM, Disk), and a horizontal bar that tints amber at 60%, destructive at 80% (CPU/RAM), or amber at 75% / destructive at 90% (Disk). Hidden on offline nodes. |
| **Update to v…** button | Admin-only outline button that runs along the bottom of the card when an update is available. The label includes the latest version. |
| **Update dev build** button | Admin-only, brand-colored button that replaces **Update to v…** when the local node is pinned to `sencho-dev:dev` and a newer integration build has been published. Pulls and recreates without rewriting the pinned tag. |
| **Stack details** trigger | Footer button that toggles the stack drill-down. The label carries the stack count for the node. |
Offline nodes render dimmed, with no stats grid, no usage bars, and no update affordance.
### Node actions menu
Every card carries a three-dot **Node actions** kebab in the top-right corner:
| Action | Notes |
|--------|-------|
| **Node details** | Opens an info sheet with the node's connectivity, live capacity, Compose workload, version and update compatibility, and governance info (labels, cordon reason and date, default-node status, Compose directory, registration date). Available to anyone who can see the card; the label picker inside the sheet stays editable only for whoever holds `node:manage` on that node. |
| **Edit node** | Opens the Edit dialog prefilled with the node's connection details. For proxy-mode remotes, saving with a changed API URL or token re-runs the connection test automatically. |
| **Delete node** | Opens a destructive confirmation. The local (default) node has no Delete option. Deleting a remote only removes it from this console; the remote instance and its containers are untouched. |
| **Cordon node** / **Uncordon node** | Marks the node unschedulable so new blueprint deployments skip it. Existing deployments keep running. Requires the `node:manage` permission (admin, or node-admin when scoped to that node). |
| **Mute** submenu | Mute node notifications, mute update notifications, mute monitor alerts for this node, or open the full mute-rule manager. Shown to whoever can manage mute rules for the node. See [Alerts & Notifications](/features/alerts-notifications). |
Edit, delete, cordon, and mute stay gated on `node:manage` or mute permission as before. Every card shows the kebab with at least **Node details**, even for a viewer with no manage permissions.
### Topology view
Switch the segmented control to **Topology** to swap the grid for a hub-and-spoke graph. The local node sits in the centre with a cyan ring and the remotes radiate outward; the layout is computed automatically from the registered fleet.
<Frame>
<img src="/images/fleet-view/fleet-topology.png" alt="Fleet Topology view in Hub layout. The Local node card sits at the centre with an Online pill, a 'prod' label pill, CPU / MEM / DISK bars, and '15 stacks · 16 running' below. Three remote cards radiate to the right (SLX-Mars, Pitt-Moba, Opsix), each with an Online pill, a 'dev' label pill where assigned, CPU/MEM/DISK bars, a stack/running summary, and a millisecond latency reading. Cyan connector lines run from the local node to each remote. ReactFlow zoom controls sit at the bottom-left and the navigator minimap is pinned to the bottom-right." />
</Frame>
Each node card in the topology graph carries:
- A **status pill** at the top reading `Online`, `Critical`, or `Offline`, with a coloured dot.
- A **Local** or **Remote** label in the top-right.
- A **lightning icon** when the node is critical; a clock icon when a pilot-tunnel node's heartbeat has gone stale; the card frame mutes when the node is offline.
- The node name with a server icon.
- A row of **node label** pills, when the node carries any (matching the labels set in **Settings · Infrastructure · Nodes**), with a `+N` overflow tooltip past the first few.
- Three compact **CPU / MEM / DISK** bars with the percent value at the right of each row.
- A footer line summarising stacks and running containers, e.g. `3 stacks · 4 running`, plus a round-trip latency in milliseconds for online remotes.
**Connector lines** colour by the link's health: cyan when both ends are online, warning when the remote is critical, dashed and muted when the remote is offline.
The graph is interactive: drag the canvas to pan, scroll to zoom, drag a node to reposition it, and use the **navigator minimap** in the bottom-right corner to jump around large fleets. The ReactFlow controls in the bottom-left expose explicit zoom in / zoom out / fit-to-view buttons. Click any node card to navigate into its node detail.
The graph re-lays out only when nodes are added, removed, or change type. Live metric updates on existing nodes do not move the layout, so an operator who has dragged nodes into a custom arrangement keeps it.
#### Topology layout modes
A toolbar at the top of the topology canvas offers three layouts. Pick the one that matches how you reason about your fleet.
| Mode | What it does |
|------|--------------|
| **Hub** | Local node on the left; remotes radiate to the right. Best when you have a handful of nodes and want the gateway anchored. |
| **Grouped** | Remotes cluster by their primary node label (alphabetically first when a node has multiple). The local node sits in its own cluster. Use this to read environment, region, or role groupings at a glance. Nodes without any label fall into an Unlabeled cluster. |
| **Free** | Drag any node anywhere on the canvas. Positions persist in this browser, so the arrangement is there when you come back. |
Assign labels in **Settings · Infrastructure · Nodes** to drive Grouped mode. Each node accepts multiple labels; the alphabetically first one is its cluster.
### Stack drill-down
Click **Stack details** on any online node card to expand the stack list. The right side of the row carries a `<n> stacks` counter once the list has loaded.
<Frame>
<img src="/images/fleet-view/fleet-drill-down.png" alt="Fleet grid with the Opsix node card's Stack details expanded, showing three stacks (saelix-app-ca, saelix-app-wa, saelix-db) each with a fleet label dot. saelix-db is expanded to show one container row with a green running dot, the container name, a 'running' state badge, the image tag mariadb:10.11, and an 'Up 6 weeks' status string." />
</Frame>
Each stack row carries the stack name, an expand chevron, a `running / total` container count, and any **fleet label dots** colour-coded from the node's label palette (when labels are configured for the stack).
Click a stack row to expand the per-container drill-down. Each container shows:
- A **state dot** that is green for running, warning for restarting, destructive for exited.
- The **container name** and a **state badge** (`running`, `restarting`, or `exited`).
- The **image tag** if known (e.g. `linuxserver/plex:latest`).
- A **status string** carried by Docker (e.g. `Up 4 days`, `Restarting (1) 3s ago`).
- An **external-link button** that appears on hover and navigates straight to that stack's editor on the corresponding node.
Container data is fetched fresh each time you expand a stack, so you always see the current state. Stack data for a node is fetched once on first open and cached for the session; re-expanding a stack does not refetch the stack list.
### Auto-refresh cadence
The Overview data refreshes on its own:
| Loop | Cadence |
|------|---------|
| Per-node health, stats, stack counts | every 30 seconds |
| Sencho update-status check | every 2 minutes |
| Fast poll while any node is actively updating | every 5 seconds |
A subtle `Auto-refreshing every 30 seconds` line at the bottom of the page confirms the loop is alive. The **Refresh** button in the top-right forces an immediate poll without waiting.
## The Status tab
The **Status** tab gives you a fleet-wide rollup of which automations and security features are configured on every node, without having to open each node's settings individually.
<Frame>
<img src="/images/fleet-view/fleet-status-tab.png" alt="Fleet Status tab with four node cards. The Local card shows the full eight-row summary (Agents 1 active, Alert rules None, Auto-heal 1/1, Webhooks None, MFA Off, Scanning None, Backup Enabled, Crash detect On). Opsix, Pitt-Moba, and SLX-Mars each show a seven-row summary that omits MFA and Backup but adds a Policy sync row reading 'In sync'." />
</Frame>
Online nodes render a two-column summary grid with up to nine rows:
| Row | What it shows | Visibility |
|-----|---------------|------------|
| **Agents** | Active notification agents, formatted `<n> active` or `None` | Always |
| **Alert rules** | Per-stack alert rule count, formatted `<n> rule(s)` | Always |
| **Auto-heal** | Enabled / total auto-heal policies, formatted `<enabled>/<total>` | Always |
| **Webhooks** | Active outbound webhooks, formatted `<n> active` | Always (when not gated) |
| **MFA** | `On`, `Off`, or `Not set` | Local node only |
| **Scanning** | Active vulnerability-scan policies, formatted `<n> policy/policies` | Always (when not gated) |
| **Backup** | `Enabled` or `Disabled` | Local node only |
| **Crash detect** | `On` or `Off` | Always |
| **Policy sync** | `In sync`, or a `degraded` / `paused` pill with a tooltip naming the cause. See [Fleet Sync](/features/fleet-sync). | Remote nodes that have received at least one sync push |
Offline nodes show a muted card with the heading and the message `Node is unreachable. Configuration unavailable.` The Online or Offline indicator in the card header confirms reachability at the time of the last fetch.
The tab fetches each node's configuration in parallel; a single dead node does not block the rest from rendering.
## The Map tab
The **Map** tab draws a read-only diagram of how everything on the fleet fits together: each node, the stacks on it, and how those stacks' services connect to networks, volumes, and published ports. It is built for troubleshooting ("which stacks share this network?", "what is still claiming port 8080?", "why is this volume sitting unused?") rather than for editing anything. Nothing on this tab changes your stacks.
The map is assembled from what Docker and your compose files already describe, so it reflects the live state every time you open the tab or press **Refresh**.
### What the map shows
The graph reads left to right: each **node** branches into its **stacks**, and a stack expands into its **services**, with each service linked to the **networks** it joins, the **named volumes** it mounts, and the **host ports** it publishes. A dashed link between two services marks a declared `depends_on` relationship.
To stay readable on large fleets, the graph starts collapsed: you see the nodes and their stacks, and the element count stays small no matter how big the fleet is. Click any stack to expand its services and resources; click again to collapse it.
### Anomaly flags
A summary strip above the map counts four kinds of issue, and the affected elements are ringed in the graph (and tagged in the list):
| Flag | What it means |
|------|---------------|
| **Missing deps** | A service declares a `depends_on` target, network, or volume that is not present at runtime. The service that points at the missing thing is flagged. |
| **Port conflicts** | Two services claim the same host interface, port, and protocol on the same node, so they cannot both bind it. |
| **Orphans** | A network or volume that no container uses, or a stack that is running on the node but is not managed by Sencho. |
| **Shared** | A network or volume that more than one stack uses. Often intentional, but worth knowing before you change or remove it. |
### Filtering and the list view
- **Search** narrows the map to matching stacks, services, and resources, expanding the stacks that contain a match.
- **Node** chips limit the map to one or more nodes.
- Clicking a flag in the summary strip filters to just the elements carrying that flag.
- The **Graph / List** toggle swaps the diagram for a flat table (Node, Stack, Type, Name, State, Flags). The list has no size limit, so it is the surface to reach for on very large fleets, or when the graph suggests narrowing with a filter first.
### When a node can't be reached
The map is fleet-wide: the control instance gathers each node's view and merges them. If a node is offline or unreachable, a banner names it and the rest of the fleet still draws, so one dark node never blanks the whole map.
## Node Updates
Click **Node Update** in the Overview tab's toolbar to open the **Node Updates** sheet. From here you can read every node's current Sencho version, see which nodes have an update available, and trigger updates one node at a time or across the whole fleet. The sheet has two tabs: **Nodes** (the node table and summary) and **Changelog** (release notes for the available Sencho version, fetched from the GitHub Releases page).
<Frame>
<img src="/images/fleet-view/fleet-node-updates.png" alt="Node updates sheet titled 'Node updates · 4 nodes · 0 updates available'. A Recheck button sits in the header. Below it, four summary cards read '4 Up to date', '0 Available', '0 Updating', '0 Failed'. A 'Filter nodes…' search box sits above the node table, which has columns Node / Type / Current / Latest / Status; all four rows (Local, Opsix, Pitt-Moba, SLX-Mars) show a green 'Up to date' badge at v0.95.0. The footer reads 'LATEST VERSION v0.95.0'." />
</Frame>
### Header actions
- **Recheck** re-queries every node's `/api/meta` endpoint and re-resolves the latest available Sencho version from GitHub Releases (with a Docker Hub fallback if the GitHub API is unreachable). The button disables and shows a spinner while the check is in flight.
- **Update all (n)** carries the count of remote nodes with a pending update. Clicking it dispatches the update to every remote with a known current and latest version. Nodes with a skipped version are excluded.
### Summary cards
Four tiles tell you how the fleet is split across update states:
| Card | Meaning |
|------|---------|
| **Up to date** | Nodes whose current version equals the latest published Sencho release. |
| **Available** | Nodes with a published update they are not yet running. Combines a pending stable release with a pending dev build on a `:dev`-pinned local node. |
| **Updating** | Nodes that are currently pulling and recreating with the new image. |
| **Failed** | Nodes whose most-recent update attempt did not succeed (failure or timeout). |
### Node table
The table lists every registered node, filtered by the search box at the top. Columns:
| Column | Content |
|--------|---------|
| **Node** | Node name with a Monitor icon for local nodes and a Globe icon for remotes |
| **Type** | `local` or `remote` outline pill |
| **Current** | The node's reported Sencho version, in mono. Reads `unknown` if the node has not reported (offline, unreachable, or never connected). |
| **Latest** | The newest published Sencho release. Highlighted when newer than Current. Reads **Integration build** on a `:dev`-pinned local node, since a dev build has no version number to compare against. |
| **Status** | Either an `Up to date` success badge, an `Update` button when a newer release or dev build is available, an icon-only **Reapply configuration** control (tooltip) for Compose-managed nodes (including up-to-date rows), an in-progress / failed badge with retry and dismiss controls, or a `Skipped` badge when the version has been deferred. |
The latest-version label is resolved from the GitHub Releases API (with a Docker Hub fallback) and cached for 30 minutes. **Recheck** flushes the cache and re-resolves immediately. See [Remote Updates · Reapply configuration](/features/remote-updates#reapply-configuration) for what reapply does and when to use it.
### Skipping a version
When a node has an update available and both its current version and the target version are known, a **Skip** button appears next to the Update button. Skipping a version hides the update notification for that node until a newer version is released. The node is also excluded from **Update all** while the skip is active.
Skipped nodes show a **Skipped vX.Y.Z** badge and an **Unskip** button. Unskipping re-enables the update prompt and returns the node to the **Update all** pool. Skipped state persists across restarts and reloads.
### Changelog tab
The **Changelog** tab shows the release notes for the latest published Sencho version, fetched from the GitHub Releases page. It displays the raw release notes alongside a link to view the full release on GitHub. If the release notes cannot be loaded, a message is shown in place. A pending dev build does not light the changelog indicator or appear here: it has no release notes to show.
### What happens when you click Update
When you click **Update** on a remote row, Sencho dispatches the update command to that remote's API. The remote pulls the latest Docker image directly, then spawns a short-lived helper container that bind-mounts the compose working directory from the host and runs `docker compose up -d --force-recreate <service>`. The remote restarts with the new image; the table cell flips from **Updating** to a green **Updated** badge once the gateway detects the version change. The **Updated** state stays visible for a few seconds before the row settles back to **Up to date**.
When you click **Update** on the local row, a confirmation dialog appears first ("Update local node"). Confirming kicks off the same pull-and-recreate flow on the local Sencho instance. Because the gateway is restarting itself, a full-screen reconnecting overlay takes over the browser tab. The overlay polls `/api/health` every 3 seconds and dismisses itself once the new gateway answers. If the gateway has not returned within 5 minutes the overlay switches to a *Taking longer than expected* state with a **Reload to check** button rather than waiting indefinitely. A large image pull can legitimately run past that window, so this state is not a failure on its own.
<Note>
**Recheck**, triggering updates (per-row and **Update all**), skipping versions, and unskipping versions all require the **admin** role. Viewer and operator roles can read update status, the changelog, and the skipping state but cannot change them or force a recheck.
</Note>
## How fleet data is fetched
Fleet View queries every registered node in parallel. Each node responds independently; one slow or offline node does not block the rest. Local node data comes from the host Docker socket and system stats directly. Remote node data is fetched over the [Distributed API proxy](/features/multi-node) using each node's bearer token.
<Note>
Fleet View always runs on your control (local) Sencho instance. It is never proxied through a remote node.
</Note>
## Troubleshooting
<AccordionGroup>
<Accordion title="A remote node's Current version reads 'unknown'">
The control instance could not resolve the remote's `/api/meta` endpoint. Open **Settings → Infrastructure → Nodes**, click **Test connection** on the row, and read the toast. The most common causes are a wrong API URL or scheme, a token that has been rotated on the remote (issue a new one and update the saved row), or a firewall or reverse proxy that is not forwarding to the remote's Sencho port. Once the remote is reachable, its version surfaces on the next fleet refresh and the **Update** button becomes available again.
</Accordion>
<Accordion title="'Update all' says no updates available">
The bulk action only triggers updates on remotes that report a valid current version *and* a valid latest version. If a remote's Current reads `unknown`, it is excluded from **Update all** because the gateway cannot perform a safe version comparison. The per-row **Update** button is still available and will run a one-shot update with the latest version it could resolve.
</Accordion>
<Accordion title="Update fails with 'Remote node is unreachable'">
The gateway could not connect to the remote's API while issuing the update. Check that the remote is powered on, that the API URL saved for the row is still correct, that the network allows traffic between the control instance and the remote on the configured port, and that `docker ps` on the remote shows the Sencho container running. Once the remote answers `/api/health`, retry from the row's **Update** button.
</Accordion>
<Accordion title="The local-update overlay says 'Taking longer than expected'">
When you update the local (control) node, Sencho restarts itself and a full-screen overlay waits for the new gateway to answer. If it has not come back within five minutes the overlay stops waiting and offers a **Reload to check** button. On its own this is not a failure: a large image pull can take longer than the reconnect window. Give the pull a little more time and click **Reload to check**. If the page still does not load after the Sencho container has restarted, inspect it on the Docker host with `docker ps` and `docker logs`.
</Accordion>
<Accordion title="Free-mode topology positions don't persist">
Free mode stores node positions in your browser's local storage. The arrangement is per-browser; it does not sync across devices or users. If positions reset on reload, the most common causes are a browser session that has local storage disabled (private or incognito windows, strict tracking-prevention modes, which fall back to in-memory state for the session only), storage cleared by browser cleanup tools or a profile reset, or simply a different browser or profile. To verify storage is writable, open DevTools → Application → Local Storage on the Sencho origin and confirm the `sencho-topology-preferences` key exists after switching modes or dragging a node.
</Accordion>
<Accordion title="Grouped topology shows everything in one Unlabeled cluster">
Grouped mode clusters remote nodes by their primary node label. When no remotes carry any labels, every remote falls into the **Unlabeled** cluster and the canvas looks similar to Hub mode. A hint banner above the canvas points to **Settings · Infrastructure · Nodes** where labels are managed. Add at least one label to two or more remotes and reopen the topology to see the clusters split.
</Accordion>
</AccordionGroup>