Files
sencho/docs/features/dashboard.mdx
T
Anso b1decbb32a docs: v1 docs refresh (batch 7) (#1627)
* docs(introduction): refresh screenshots and correct stale nav/tab coverage

Replace all 5 screenshots with fresh 1920x1080 production captures.
Document the shipped Take down stack action, the Compose Labels stack
tab, and the Docker Labels Fleet tab, none of which were mentioned.
Correct the Console nav item to note it is a limited-availability
surface rather than a plain role/tier-gated view.

* docs(quickstart): refresh screenshots and correct preflight, dashboard, and menu drift

Replace all three quickstart screenshots with fresh captures and correct
several claims that drifted from the current UI:

- Document the environment preflight's 7th check (Sencho compose
  location), previously missing entirely from both the text and the
  screenshot alt text.
- Add the Stack health table's Source and Port columns, and note that
  columns are sortable and default to load order.
- Note the masthead's running-container count, live CPU/memory readout,
  and alert count.
- Fix the profile menu list: replace the nonexistent "Feedback" entry
  with "Open New Issue", drop "Appearance" (it lives in Settings, not
  the profile menu), and add the conditional "Billing" entry.
- Note that the sidebar groups stacks by Docker Compose label once
  labels are assigned, with pinned stacks first.

* docs(configuration): document missing env vars and fix JWT secret wording

Add SENCHO_UPLOAD_DIR, TRIVY_CACHE_DIR, SENCHO_MESH_RECONCILE_INTERVAL_MS,
SENCHO_MESH_PROXY_TUNNEL_IDLE_MS, SENCHO_COMPOSE_COMMAND_TIMEOUT_MS, and
SSO_OIDC_CUSTOM_ENABLED, all real environment variables that were missing
from the reference tables. Clarify that the JWT signing secret has no
environment variable at all, generated and stored in the database only,
rather than implying an unused JWT_SECRET var exists.

* docs(stack-management): refresh against current live app and source

Rewrites the Stack Management page against the production node and
frontend source: adds the Compose Labels anatomy tab (new since the
last refresh), the Mute action on the stack header and sidebar
context menu, corrects the update-available banner wording and the
sidebar context menu's lifecycle ordering, notes the Doctor tab's
severity dot and the scan-status banner, and replaces all 14
screenshots with fresh production captures.

* docs(editor): refresh against current live app and source

Replace all 8 screenshots with fresh 1920x1080 production captures (mobile
shot at phone viewport). Document the self-stack protection dialog that
blocks deploy/delete actions on the stack running the current Sencho
instance, the multi-container summary strip and Compact/Detailed density
toggle, the mutually exclusive Expand containers / Expand logs controls,
the Files tab full-screen toggle, and the post-deploy scan-status banner
on the Anatomy panel, none of which were previously covered.

* docs(editor): recapture mobile compose screenshot without redaction

The previous mobile screenshot used the plex stack, whose compose volume
mount includes a real host path that had to be blacked out and overlaid
with placeholder text, leaving a visible seam. Recapture against the
dozzle stack instead, whose only volume is the Docker socket, so the
screenshot needs no editing.

* docs(stack-file-explorer): refresh page against current app and code

Renames "Files & Volumes tab" references to the current "Files" tab
label, documents the full-screen toggle and word-wrap control, expands
the non-browsable-volume reasons to match the current containment
logic (single-file binds, the full protected host-path list, and the
Docker-unreachable named-volume case), documents the 100-item bulk
selection cap and the 5000-file/1 GiB archive limits, corrects the
non-existent DISK_FULL error code, notes that override compose files
are unprotected, and scopes the atomic-write claim to fs-backed roots.
Replaces all 9 screenshots with fresh production captures taken in the
Files tab's full-screen mode, so no other stack UI (health metrics,
logs) appears in the background.

* docs(resources): refresh against current live app and source

* docs(dashboard): refresh against current live app and source

* docs(app-store): refresh against current live app and source

Replaces all 6 screenshots with fresh production captures and corrects
the Environment Variables and About-panel metadata claims to match what
the bundled LinuxServer.io registry actually returns today.

* fix(docs): recapture app-store advanced-tab screenshot without scroll cut-off

The prior capture was taken mid-scroll to reveal the Security checkbox,
leaving the template logo and About text cropped awkwardly at the top.

* docs(global-search): refresh against current live app and source

* docs(deep-links): refresh against current live app and source

- Add App Store, Logs, Update, Console, and Audit to the URL table (previously undocumented)
- Document hub-only URL redirect behavior and cross-link to Multi-Node Fleet
- Note Fleet tab URL segments do not always match their on-screen label (Status/configuration, Map/dependencies)
- Document the no-env-files edge case (Env tab absent, falls back to compose)
- Clarify Settings section-list state is phone-only; desktop always normalizes to /settings/appearance
- Note which top-level views share identical URLs between desktop and phone
- Remove greenfield-violating temporal phrasing ("now has", "as today")
- Replace em dashes in touched bullet list

* docs(appearance): refresh against current live app and source

* docs(stack-activity): refresh against current live app and source

Add the missing "Stack taken down" event category, cross-link the
Suppressed badge to notification mute rules, list the new Compose
Labels tab in the Anatomy panel strip, and replace both screenshots
with current production captures.

* docs(dossier): fix generated-facts wording, add rollback readiness section

Corrects three Dossier tab Generated Facts rows against current frontend
behavior: the Ports count is not filtered to host-published mappings, the
Network row's "bridge" label is fixed text rather than a derived driver,
and missing env var names appear only in the Markdown export, not the
live tab. Adds a Rollback Readiness subsection under Connected Features
(previously only an orphaned Limitations bullet), links to Compose
Networking and Storage Portability, and clarifies that a missing export
section can mean either an older node build or a failed on-demand fetch,
not only a version gap. Recaptures all three screenshots against
production.

* docs(stack-drift): refresh against current live app and source

Fixes the stack detail tab bar description (Compose Labels tab was
missing, Files/Edit compose were wrongly listed as tabs), documents
that re-check only requires read access to the stack (so viewer and
auditor roles can trigger a ledger write), and cross-links the network
findings to the Networking tab's runtime drift section and the Fleet
Overview's Drift filter. Replaces all four screenshots with fresh
1920x1080 production captures.

* docs(compose-doctor): refresh against current live app and source

Corrects the rule count from 31 to 32 and documents the previously
undocumented self-managed-stack guard rule and enforcement, the
per-browser dismiss control for the summary card and Doctor tab dot,
the acknowledged summary state, the full Health-Gated Updates verdict
mapping, and the temporary exposure intent option. Replaces all four
screenshots with production captures.

* docs(compose-networking): refresh against current live app and source

Adds cross-links to the new node-wide Networking operator page, documents
the "view node networking" link and the unset/inherit exposure-intent pill
labels, corrects the stale Resources Hub network-tab reference, and
replaces all screenshots against the production node.

* fix(docs): unbreak stack-activity page render

The screenshot alt text used backslash-escaped quotes, which is
invalid MDX/HTML attribute syntax. Mintlify failed to parse the file
and silently dropped it from routing, so the page 404'd despite being
listed in docs.json navigation.

* docs(environment-guardrails): refresh against current live app and source

Fix the ENV FILES section description: it lists env_file: entries and
project files with an existence problem, not every declared env file.
A cleanly-resolving project file is already named in the Project
Environment File panel above it. Replace all three screenshots with
current production captures.

* docs(docker-label-audit): refresh against current live app and source

Add production screenshots (page had none), a Capability gating
section, a Limitations section, and a Troubleshooting accordion.
Tighten the secret-redaction heuristic description and cross-link
Stack Labels, Compose Doctor, Environment Guardrails, and Node
Compatibility.

* docs(compose-storage): refresh against current live app and source

Captured fresh production screenshots (both prior images were stale)
and fixed a leftover pre-rename card title pointing at the Files tab.

* docs(stack-labels): refresh against current live app and source

Removed the stale trailing-dot sidebar claim (that rendering was
removed from the sidebar in a prior UI pass), documented the new
label-scoped notification muting available from the sidebar, stack
context menu, and Settings panel, corrected the Fleet Actions card
copy to match the redesigned cards, added the previously undocumented
bulk-assign size cap and the node-scoped label bulk-action API, and
replaced all 8 screenshots with fresh production captures.

* docs(sidebar): refresh against current live app and source

Removes the stale per-row label dot claim (removed from StackRow), corrects the update indicator and Updates chip color from orange to fuchsia, and documents Mute/Take down in the context menu, the label group mute kebab, Ctrl+A/Esc bulk-mode shortcuts, pin-eviction toast, and the offline-node skip behavior in cross-node search. Replaces all 8 screenshots with fresh production captures.

* docs(deploy-progress): refresh against current live app and source

Corrected the Scan entry point (node-wide Scan this node from Security
Overview, not a per-stack config scan), added the missing Take down
verb and entry point, broadened recovery actions to cover restart and
rollback failures, fixed the modal status text and raw-output color
claims to match the live UI, and replaced all screenshots with fresh
captures from a throwaway demo stack on the production node.

* docs(atomic-deployments): refresh against current live app and source

* docs(health-gated-updates): refresh against current live app and source

* docs(deploy-enforcement): refresh against current live app and source

Correct the tier note to every-tier (verified against the current
policy gate, which no longer has a paid-only blocking switch), add the
Security Overview deploy-enforcement summary card, tighten the honor-
suppressions default wording, cross-link the separate pre-deploy scan
advisory dialog to avoid confusion with the block dialog, and replace
all three screenshots against the current production UI.

* docs(blueprint-model): refresh against current live app and source

Replaced all 11 screenshots with fresh production captures. Corrected
the volume-destroying-drift Enforce-downgrade claim (appeared three
times): that code path has no call sites in the runtime reconciler, so
any compose edit on a stateful or state-unknown blueprint always
re-enters Awaiting confirmation regardless of drift mode. Corrected
the Create workflow (creation does not trigger an immediate
reconciliation tick) and the Delete workflow (now requires typing the
blueprint name to confirm).

* docs(scheduled-operations): refresh against current live app and source

* docs(blueprint-model): drop tier framing, treat as Community-native

Blueprints carry no tier gate, so "available on every tier" implied a
comparison that doesn't exist. Removed the tier framing from the intro
note, the Prerequisites table, and the Security section; the role
requirement (admin to write, every role to read) already says
everything that matters.

* docs(auto-update-policies): refresh against current live app and source

Documents that every update trigger on this page, Apply now and
scheduled Auto-update tasks alike, runs through the same atomic backup,
build-aware rebuild, and post-update health gate as a manual update
from the stack editor, and is subject to the same deploy enforcement
policy gate. Corrects the readiness-computation description to scope it
to registry-image services and cross-links to Health-Gated Updates,
Atomic Deployments, Deploy Enforcement, and Stack Management. Replaces
the readiness board screenshot with a fresh production capture and
removes three orphaned screenshots left over from the page's prior
CRUD-policy layout.

* docs(auto-heal-policies): refresh against current live app and source

Correct the admin-only prerequisite: viewing policies and history is
open to every signed-in role, only creating/toggling/deleting requires
admin. Rewrite the matching troubleshooting entry to match, note the
overlap behavior between an All-services and a named-service policy on
the same container, and replace all three screenshots with a real
production policy.

* docs(webhooks): refresh against current live app and source

Replaced all three screenshots with fresh production captures (single
webhook create/reveal flow, two-card configured list), verified every
claim against WebhookService/StackOpLockService/registry.ts, and added
the previously undocumented per-stack lock skip behavior for
non-Git-source-sync actions with a matching troubleshooting entry.

* docs(global-observability): refresh against current live app and source

* docs(audit-log): refresh against current live app and source

- Fix nav terminology: Audit is a top navigation tab, not a sidebar tab.
- Correct retention claim: Admiral shows full retained history (default
  90 days), not a fixed 14-day window; the 14-day clamp only applies to
  the Community API, which has no navigation entry point at all.
- Fix settings path: Settings - Operations - Data Retention (previously
  pointed at the now-split Developer Diagnostics section).
- Restore current Recovery Vault naming (was stale Sencho Cloud Backup).
- Expand the tracked-actions list: stack take-down, label management,
  MFA reset, and notification suppression rules were missing.
- Note that assigning the Auditor role itself requires Admiral.
- Replace all four screenshots with fresh production captures (Stream,
  Table, expanded row, Data Retention) reflecting the current five-row
  retention card.

* fix(audit-log): reframe data-retention screenshot to card content only

The prior crop included the Settings sub-navigation panel, which isn't
part of what the surrounding text describes. Retake tightly cropped to
match the page's other screenshots (card content only, no chrome).

* docs(multi-node): refresh against current live app and source

Rewrites the Pilot Agent enrollment flow (now a generated compose.yaml
plus docker compose up -d, not a single docker run command) and the
Settings scope table (current registry: Stacks, Container Alerts,
Docker & Storage, Data Retention, Mute Rules, Recovery, and the
Admiral Account / Recovery Vault renames). Adds a Sencho Mesh
cross-link matching the live add-node form copy. Replaces all eight
screenshots with sanitized production captures.

* docs(pilot-agent): refresh against current live app and source

* docs(readme): refresh GitHub README against current live app

Replace all 9 screenshots (stale top nav showing a removed Console
item and missing the new Networking page); document Take down,
Drift Detection, Environment and Secrets Guardrails, Storage
Portability, Docker Label Audit, Remote Updates, and the node-wide
Networking dashboard; fix a broken non-root-user anchor link; correct
the notification channels list (no email channel exists yet) and the
global search scope (pages, nodes, and stacks, not containers or
services); rename "scan policy packs" to the current "scan policies"
terminology; broaden the App Store bullet to cover custom registries.

* docs(readme): reposition intro copy and refresh badge row

Lead with DevOps/platform/sysadmin framing instead of homelab-first,
drop pre-1.0 "production" language, state plainly that Sencho
provides a UI instead of promotional "does the work you do" phrasing,
and call out that multi-node was part of the architecture from the
start. Swap the Discussions badge for CodeQL, last-commit, open
issues, and live website/docs status badges, and add a blog link.

* docs: retake Dashboard and Global Search screenshots for accuracy

Configuration Status now shows Recovery Vault instead of the stale
Cloud Backup label, and the search palette capture includes the
Networking page entry added after the prior screenshot batch.

* docs: refresh Fleet View page against current app

Fleet View has drifted since its last rewrite: the Docker Labels tab,
Export Dossier action, Networking filter/badge, node label pills and
latency in topology, the Policy sync status row, and the node card
Mute submenu were all shipped but undocumented. The Check Updates and
Add Node actions also moved into the Overview toolbar (renamed Node
Update / Manage Nodes) and no longer sit in the shared action row.

Replaces all five screenshots with fresh production captures and
corrects the masthead's motion description (a calm shimmer on Healthy,
a steady glow on Degraded/Critical, not a pulsing dot). Also fixes a
Fleet Sync limitation that claimed no first-party sync-status panel
exists, now that the Status tab surfaces one.

* docs: refresh Fleet Dossier page against current app

Adds the page's first screenshot, documents the network exposure summary now included per stack, documents the storage-summary omission versus the Stack Dossier export, and adds two troubleshooting entries.

* docs(fleet-federation): deep rewrite for the rollout-approval gate

Federation's cordon and pin controls no longer take effect by
themselves: the reconciler requires a confirmed rollout preview
(Apply now -> Confirm Apply) before it mutates the fleet, and pinning
or unpinning always clears a blueprint's approval. Rewrote the mental
model, step-by-step, lifecycle table, security section, limitations,
practical workflows, troubleshooting, and cross-links to reflect that
gate. Also corrected the audit-log claims: cordon/uncordon/pin are not
filterable by a node.cordon/blueprint.pin action taxonomy, only by the
free-text search box against the method, path, and summary.

Replaced 4 screenshots and added 3 new ones (production fleet plus an
isolated instance to capture the rollout preview dialog in both a
safe and a blocked state), and fixed one stale pin-workflow sentence
in blueprint-model.mdx that the same audit surfaced.

* docs(fleet-federation): drop tier-availability framing

The feature isn't tier-gated, so calling out "every tier" reads as an
unnecessary advertisement. Removed the tier note and the Community/
Admiral prerequisite row, and reworded two role-visibility sentences
that had conflated tier with role.

* docs(fleet-actions): deep rewrite for the v2 action-card redesign

Rewrites the page against the shipped Fleet Action Card redesign: a
single cyan rail with per-card action-class chips instead of the old
per-card rose/purple/amber rails, plus the live blast-radius preview,
dry-run, and confirmed-target binding mechanics that replaced the old
static warning banners. Documents two previously-undocumented
endpoints (match-preview, prune/estimate), the preview-confirm-execute
model shared by all three cards, per-node locking, timeouts, and the
version-gated confirmed-target contract for mixed-version fleets. All
five screenshots recaptured against the live production UI.

* docs(fleet-actions): fix nested quote in screenshot alt text

* docs(fleet-sync): refresh against current app and document proxy gap

Retake all screenshots against the live fleet (control authoring view,
a genuine replica showing the read-only banner, and replicated
suppression rows) and document a previously-unwritten behavior: scan
policies, CVE suppressions, and misconfig acknowledgements are fetched
localOnly and are not proxied through the node switcher like most
other Security page tabs, so viewing a remote's Fleet Sync state
requires signing into that instance directly.

* docs(fleet-backups): refresh against current app and note new integration points

Replaces all 6 screenshots with fresh 1920x1080 production captures and
corrects several drifted details: the Cloud Snapshots panel's pagination
and per-item delete action, the exact Recovery Vault storage-mode label,
and the real per-file size-cap behavior (an oversized compose.yaml skips
the whole stack; an oversized .env is dropped but the stack still
captures). Documents two entry points that were missing from the page:
the pre-update snapshot checkbox in Health-Gated Updates and the snapshot
coverage nudge on the Storage Portability tab. Expands the single warning
banner description in the detail view to cover all four banner types,
notes the Snapshots tab and Recovery Vault are hub-only, and adds a
Where Fleet Backups fits table cross-linking the six adjacent features.

* docs(remote-updates): refresh against current app and add hardened-channel notes

Replace all screenshots with fresh 1920x1080 production captures and fix the
Node Updates trigger label (Check Updates -> Node Update). Add the Changelog
tab, correct the reconnect overlay's stale 'Update timed out'/Try Reloading
copy to the current Taking longer than expected/Reload to check text, and
document the Admiral Hardened Build channel's carve-outs (pinned-but-not-blocked,
entitlement-gated local update path). Disambiguate this page from the
unrelated top-level Update (Health-Gated Updates) nav tab.

* docs(node-compatibility): refresh against current app and expand capability list

Replace all screenshots with fresh production captures (node switcher,
capability lock card, connection test panel) and swap the lock-card example
from a now-hidden Host Console to Audit Log, which is directly reachable.
Bring the capability table from 26 to 35 entries to match the current
registry, add a Mute Rules row, and split out fine-grained fallback
capabilities (update-guard, service-scoped-update, cross-node-rbac,
stack-down-remove-volumes) with their non-lock-card fallback behavior.
Fix the compose-networking row, which incorrectly claimed to also gate the
node-level Networking overview. Note that Pilot Agent nodes never advertise
host-console regardless of version.

* docs(security): refresh against current app and document scan-node launcher

Replaces all screenshots with fresh production captures, documents the
Scan this node launcher and the third Scanner setup toggle (pre-deploy
scan advisory), tightens the Compose risks example list and the Policies
block-condition description to match the live app, and adds an On a
phone section.

* docs(vulnerability-scanning): refresh for exploit intel and risk-based policies

Documents exploit intelligence (KEV/EPSS evidence tags), the redesigned
risk-based scan policies (severity/known-exploited/fixable block
conditions replacing the single max-severity field), and the new
Findings badge state. Replaces all screenshots with fresh production
captures and adds a Secrets tab example.

* docs: refresh CVE suppressions page

Retook all screenshots against production (prior ones predated the
triage-status/OpenVEX UI and the edit capability). Documented the
suppression edit flow, which the page previously described as
delete-and-recreate only. Corrected the remote-node banner copy on the
Suppressions tab and added a screenshot of it. Cross-linked the
Misconfig acknowledgements panel that now sits on the same tab.

* docs(two-factor-authentication): refresh screenshots and document rate-limit layers

Replace all 11 non-count-variant screenshots with fresh production captures.
Document the throttled sign-in state precisely (kicker/hero/caption change,
not just the input), that a replayed TOTP counts toward the lockout counter,
and the separate per-network-address sign-in rate limit that sits in front
of the per-account MFA lockout.

* docs(api-tokens): refresh screenshots and document service-scoped deploy actions

Deploy Only now authorizes eight lifecycle POSTs (was six): the per-service
update and restore endpoints were missing from the scope table. Universal
restrictions gained a row for image channel management, which was already
rejecting API tokens in code but undocumented. Replaced all five screenshots
with fresh captures showing the current three-scope create flow and a
populated list with one token of each scope.

* docs(upgrade): refresh against current app and remove legacy-version framing

Fixes migration-list inaccuracies (registry credentials were never
plaintext, unlike node API tokens), rephrases the SSH/TLS and JSON-config
migration bullets to drop version-numbered legacy framing, adds a GHCR
mirror note matching the quickstart pattern, and cross-links the Hardened
Build entitlement-gated update path so digest-pinned installs aren't sent
down the manual docker pull steps.

* docs: refresh Backup & Restore against current backend

Broadens the encryption.key warning to cover everything CryptoService
actually encrypts (node tokens, Git source and Fleet Secrets, SSO/MFA,
Recovery Vault, fleet snapshot contents), not just registry credentials.
Adds the built-in backupData CLI as a no-host-tooling alternative to the
sqlite3 .backup command, cross-linked to Emergency command-line recovery.
Corrects node-token migration guidance: tokens are signed with a secret
stored in sencho.db and remain valid after a host move, so no
regeneration is required.

* docs: refresh Recovery guide against current Settings page and CLI

Documents the Settings · Recovery page's full System health / Environment /
Safe actions / Command-line hub instead of only its environment-preflight
section, adds the missing "Sencho compose location" preflight check, adds the
SSO/OIDC/LDAP lockout scenario now that disableSso.js exists, and corrects the
rollback description (exact UI label, backup-required and image-layer caveats).

* docs(emergency-cli): sync in-app page description, add CLI operational details

Aligns the "in-app Recovery page" summary with the freshly refreshed Recovery
guide (System health / Environment / Safe actions / Command-line recovery, SSO
providers, one-click command download). Adds constraints verified against the
CLI source: password/username minimums, the valid SSO provider identifiers with
an example, and a note that diagnostics.js always reports Docker as unreachable
(it runs without a live Docker connection, unlike the in-app page).

* docs(troubleshooting): refresh against current backend and nav

Corrects several claims that had drifted from the implementation:
network name validation (underscore is not a valid leading character),
the remote-update delayed-failure window (3 minutes, not 90 seconds),
and an unsubstantiated version-compatibility check on the Admiral 403
entry. Removes two entries describing legacy pre-v0.39 node behavior
that no longer applies to any currently shipped build. Replaces stale
"Profile > Settings > X" navigation with the current Settings hub
paths, and "wifi icon" with the current Test Connection label. Adds
Pilot Agent awareness to the remote-offline entry with a cross-link to
its dedicated troubleshooting section, and trims the docker-run
converter entry to point at the fuller, already-current version on the
Stack Management page instead of duplicating it with a broken anchor.

* docs(verifying-images): document the :dev GHCR integration tag

Verified the existing cosign/SBOM/VEX/tag-policy content against
docker-publish.yml and docker-preview.yml; all accurate. Added the
previously undocumented :dev/:dev-<sha> integration tag published on
every push to main via docker-dev.yml.

* docs(trivy-setup): sync install guide with node-scoped scanner and current Trivy upstream docs

Documents that Trivy installs independently per node, adds the TRIVY_BIN
override and non-PATH mounting option, fixes the deprecated apt-key install
flow and RHEL repo gpgkey placement, notes the Exploit intelligence toggle
for air-gapped hosts, and replaces the stale Scanner setup screenshot.

* docs(two-factor-admin): document lockout recovery and self-reset gotcha

Verified the reset flow, CLI recovery, and SSO toggle claims against
current source and replaced the stale admin-reset modal screenshot.
Added the failed-attempt lockout behavior (undocumented despite being
promised in the page description) and a warning about resetting your
own 2FA from the Users list, which signs you out immediately unlike
the self-service disable flow.

* fix(docs): make Settings and Self-hosting discoverable in the sidebar

The Reference and Operations groups used root pages that were only
reachable by clicking the ambiguous group label, and that click also
triggered an unwanted double action (expand plus navigate). List both
pages as explicit sidebar entries and disable global drilldown so group
headers only expand or collapse.

* docs(settings): sync reference page with current Settings Hub

Renamed License to Admiral Account throughout (Hardened Build channel
switch, image channel display, DURATION pill), corrected the password
policy, documented the new Appearance navigation and log-chip-color
controls plus font options, noted Mesh data plane is not on every
installation, fixed the label cap (50, not 100), and replaced every
screenshot with a fresh capture from the production node.

* fix(docs): apply the sidebar toggle-only fix to Start here and Product guide

These two groups had the same click-to-navigate-and-expand pattern as
Reference and Operations. Flattening them is safe here too: the
Product guide root page has no directory listing depending on it, and
the Start here root page's own hand-authored Next steps CardGroup
already covers the same links the auto-generated listing duplicated.

* docs(licensing): refresh licensing page for Hardened Build and sales-led Admiral pricing

Admiral pricing moved from self-serve checkout to a contact-sales model, and
a new Hardened Build image-channel switcher shipped in the Admiral Account
settings page; neither was reflected in the docs. Also documents the
lifetime-license edge case (no Manage subscription button, no Billing row)
and refreshes all four screenshots against the current UI.

* docs(security): sync reference page with current API-token and encryption scope

Verified every claim against the live implementation: the API-token universal
restrictions list was missing MFA management, Recovery Vault, and image
channel management; the encryption-at-rest field list was missing Recovery
Vault credentials; added a note on the password strength indicator's
recommended-vs-enforced distinction. Refreshed the SSO settings, API tokens,
and audit log screenshots against the production node.

* docs(settings): fix password policy and session-invalidation claims

Cross-checked against auth.ts while verifying the same claims on the
security reference page: the enforced minimum is 8 characters (the
12+/mixed-case/number text is a frontend strength hint, not a validated
rule), and a password change invalidates every other session for the
account rather than leaving them valid.

* docs(contact): sync contact channels with sales-led pricing and current support gating

Replace the retired contact@sencho.io with hello@sencho.io (the address actually
used in the website footer and the pricing page's Get in touch CTA), narrow the
licensing@sencho.io scope to existing-license activation/billing/refunds now that
new Admiral conversations route through hello@sencho.io, drop the stale LICENSE-file
and in-app upgrade-prompt claims (the AGPLv3 relicense removed both), correct the
Support settings path and add the published response-time targets, and remove the
unsubstantiated bug bounty mention.
2026-07-21 09:13:12 -04:00

250 lines
23 KiB
Plaintext

---
title: Dashboard
description: Real-time system health, stack load, configuration overview, fleet activity, and recent alerts for the active node.
---
The **Home** tab is the first thing you see after logging in. It surfaces the active node's overall health, live system metrics, the load on every stack, the on/off state of every automation and security feature, a fleet- or restart-activity panel, and the running alert tape, all on a single scrollable page.
<Frame>
<img src="/images/dashboard/dashboard-overview.png" alt="Sencho Home dashboard showing the Healthy state masthead with its left-edge accent rail, four-tile resource gauge strip, paginated stack health table with Source and Port columns, Configuration Status card paired with Fleet Heartbeat, and the Recent Alerts feed." />
</Frame>
## Status masthead
The masthead at the top of the dashboard is the single place to read the node's current condition.
<Frame>
<img src="/images/dashboard/status-masthead.png" alt="Status masthead in the Healthy state with a teal left-edge accent rail, the editorial state word, the meta line 'LOCAL · 4 NODES · LAST SYNC 0S', the reasons line 'All systems nominal', the RUNNING / CPU / MEM stat tiles, and an 8-alerts counter." />
</Frame>
It carries:
- A **state word** (Healthy, Degraded, or Critical) set in the editorial display face so the reader sees it first.
- A thin **accent rail** down the card's left edge that mirrors the state color: brand teal when Healthy, amber when Degraded, rose when Critical. The rail shimmers when Healthy and glows steadily otherwise.
- A **meta line** in uppercase mono tracking with the active node's name, the number of nodes registered to this Sencho instance, and the time since the last successful poll, for example `LOCAL · 4 NODES · LAST SYNC 1S`.
- A **reasons line** that names exactly which signals moved the state away from Healthy, for example `RAM 95% · 18 unread errors`. When the state is Healthy the reasons line reads `All systems nominal`.
- Three quick stat tiles on the right edge of the bar, hidden below the `md` breakpoint: **RUNNING** (`active/total`), **CPU**, and **MEM**.
- An alerts counter pinned to the far right, showing the number of unread notifications next to a bell icon and the word `alert` or `alerts`. The bell and count tint amber while at least one alert is unread.
The masthead's CPU stat tile tints amber once host CPU crosses 80% and stays amber even at 90% or higher. The MEM and RUNNING values stay neutral; the reasons line and the gauge strip below are where you read severity.
## Resource gauges
A single rail of four tiles shows the numbers that change minute-to-minute.
<Frame>
<img src="/images/dashboard/resource-gauges.png" alt="Four-tile resource gauge strip: a CPU hero tile at 0.4% with the 'avg 0% last 10m · peak 0% @ 08:05 PM' caption and a 10-minute sparkline; a MEMORY tile at 15% in brand cyan with '2.4 GB / 15.6 GB' below; a DISK tile at 59% in brand cyan with '165.6 GB / 291.7 GB'; and a NETWORK tile reading '3.3 KB/s' with the rx/tx split and a rhythm sparkline." />
</Frame>
| Tile | What it shows |
|------|---------------|
| **CPU** (hero) | Current usage with the host's core count in the kicker, a 10-minute sparkline, and a caption with `avg X% last 10m · peak Y% @ HH:MM` |
| **MEMORY** | RAM usage as a percentage with a gauge bar and the exact `used / total` split below |
| **DISK** | Mount usage as a percentage with a gauge bar and the exact `used / total` split below |
| **NETWORK** | Total throughput per second with a `↓ <rx>/s · ↑ <tx>/s` breakdown and a live rhythm sparkline |
The gauge bars (and the corresponding numeric values) pick up amber at 80% and rose at 90%. The CPU sparkline uses brand cyan as the data color and does not mark the peak, since the caption already names the peak value and time. The network strip falls back to a dashed baseline when there is no traffic to plot.
While the dashboard is loading the CPU tile reads `--` and the caption shows `collecting metrics…`; bars and sparklines render once the first sample arrives.
<Note>
**ZFS hosts:** the memory tile and host RAM alerts are ZFS ARC-aware. Reclaimable ARC cache is treated as available memory rather than used, so a large ARC does not inflate the gauge or trigger false low-memory alerts. See [ZFS ARC-aware host memory](/getting-started/configuration#zfs-arc-aware-host-memory) for how to expose ARC stats to a Docker install.
</Note>
## Stack health
A mono table of every stack discovered in the active node's `COMPOSE_DIR`, sorted so the stacks demanding attention sit at the top.
<Frame>
<img src="/images/dashboard/stack-health.png" alt="Stack health table titled 'Stack health · 15 STACKS · SORTED BY LOAD' with pagination chevrons reading 1 / 2 and STACK / SOURCE / PORT / UP / CPU / MEM / CPU · 10m column headers, listing eight stacks (cloudflared, swag, plex, bazarr, radarr, sonarr, prowlarr, tautulli) each with a row tint, Local source, published port, uptime, current CPU and MEM, and a 10-minute CPU sparkline." />
</Frame>
The row itself carries the health tint (see below); there is no separate status-dot column. The columns are:
| Column | Description |
|--------|-------------|
| **STACK** | Stack name with an orange "Update available" badge when a newer image has been detected. When per-service status is known, the badge narrows to the outdated service name or a count (for example `2 updates`); hover for the full breakdown. The badge appears regardless of the sidebar indicator setting. Sortable. |
| **SOURCE** | `Git` when the stack is linked to a [Git source](/features/git-sources), `Local` otherwise |
| **PORT** | The stack's main published port, or `--` when it does not publish one |
| **UP** | How long the oldest running container has been up, in compact units (`s` / `m` / `h` / `d`); a stopped or never-started stack reads `--`. Sortable. |
| **CPU** | Latest aggregate CPU across the stack's containers. Sortable. |
| **MEM** | Latest aggregate memory across the stack's containers, formatted in MB or GB. Sortable. |
| **CPU · 10m** | Per-stack 10-minute sparkline tinted to match the row state; warn and error rows mark the peak with a contrasting accent color |
By default, rows sort by state (errors → warnings → healthy) and then by 10-minute peak CPU descending, so an exited stack always rises to the top and the noisiest healthy stacks float above the quiet ones (the header reads `sorted by load`). Click the **STACK**, **UP**, **CPU**, or **MEM** column header to sort by that column instead; the `sorted by load` label disappears once you do. Click the same header again to flip between ascending and descending. A row is tinted amber when its 10-minute peak CPU is at or above 80% or the stack is partially down (some but not all containers exited), or rose when the stack is fully exited or its peak CPU is at or above 90%. Click any row, or focus it and press <kbd>Enter</kbd> or <kbd>Space</kbd>, to jump to that stack's editor.
The table paginates at eight rows; chevrons appear in the header along with an `N / M` indicator when there is more than one page. When the active node has no stacks at all the card renders an empty state with a layered-disks glyph and the message `No stacks found. Create one from the sidebar.`
## Configuration Status
The Configuration Status card is the at-a-glance audit of every toggleable automation and security feature on the active node, so nothing is silently off when you expect it to be on.
<Frame>
<img src="/images/dashboard/configuration-status.png" alt="Configuration Status card with four sections: Notifications (Channels 'Discord', Alert rules None, Routing None, Mute Rules None), Automation (Auto-heal policies '1 / 1 active', Auto-update schedules '1 / 1 active', Webhooks None, Scheduled tasks '1 active'), Security (MFA Off, SSO Off, Trivy installed Installed, Scan policies None), and Backups & Thresholds (Recovery Vault active, Alert thresholds Off, Crash detection On)." />
</Frame>
The card is divided into four sections.
### Notifications
| Row | What it shows |
|-----|---------------|
| **Channels** | The list of enabled delivery agents from `Discord`, `Slack`, `Webhook`, and `Apprise`, joined by commas; reads `None` when no agent is enabled |
| **Alert rules** | Total per-stack alert rules in effect, formatted `<n> rules` |
| **Routing** | Number of enabled routing rules that direct categories to specific agents, formatted `<n> routes` |
| **Mute Rules** | Number of enabled notification-suppression rules, formatted `<n> rules`. See [Mute Rules](/features/alerts-notifications#mute-rules). |
### Automation
| Row | What it shows |
|-----|---------------|
| **Auto-heal policies** | `<enabled> / <total> active` for crash-recovery policies across all stacks; reads `None` when no policies exist |
| **Auto-update schedules** | `<enabled> / <total> active` count of `Auto-update Stack` / `Auto-update All Stacks on Node` rows configured for this node; reads `None` when none are configured |
| **Webhooks** | Active inbound deploy webhooks tied to Git Sources or stacks, formatted `<n> actives` |
| **Scheduled tasks** | Active scheduled operations (backups, restarts, scripts), formatted `<n> actives` |
### Security
| Row | What it shows |
|-----|---------------|
| **MFA** | `On` when TOTP is configured for the signed-in operator, `Off` when configured but disabled, `Not set up` when there is no MFA secret on file |
| **SSO** | The active SSO provider name (`OIDC`, `Google`, `GitHub`, `Okta`, `LDAP`); reads `Off` when SSO is not enabled |
| **Trivy installed** | `Installed` when a Trivy scanner binary is available on this node, `Not installed` otherwise |
| **Scan policies** | Count of enabled vulnerability scan policies on the active node, formatted `<n> policies`; reads `None` when no policy is enabled |
### Backups & Thresholds
| Row | What it shows |
|-----|---------------|
| **Recovery Vault** | The active recovery target: `Recovery Vault` (Admiral), `Custom S3` (with ` (auto)` appended when auto-upload is enabled), or `Disabled` |
| **Alert thresholds** | The current host thresholds, formatted `CPU x% · RAM y% · Disk z%`, or `Off` when host threshold alerts are disabled |
| **Crash detection** | `On` when global container-crash notifications are enabled, `Off` otherwise |
Click any row to jump directly to the settings section that manages it.
The card refreshes automatically every 60 seconds. It also refreshes immediately whenever a scheduled task is created, edited, toggled, or deleted, so the **Scheduled tasks** row (and the rest of the card) reflects that change within a second. Other changes covered by this card, such as switching the cloud backup provider or enabling SSO, update on the next 60-second tick.
## Activity panel
To the right of Configuration Status, the dashboard shows one of two activity panels:
- **Fleet Heartbeat** when at least one remote node is registered and your role has the `node:read` permission.
- **Stack Restarts (7d)** when this Sencho instance manages only the local node, or when your role lacks `node:read` (so a deployer-level role never sees a fleet card it cannot load).
### Fleet Heartbeat
<Frame>
<img src="/images/dashboard/fleet-heartbeat.png" alt="Fleet Heartbeat card with the radio-tower kicker '4 NODES' on the right, listing four entries with green status dots: Local with the brand 'LOCAL' pill and 16 containers; Opsix with 4 containers and 75 ms latency; Pitt-Moba with 3 containers and 65 ms latency; SLX-Mars with 4 containers and 243 ms latency." />
</Frame>
The header carries a radio-tower icon, the total node count, and a `<n> unreachable` callout in destructive red whenever any node is offline. Each row shows:
- A **status dot**: green when the node has answered a recent heartbeat, amber when its state is unknown, rose when offline or unreachable.
- The **node name** in mono.
- A **LOCAL** brand pill on the local node.
- The **active container count** for that node, returned alongside the latency in the fleet-overview payload.
- A **latency** column on the right of online proxy nodes, in milliseconds. Pilot-agent nodes (which receive their commands over an outbound tunnel rather than being polled) read `n/a` since round-trip latency is not meaningful for them.
- A **last-seen** timestamp on offline rows; pilot-agent nodes use their tunnel heartbeat, regular proxy nodes use the local Sencho instance's last successful contact.
The card refreshes every 30 seconds. When you have not registered any node it shows a `No nodes registered.` empty state.
### Stack Restarts (7d)
When the dashboard renders this variant, the same column shows a 7-day rolling restart map. Each stack with at least one restart in the window appears as a row carrying:
- The **stack name** in mono.
- A **category badge** for the dominant restart reason: `crash` (rose), `auto-heal` (green), or `manual` (brand). Ties resolve in favor of `crash` over `auto-heal` over `manual`.
- The **total restart count** on the right, with a proportional fill bar tinted to match the badge.
Stacks with zero restarts are not listed individually; they are summarized in a single footer row that reads `<n> stacks stable`. When no stack has restarted in the window the card shows a green check and `No restarts in the last 7 days.` This panel refreshes every five minutes.
## Recent Alerts
The bottom-most card on the dashboard collects every triggered notification on the active node, sorted by recency.
<Frame>
<img src="/images/dashboard/recent-alerts.png" alt="Recent Alerts card listing eight local, unbadged info-level rows about stack image updates and auto-update runs for plex, radarr, and swag, each carrying a relative timestamp such as '4h ago', with a 'Clear All Notifications' button anchored at the bottom right. No pagination chevrons appear because the feed holds exactly eight entries." />
</Frame>
Each row shows:
- A **severity icon**: an info circle (brand) for `info`, a warning triangle (amber) for `warning`, an octagonal alert (rose) for `error`.
- An optional **node badge** when the alert came from a remote node, showing the source node's name. Local-node alerts skip the badge.
- The **alert message**, dimmed once the alert has been read.
- A relative **timestamp** on the right, for example `2h ago`.
A **Clear All Notifications** button is anchored at the bottom right of the card. Clicking it deletes every alert in the feed by issuing a `DELETE /api/notifications` against each node that contributed an entry; while the request is in flight the button label switches to `Clearing...` and the button disables. The button stays visible while there are alerts; when none remain the card collapses to a green check and the message `No recent alerts.`
The list paginates at eight rows; chevrons and an `N / M` indicator appear in the header when there is more than one page.
Per-stack alert rules live on each stack's **Monitor** sheet; delivery channels and severity routing are configured in **Settings · Notifications**. See [Alerts and notifications](/features/alerts-notifications) for channel setup, per-stack rules, mute rules, and the routing engine.
## On a phone
<Frame>
<img src="/images/dashboard/dashboard-mobile.png" alt="Mobile dashboard masthead reading 'Healthy' with a status dot, node switcher, and '15 stacks · 15 up · 0 dn · sync 0s' meta line, followed by a CPU hero card, a mem/disk/net three-up strip, and a Stack Health list of six rows with a 'view all' link." />
</Frame>
The phone layout condenses the dashboard to what fits a single thumb-scroll: the masthead (state word, status dot, and a compact `<n> stacks · <n> up · <n> dn · sync <age>` meta line), the CPU hero card with its sparkline, a three-up memory/disk/network strip, and a **Stack Health** list capped at six rows with a **view all →** link to the full stack list. Rows show only the stack name, its host, a 10-minute sparkline, and current CPU; tap a row to open that stack's editor.
Configuration Status, the Fleet Heartbeat / Stack Restarts panel, and Recent Alerts are desktop-only; open Sencho on a wider screen to see them. Notifications stay reachable from the bell icon in the mobile masthead.
## How health is derived
The masthead's state word and reasons line both come from a single rule against five signals: host CPU, host RAM, host disk, exited container count, and unread `error`-level notifications.
| State | Meaning |
|-------|---------|
| **Healthy** | All five signals nominal. CPU, RAM, and disk are all under 80%, no container is in the `exited` state, and there are no unread error alerts. |
| **Degraded** | At least one of CPU, RAM, or disk is at or above 80%, OR at least one container is in the `exited` state, OR at least one unread error alert exists. |
| **Critical** | At least one of CPU, RAM, or disk is at or above 90%, OR at least one container is in the `exited` state AND at least one unread error alert is present at the same time. |
Every signal that contributes to a non-Healthy state appears in the reasons line, separated by middle dots (`·`), so you can read the cause directly without hovering anything.
## Refresh cadence
The dashboard is built from several independent polling loops so that fast-moving values stay live without flooding the API for slowly-changing ones.
| Surface | Polling cadence |
|---------|-----------------|
| Container counts and host system stats (masthead, gauges) | 5 seconds |
| Stack statuses (status dots and uptime in the table) | 10 seconds |
| Historical metrics (CPU sparklines and per-stack 10m series) | 60 seconds |
| Configuration Status card | 60 seconds |
| Fleet Heartbeat | 30 seconds |
| Stack Restarts (7d) | 5 minutes |
| Recent Alerts | Pushed live over the notifications WebSocket, with a 60-second safety-net reconcile poll |
In addition, every Docker container event (start, stop, die, restart, health-status change) on the active node fires a `sencho:state-invalidate` browser event that immediately refetches container counts, system stats, and stack statuses, so the masthead, gauges, and stack health table react in well under a second instead of waiting for the next polling tick. The Configuration Status card ignores these container-level events; it only reacts early to scheduled-task changes (see the section above), so it stays on its 60-second cadence for everything else.
When you switch the active node from the node switcher, the dashboard resets every panel to its loading state and starts a new round of fetches against the new node so you never see stale data from the previous one.
## Troubleshooting
<AccordionGroup>
<Accordion title="Stack health is empty even though I have stacks on disk">
The table reads from the active node's `COMPOSE_DIR`. If you mounted the compose volume to a different path inside the container, set the environment variable in your `docker-compose.yml` so it matches what is on the host. Per-node overrides live on the **Compose directory** field of the node entry; the resolution chain is `node.compose_dir` → `process.env.COMPOSE_DIR` → `/app/compose`. After fixing the mount or the field, the table populates on the next 10-second poll.
</Accordion>
<Accordion title="Fleet Heartbeat shows a node as 'unreachable' or with a red dot">
The Heartbeat card calls `/api/fleet/overview` on the local Sencho instance every 30 seconds; that endpoint records the last contact for each remote node. A red dot means the local instance has not been able to reach that node's `api_url`, or that the long-lived API token attached to the node entry has expired or been revoked. Re-test the connection from **Settings · Nodes**, regenerate the token on the remote instance, paste the new value into the node form, and the heartbeat clears on the next tick. See [Multi-node management](/features/multi-node) for the full setup flow and [Pilot agent](/features/pilot-agent) when the offline node is in agent mode.
</Accordion>
<Accordion title="A pilot-agent node shows 'n/a' for latency">
Pilot-agent nodes connect outbound to the primary over a reverse tunnel, so there is no synchronous request-response loop to measure. The Heartbeat card uses the tunnel's last heartbeat to set the dot color, and intentionally renders `n/a` in the latency column. If the dot is red, check the agent container's logs for the first `[Pilot]` line and confirm the agent can reach the primary URL it is dialing. See [Pilot agent](/features/pilot-agent).
</Accordion>
<Accordion title="Configuration Status still shows the old value after I changed a setting">
The card refreshes every 60 seconds on its own schedule. Creating, editing, toggling, or deleting a scheduled task also dispatches a live invalidation that the card picks up within a second. Other settings changes (cloud backup provider switch, SSO provider change, alert rule edit, agent toggle) update on the next 60-second tick. If the row stays stale after a minute, hard-reload the dashboard tab to force a fresh fetch.
</Accordion>
<Accordion title="The masthead shows a 'metrics stale' chip next to the node name">
The chip appears after three consecutive failures of the live metrics endpoints (`/api/stats` or `/api/system/stats`), which usually means the Docker socket is unreachable from this Sencho instance or the metrics service has stopped. The dashboard keeps polling on every cycle; the chip describes the freshness of the visible numbers, not the polling cadence. The chip clears on the first successful response once both endpoints are within the threshold. Check the Docker daemon status on the active node first, then the Sencho container logs for `[Dashboard]` lines if the chip persists.
</Accordion>
<Accordion title="Recent Alerts shows '1 / 12' but I want to scrub through history">
The card paginates at eight rows. Use the chevrons in the header to flip pages. The full alert history (with filters and search) lives in the bell menu in the top bar; the dashboard card is a tape view of the latest entries on the active node. **Clear All Notifications** deletes the whole feed in one call, fanned out across every node that contributed an entry.
</Accordion>
<Accordion title="The masthead reads 'Critical' but my CPU is low">
Critical can fire on RAM, disk, or the combination of any exited container with an unread error alert, not only on CPU. Read the reasons line right under the meta line; it lists every signal pushing the state up. Common cases: RAM at or above 90% from an oversized service, disk at or above 90% from runaway log retention, or a crashed container that no one has acknowledged in the alerts panel.
</Accordion>
<Accordion title="The dashboard feels sluggish on a large deployment">
Toggle **Developer mode** under **Settings · Operations · Developer Diagnostics**. Both dashboard endpoints then emit a `[Dashboard:debug]` line in the Sencho container logs for every request, reporting elapsed milliseconds and the contextual fields (`nodeId`, row count, `days` window). Use the timings to identify whether the Configuration Status payload or the stack-restarts query is the bottleneck. Disable developer mode when you are done to keep the log volume manageable.
</Accordion>
</AccordionGroup>