docs: v1 docs refresh (batch 2) (#988)
* docs(atomic-deployments): refresh page around current UI and behavior
Rewrites the page to match the v1 docs refresh template. Corrects
several factual errors against the current code, fills in missing
detail, and adds a screenshot of the rollback overflow menu.
Notable corrections:
- Scheduled tasks do not run atomically; only stack editor Deploy and
Update, App Store installs, webhook triggers, and image auto-updates
pass the atomic flag through to ComposeService.
- Rollback lives in the stack editor's More actions overflow menu, not
on the action bar directly. The backup timestamp renders as a
sub-line of the menu item.
- Health probe is a 3-second window with an exit-code check on every
container labelled with the compose project name; describe this
exactly rather than as 'waits briefly'.
- Document where backups live (DATA_DIR/backups/<stack>/), why they
are kept outside the compose folder, and that the slot is one per
stack with overwrite semantics.
- Document the four streamed log markers users see in the deploy
progress modal during the atomic flow.
- Add a troubleshooting accordion group covering missing menu entry,
late crashes outside the probe window, manual-intervention message,
and the single-slot retention edge case.
* docs(deploy-enforcement): refresh page for v1 and align with current enforcement paths
Update the page to match the current pre-flight gate behavior, the v1 modal chrome on the
block dialog, and the AccordionGroup troubleshooting pattern used across the v1 docs.
Drift items corrected:
- Replace the broken vulnerability-scanning/deploy-blocked-dialog.png reference with three
fresh captures under docs/images/deploy-enforcement/ (policy list, policy editor, block
dialog).
- Drop "Recreate from the stack actions menu" and the git-source apply pre-flight claim;
neither path runs the gate.
- Add bulk label deploy and the auto-update scheduler to the enforced code paths, with a
dedicated subsection for the auto-update interaction (alert-and-skip, not 409).
- Drop the false claim that severity chips in the block dialog are clickable; the dialog
is informational.
- Document the compose-parse-fails-closed branch with its synthetic violation label.
- Refresh dialog copy to reflect the v1 ModalDestructiveHeader (kicker, title, button
variants).
- Convert the troubleshooting Q&A into AccordionGroup blocks and add accordions for the
compose-parse-error case and the auto-update-skipped case.
- Quote the verbatim audit-log summary format.
* docs(blueprints): refresh against v1 UI and add federation/state-review coverage
* docs(git-sources): refresh page against v1 UI and current behavior
Rewrites the page against the v1 docs refresh template (Note tier-gate,
sectioned anatomy, AccordionGroup troubleshooting), aligning prose with
the live UI labels and the current code paths.
Corrections:
- Authentication toggle reads "Public (no auth)" / "Personal Access
Token" (not "None"), and apply mode "Auto-write files" (not
"Auto-write").
- Diff dialog kicker is GIT . PULL PREVIEW; local-edits state opens an
Overwrite local edits? confirmation modal whose primary button is
Overwrite and apply.
- Sidebar pending indicator is a small GitBranch icon, not a brand-color
dot, and the image-update dot takes priority over it on the same row.
- Pending update banner appears in the panel; Review re-fetches the
commit and opens the diff (no client-side payload caching).
Adds coverage for:
- Anatomy of the panel (pending banner, form, last-applied stat strip,
footer actions).
- 10-second webhook debounce window.
- Pending compose/env content is encrypted at rest in the database, not
just the token.
- Auth/host failures map to HTTP 400, never 401, so they do not sign
the user out.
- Per-stack lock serializes pull, apply, and create-from-git so a
webhook firing during a manual apply waits rather than racing.
- Compose validation has a 10-second budget; clone fetches have a
30-second timeout.
- New troubleshooting accordion for Pending commit has changed since
this pull was fetched.
Recaptures all five screenshots from the v0.74.x production node,
signed in as admin: panel, create-from-git tab, pull-preview diff
dialog, sidebar GitBranch pending icon, webhook Action select with
Git source sync highlighted.
* docs(stack-labels): refresh page for v1 sidebar grouping and fleet-action surface
- Lead with the v1 behavior the previous page did not cover: the sidebar
groups stacks under collapsible label headers (PINNED first, label
buckets sorted by stack count desc then name asc, UNLABELED last)
with a count chip per group. Trailing colored dots on each row
(max 3 + N overflow, paid-only) supplement the headers.
- Drop the stale claim that a label-pill filter bar lives between
search and the stack list; that UI no longer exists.
- Drop the right-click-on-pill bulk actions table (Deploy all / Stop
all / Restart all). The legacy per-node action endpoint stays in
the backend but no longer has a UI binding, so the page documents
only what users can click today.
- Document the two Skipper+ Fleet Action cards: Stop fleet by label
(name match across nodes, autocomplete, per-node breakdown,
HTTP 429 on per-node concurrency) and Bulk label assign (per-node,
replace semantics, clear on empty selection).
- Document the inline 'New label' form inside the stack right-click /
three-dot Labels submenu, the Settings - Advanced - Labels masthead
N/50 stat, the LABELS - NEW / EDIT modal kickers, and the
LABELS - DELETE - IRREVERSIBLE confirmation copy verbatim.
- Document the Fleet Overview Tags multi-select filter (filters by
stack labels aggregated across nodes), with cross-link to fleet-view.
- Capture every screenshot fresh from production signed in as admin:
sidebar-grouping, context-menu-labels, inline-create-form,
settings-labels, create-label-dialog, fleet-tags-filter,
fleet-actions. Drop the now-stale sidebar-with-labels,
sidebar-filtered, and bulk-actions-menu captures.
* docs(dashboard): refresh page for v1 layout (status masthead, gauges, fleet heartbeat, restart map)
Aligns docs/features/dashboard.mdx with the redesigned Home tab. Replaces the obsolete
Recent Activity feed coverage with the actual DashboardActivityCard split (Fleet Heartbeat
when remote nodes are registered, Stack Restarts (7d) otherwise) and recaptures every
screenshot from the v0.74.x production node.
* docs(global-search): refresh page for v1 palette
- Note tier and role gating on the Pages list (Auto-Update, Console,
Schedules, Audit) so the prose matches what the top bar exposes.
- Document the ACTIVE chip on the currently active node row.
- Document the 50-result cap counter and the Searching... loading state.
- Mention the ~250 ms debounce and clarify that filename matching
includes the file extension.
- Replace stack screenshot with a redesigned capture and add empty-state
Pages and Nodes captures showing the ACTIVE chip.
* docs(global-observability): refresh page for v1 layout (masthead, signal rail, filter strip, paused-resume chip)
Full rewrite against the current Logs tab and the v1 docs refresh template
(hero Frame, sectioned anatomy, AccordionGroup troubleshooting, refresh-cadence table).
Replaces the single overview screenshot with seven captures under
docs/images/global-observability/ (overview, masthead, signal-rail,
filter-strip, feed-bands, paused-resume-chip, error-only-filter), all from
the v0.75.x production node signed in as admin with PII scrubbed
(profile chip patched to AD, in-feed LAN IPs and third-party hostnames
substituted via DOM injection while the stream was paused).
Aligns prose with the actual UI labels and code:
- Masthead kicker reads LIVE LOGS · NODE · <NAME> with LOCAL for the
local node; state word toggles Streaming / Idle / Offline; SESSION
uses uppercase letter suffixes (1H 43M / 0M 12S) per formatUptime.
- Signal rail tile counts are scoped to the 2000-entry buffer and reset
with Clear; CONTAINERS is buffer-bound, not a monotonic accumulator.
- Filter strip controls quoted verbatim (Stacks · All / Stacks · n,
segmented controls All / Out / Err and All / Info / Warn / Error).
- Feed row anatomy: severity dot, timestamp, brand-cyan container name
with stack/container tooltip, message tinted by source. Row tint
follows detected level, which is regex-based, so an STDOUT line
containing ERROR: still classifies as ERROR.
- Day bands: NOW, Nm AGO, Nh AGO, calendar date.
- Empty states: two-tier kicker over caption (Awaiting events / No matches).
- Pause keeps the SSE buffer filling up to the 2000-entry cap; resume pill
reads <n> NEW · RESUME and counts the queue, not total arrivals during
the pause.
- Download filename and row format quoted: sencho-logs-<ISO8601>.txt and
[<ISO>] [<stack>/<container>] <LEVEL>: <message>.
Documents behavior the previous page never covered:
- Active-node scoping; node switch resets the stream and the buffer.
- SSE primary transport with 30-second server heartbeat and a 5-second
polling fallback against /api/logs/global (server-capped at 500 lines
per snapshot).
- Initial replay of the last 500 lines per container when the SSE
connection opens, so the feed has context immediately.
- Display limits (2000 client buffer, 300 rendered rows, Showing last
300 of N overflow notice).
- Refresh cadence table covering UI tick, flush cadence, polling
cadence, SSE heartbeat, sparkline window, and the Idle threshold.
Adds a seven-accordion troubleshooting block (Offline state, gray Idle
dot, ERROR-without-tint, growing Resume pill, Clear-cutoff lag,
node-switch buffer drop, fleet-wide aggregation expectations).
Tightens the closing Note so it makes clear that Notification Log
Retention does not govern this live container stream.
* docs(alerts-notifications): refresh page for v1 and absorb notification-routing
Full v1 template rewrite of /features/alerts-notifications. Bundles in
the entire Notification Routing page so a reader sees channels, routing,
per-stack rules, and retention in one place; deletes the standalone
notification-routing.mdx and points all five cross-link sites at the new
in-page anchor.
* docs(alerts-notifications): drop "What's not in scope" section
The page should describe what Sencho does, not enumerate what it does
not ship. Users find missing integrations through the Webhook section
and the routing matcher reference; the explicit disclaimer added noise
without adding guidance.
* docs(audit-log): refresh page for v1 layout, expanded action list, troubleshooting accordion
- Clarify that the search/method/date filter strip lives in Table view only.
Stream view always shows the unfiltered chronological feed.
- Fold the total-entries readout into the card subtitle wording where it
actually renders, instead of describing it as a separate header element.
- Sharpen the Peak hour off-hours window to the literal 08:00 to 17:59
working window the tile keys off, plus the 5% / 20% failure-rate tints.
- Note that the Actors tile names a sample actor alongside the new-IP count.
- Expand the example actions list to cover surfaces that have shipped since
the last edit: per-service stack lifecycle, node cordon/uncordon, fleet
replica role changes, Sencho Cloud Backup operations, Fleet Secrets, and
blueprint federation pin updates.
- Correct the Settings path: Settings · Developer · Data retention card,
Audit log input, Save settings button.
- Add a Troubleshooting AccordionGroup matching the rest of the v1-refresh
pages: missing tab, filter scope, anomaly thresholds, export cap, and
retention pruning.
- Replace all four screenshots with fresh captures of the current UI.
* docs(multi-node): refresh page for v1 layout, pilot agent mode, refreshed table columns
Rewrites docs/features/multi-node.mdx against the current product. The previous page predated the v1 Settings hub redesign and the Pilot Agent enrollment model, so it documented only the Distributed API Proxy add-node flow and missed the new Mode, Endpoint, and Labels columns on the Nodes table.
Restructures the page into 13 sections: intro, How it works, the local node, Choose a remote mode (decision table comparing Pilot Agent vs Distributed API Proxy), Add a remote node: Pilot Agent (three steps plus re-enrollment), Add a remote node: Distributed API Proxy (three steps), Switching between nodes, the Nodes table (full column reference), What Settings apply per node (verified against settings/registry.ts), License enforcement across nodes, Editing and deleting nodes, Security (token security, transport encryption, why no application-layer TLS), and Troubleshooting (AccordionGroup matching the v1 template used on audit-log, atomic-deployments, and deploy-progress pages).
Refreshes seven screenshots against the production node signed in as admin, scrubbing IPs and usernames before capture: full Nodes panel overview, Generate Node Token card with a placeholder token, Add node modal in Pilot Agent mode, Add node modal in Distributed API Proxy mode (with the inline plain-HTTP warning visible), Edit modal showing the Regenerate enrollment token card for a pilot agent, Pilot enrollment modal with the docker run command, refreshed node switcher popover, and a close-up of the table columns. Drops the obsolete add-node-form.png, http-warning.png, and per-node-scheduling/ folder.
* docs(fleet-view): refresh page for v1 layout, expanded tabs, cordon, sheet-based updates
- Aligns the Overview, Status, and Node Updates content with today's UI:
the masthead's `The fleet` headline plus CPU / MEM / CONTAINERS stat tiles,
the eight-tab strip (Overview, Snapshots, Status, Deployments, Traffic,
Federation, Fleet Actions, Secrets) with per-tier visibility, and the
Check Updates surface that is now a system sheet rather than a modal.
- Documents the toolbar (search, sort, filter popover with Status / Type /
Severity / Tags sections) and the Grid / Topology segmented control
including the topology graph's status pill (Online / Critical / Offline),
connector colouring, ReactFlow controls and minimap.
- Documents the per-card surfaces that were missing from the prior page:
Cordoned badge with cross-reference to Fleet Federation, fleet stack
label dots in the drill-down, container drill-down rows (state dot,
badge, image, status, open-in-editor hover button), and the Admiral
three-dot Node actions menu for cordon / uncordon.
- Documents the Node Updates sheet anatomy (Recheck and Update all (n)
header actions, four summary cards, node table columns, Update flow,
reconnecting overlay timing, admin enforcement) and the GitHub Releases
with Docker Hub fallback resolution path with its 30-minute cache.
- Replaces every stale screenshot with a fresh capture (overview,
topology, drill-down, status tab, node updates sheet) and removes the
obsolete files plus the empty docs/images/fleet/ folder.
- Reformats troubleshooting as an AccordionGroup matching dashboard,
multi-node, and audit-log refreshes.
* docs(fleet-backups): refresh page for redesigned fleet and settings UI
Replace all six screenshots with current production captures. Update
content to reflect the new fleet header card, eight-tab layout, full-
page Cloud Backup settings with header stats, and corrected navigation
paths. Add cloud backup rows to the access control table.
* docs(fleet-backups): convert troubleshooting to AccordionGroup pattern
Match the foldable-accordion pattern used across the v1 docs refresh
batch. Merges the standalone Cloud Backup troubleshooting subsection
into a single Troubleshooting section at the bottom of the page with
seven accordions covering skipped nodes, two restore failure modes,
three cloud-upload failure modes, and a diagnostic logging entry.
* docs(remote-updates): refresh page for v1 sheet, accordion troubleshooting, factual fixes
Rewrites the page against the v1 docs refresh template (Note tier gate,
sectioned mechanism deep-dive, Frame screenshots with detailed alt text,
inline AccordionGroup troubleshooting), bringing it in line with the
recently-refreshed fleet-view, fleet-backups, dashboard, and audit-log
pages.
The page is repositioned as the mechanism deep-dive (prerequisites, what
runs on a node during an update, completion and failure detection,
recovery actions). The full UI tour for the Node updates sheet remains in
fleet-view so the two pages stop overlapping; remote-updates now links
into fleet-view#node-updates instead of restating the table anatomy.
Captures three screenshots from the production node, signed in as admin:
fleet-node-updates.png shows the Node updates sheet with eight nodes and
seven remote updates available; local-update-confirm.png shows the
LOCAL · UPDATE alert dialog with the Cancel and Update & restart buttons;
node-card-update-available.png shows the Opsix card with the Update
available pill and the Update to v0.76.7 outline button.
Corrects several factual claims that no longer matched the current code:
- The remote early-fail threshold is about 3 minutes, matching
EARLY_FAIL_MS in backend/src/routes/fleet.ts, not 90 seconds.
- The Recheck button sits in the sheet header, not the footer.
- The component is a SystemSheet, so the page now consistently calls it
the Node updates sheet instead of a dialog, with lowercase "Node
updates" and lowercase "Update all (n)" matching the live UI.
- Reconnecting overlay polls /api/health every 3 seconds, not "every few
seconds".
- The local Failed badge surfaces as soon as the helper writes its error
file, by the 3-minute mark at the latest.
Documents the LocalUpdateConfirmDialog kicker, title, body, and CTA
verbatim, the Triggering... loading state on the Update buttons, the
four completion signals the gateway accepts (version change, process
startedAt change, offline-then-online transition, version at or above
the comparison target after 15 seconds), and the 60-second auto-clear
of the Updated badge.
Drops references to two screenshots that never existed
(fleet-node-updating.png, fleet-node-failed.png); the in-flight and
failed states are described in prose instead, the same way fleet-view
handles them.
* docs(scheduled-operations): refresh page for v1 timeline, fleet-wide update action, sheet-based run history
Rewrites the Scheduled Operations page against the v1 template
(Note tier gate, sectioned anatomy, Frame screenshots, AccordionGroup
troubleshooting) applied to sibling pages in this batch. Captures
seven fresh screenshots against the production node signed in as
admin (timeline, all-tasks, action-picker, create-restart,
create-prune, create-scan, run-history) and removes every legacy
PNG.
Documents the new "Auto-update All Stacks" action that was absent
from the page, extends the Skipper allow-list to all four Skipper
actions (Auto-update Stack, Auto-update All Stacks, Fleet Snapshot,
Vulnerability Scan) and clarifies that the action picker hides
operations the active tier cannot run.
Corrects several factual claims that no longer matched the code:
- Scheduled scan completion is `info`/`scan_finding` on a clean run
and `warning`/`scan_finding` when findings are present (not
`info`/`system` as previously stated). Cross-link now points at
`alerts-notifications#vulnerability-scanning`.
- Lifecycle actions (auto_backup, auto_stop, auto_down, auto_start)
execute against the local Sencho instance only; only Auto-update
Stack / All Stacks have a remote-proxy code path. The page
reinstates the guidance to schedule remote lifecycle operations
from that node's own UI.
- Run history lives in a right-side sheet with a "Schedules ›
<task> › Runs" breadcrumb and a Download CSV secondary action.
- Timeline masthead is described in terms of the v1 visual
(`NEXT 24 HOURS` kicker, italic display heading, monospace date
range, right-anchored Next pill with countdown, glowing cyan now
rail, six-tick bottom axis).
* docs(rbac): refresh RBAC & user management page against v1 template
Bring /features/rbac onto the v1 docs refresh template (Note tier gate,
sectioned anatomy, Frame screenshots, AccordionGroup troubleshooting).
Recapture five screenshots from the production node signed in as admin
and remove the three stale captures under docs/images/rbac/.
Corrections vs. the prior page:
- Deployer no longer claims node:read in the permission matrix; the
backend grants only stack:read and stack:deploy.
- Add the system:registries row (container registry management).
- Document the form as inline below the Add user button (not a modal).
- Note the (you) marker on the signed-in admin's row and the disabled
delete icon on that row.
Additions:
- Settings nav location and hub-only visibility.
- 2FA reset row action with verbatim modal kicker, title, and body.
- Five-failure / 15-minute MFA lockout behavior and admin reset recovery.
- Token-version session-security table covering deletion, role change,
password change, and admin 2FA reset.
- SSO password-fields-hidden line quoted verbatim and the per-provider
Require MFA toggle.
- Audit-log emissions list for every user-management mutation.
- API tokens cross-link explaining the user-vs-machine boundary.
- Scoped permissions section retightened: scoped role picker is
Deployer / Node Admin / Admin only; resource type is Stack or Node.
AccordionGroup with eight troubleshooting entries covering missing nav,
greyed role options, seat-limit errors, unexpected sign-outs, scoped
deployer mismatches, missing shield icon, re-locking MFA accounts, and
SSO role drift at provisioning.
* docs(2fa): refresh two-factor authentication and admin guide against v1 template
Bring /features/two-factor-authentication and /operations/two-factor-admin
onto the v1 docs refresh template (Note tier gate, sectioned anatomy, Frame
screenshots with descriptive alt text, AccordionGroup troubleshooting,
verbatim modal copy with kicker callouts). Recapture every screenshot under
docs/images/two-factor-auth/ from a fresh session and add six new captures
for surfaces the prior page did not document.
Corrections vs the prior pages:
- Panel rename: Settings -> Account & Security is now Settings -> Account,
under the Identity group of the settings sidebar. Replaced every
occurrence on both pages.
- Enrol dialog titles match the current modal: Pair your authenticator,
Confirm the pairing, Save your recovery codes (was: Set up 2FA, Confirm,
Save your backup codes). Step rail 01 PAIR / 02 CONFIRM / 03 ARCHIVE
documented.
- Manual-entry affordance is the always-visible Secret manual entry row
with a copy icon, not the toggleable Can't scan Show secret key link.
- Confirm step auto-submits on the sixth digit; no submit button. Verified
in MfaChallenge.tsx and MfaEnrollDialog.tsx and called out explicitly.
- Authenticator-app list trimmed to match in-app copy (1Password, Bitwarden,
Google Authenticator, or any TOTP app). Authy and Microsoft Authenticator
dropped because the dialog does not mention them.
- Disable dialog: kicker SECURITY MFA DISABLE, title Turn off two-factor,
destructive header, Disable button. Replaces the prior Disable 2FA
paragraph that did not describe the dialog chrome.
- Regenerate dialog: two-step flow with kicker SECURITY BACKUP CODES, Confirm
identity then New recovery codes, with the verbatim PREVIOUS CODES HAVE
BEEN INVALIDATED warn rail on the show step. Documented that the dialog
only accepts a TOTP, not a backup code.
- Per-user SSO toggle label corrected: Require 2FA on SSO sign-in (was:
Require 2FA even when signing in via SSO). Added the per-provider vs
per-user distinction on both pages (admins can also enable Require MFA
on the SSO provider config, which is independent of the per-user toggle).
- Admin reset modal: verbatim USERS RESET 2FA kicker, Reset 2FA for
<username> title, full-body copy reproduced. Documented that the reset
bumps the target's token version and invalidates active sessions.
Additions:
- Sign-in throttle: five failed verifications lock the account for 15
minutes, server returns 423 with Retry-After, UI shows the Retry in MM:SS
countdown plus Rate limited label. Lockout recovery section explains
that the counter only clears on a successful sign-in, so retries after
the window expires re-lock immediately.
- Account panel anatomy section enumerates the three rows (Authenticator
app, Backup codes, Require 2FA on SSO sign-in) plus the destructive
Disable 2FA link, and the masthead 2FA on / BACKUP N left chips.
- Recovery codes section now covers all three count states (3 plus, 1 to 2,
0) with verbatim helper text, tone, and the standalone No backup codes
left callout that renders at zero. New screenshots for the 2-remaining
and 0-remaining states.
- Cross-references to the admin operations page (CLI fallback, token version
rotation, what a reset changes in the DB), the SSO page, and the RBAC
page (per-provider Require MFA toggle, SSO auto-provisioning).
Troubleshooting on the feature page rewritten as an AccordionGroup with
nine entries: clock drift, wrong account selected, QR will not scan, lost
phone with no codes, lost codes with authenticator, ran out of codes,
unexpected SSO prompt (with both toggle causes), repeated lockout after
the window expires, missing shield icon on Users panel.
The admin operations page also gains the SSO + 2FA two-toggles table so
administrators can answer the per-user vs per-provider question without
context-switching between pages.
Six new images added; six existing images replaced. Total 14 captures.
* docs(rbac,host-console): drop enforcement-boundary detail from tier-gate notes
Operator-facing docs should state tier or role requirements once, in plain
customer-facing language, and leave the enforcement chain to the source.
Two surfaces on the v1-refreshed pages over-specified the gate:
- `features/rbac.mdx::Scoped permissions`: the Note enumerated both the UI
hide on Skipper and the `/api/users/:id/roles` write rejection. The first
half ("Scoped permissions require Admiral.") is the operator-relevant
fact; the rest reads as a fence specification, which is awkward for an
open-core product where the gate is readable in source anyway. Trimmed
to just the tier claim.
- `features/host-console.mdx::Availability`: the paragraph already says
who can use the console and that the Console tab is hidden on Community
or Skipper. The trailing "Attempting to access the console endpoint
directly without the correct license or role is rejected" is the same
bypass-prevention coda. Dropped.
No functional behavior change; the gates themselves are untouched.
* docs(sso): refresh SSO & LDAP authentication page against v1 template
Rewrites docs/features/sso.mdx against the v1 docs refresh template (intro
+ tier callout, sectioned Configuration anatomy, Frame screenshots,
AccordionGroup troubleshooting), bringing it in line with the previously
refreshed two-factor-authentication and rbac pages on this branch.
Recaptures all four screenshots from the production node signed in as
admin: sso-settings (overview with the five collapsible provider cards),
sso-settings-ldap (LDAP form expanded), sso-settings-oidc (Google form
expanded), sso-settings-custom-oidc (Custom OIDC form expanded with all
eleven fields).
Refreshes the Settings UI section to match the redesigned panel: each
provider is a collapsible card with an Active badge on the header, an
enable / disable toggle pill, and a footer with Save, Test Connection
(green check or red X next to the button), and Remove (only after a
config has been saved). Documents the static callback-URL helper that
sits below all five cards.
Clarifies that the per-OIDC claim mapping environment variables
(SSO_OIDC_*_ID_CLAIM, *_USERNAME_CLAIM, *_EMAIL_CLAIM) are accepted for
Google, GitHub, and Okta, not just Custom OIDC. The Settings UI hides
those fields on the presets because the defaults match.
Converts the troubleshooting section to an AccordionGroup with five
entries (Test Connection discovery failure, issuer validation error,
wrong username or missing email after sign-in, invalid redirect URI,
SSO buttons missing on the login page). Cross-links the operations
troubleshooting page for setup-time errors.
Tightens the LDAP TLS env var note to spell out the literal string
'false' requirement. Syncs the Combining SSO with 2FA section to use
the live toggle label 'Require 2FA on SSO sign-in'.
* docs(sso): drop the Community-tier Custom OIDC workaround tip
The Tip walked through how a Community-tier operator could integrate
Google, GitHub, or Okta by pointing Custom OIDC at the provider's
discovery URL, bypassing the Skipper preset gate. Operator docs should
state the tier rule once and stop; they should not describe how to
circumvent it.
The tier matrix above the removed block already names which providers
are paid; the Custom OIDC row already lists "any spec-compliant OIDC
provider" as its scope. That is enough.
* docs(vulnerability-scanning): refresh page for v1 UI and corrected tier mapping
The page was last revised before the v1 visual redesign and before the
tier-mapping changes shipped in v0.81.2 (open Community access to
secret scanning, compose misconfig scanning, scan history, and scan
comparison). This refresh:
- Rewrites the tier matrix to match the shipped Community / Skipper /
Admiral split. Secret detection, compose misconfig scanning, scan
history, scan comparison, and misconfig acknowledgements are now
correctly marked as Community. Scheduled fleet scans, scan policies
with block_on_deploy, SBOM, SARIF, and Trivy auto-update stay paid.
- Drops two stale Notes that said secret detection and compose
misconfig scanning required Skipper or Admiral. The page now states
each tier requirement once, in plain language.
- Refreshes all six existing screenshots from the production node:
resources-badges, scan-details-sheet, scan-history-sheet,
scan-compare-sheet, security-settings, app-store-toggle.
- Adds a new scan-config-button screenshot showing the stack-page
overflow menu where Scan config now lives.
- Describes the scan drawer header accurately: Re-scan + Compare + CSV
+ SARIF as top-level buttons, with SBOM as a separate button below
the summary.
- Updates the compose misconfig flow to point at the stack overflow
menu (not the Deploy controls).
- Converts the troubleshooting section to a single AccordionGroup per
the v1 template, and audits each entry for legacy phrasing and the
removed tier claims.
- Adds a TRIVY_BIN reference to the How it works section so operators
know about the host-binary override.
* docs(cve-suppressions): refresh page for v1 UI and corrected suppression specifics
- Recapture all three screenshots from the production node signed in
as admin under `docs/images/cve-suppressions/` (`settings-panel`,
`create-dialog`, `suppressed-row`). The previous file referenced
three image paths that did not exist in the repo.
- Align prose with the actual UI labels:
- Dialog kicker `SUPPRESSIONS . NEW`, title `New suppression`.
- Field labels match the form: `CVE or advisory ID`, `Package
(optional)`, `Image pattern (optional)`, `Reason`, `Expires in
(days, optional)`.
- Remove confirmation reads `Remove suppression` with kicker
`SUPPRESSIONS . REMOVE . IRREVERSIBLE`.
- Factual corrections:
- Fleet sync truncation cap is 5,000 rows (not 10,000).
- State the admin-role requirement once in the lead Note.
- Drop references to a `Fleet . Sync status` page and a `Reanchor`
button; neither exists in the UI. The reanchor flow is an admin
API call and is documented in /features/fleet-sync.
- Sharpen the specificity scoring section (package + image scores
3, package only 2, image only 1, neither 0) so the order matches
the read-time filter logic.
- Note that the image-pattern glob is case-sensitive.
- New coverage:
- Suppressing directly from a scan result, including which fields
are read-only in that inline flow and when to fall back to
Settings to broaden scope.
- The `replicated` and `expired` row badges in the panel.
- Hovering the package column on a suppressed row to surface the
Reason.
- Two distinct read-only modes: viewing a remote node from the hub
(panel hidden, banner shown) versus signing into a replica
instance (panel visible, read-only).
- SARIF export carries suppressions through as
`kind: external, status: accepted`, cross-linked to the
Vulnerability Scanning page.
- Convert troubleshooting to AccordionGroup with six entries; update
the truncation entry to reflect the 5,000-row cap.
* docs(private-registries): refresh page for v1 UI and fleet-wide credential model
Rewrites the page against the v1 docs template (Note tier gate, opening Frame,
sectioned anatomy, AccordionGroup troubleshooting) and replaces every
screenshot with a fresh capture taken against the current product.
Corrects several factual claims that no longer matched the current code:
- Registries are stored once on the control instance and applied fleet-wide,
not configured per node. The old Multi-node behavior section and the
matching troubleshooting entry described a per-node model that the product
no longer has.
- The Registries section is hidden on remote nodes (global scope) and on
Sencho versions that do not surface the feature. New troubleshooting
entries explain both visibility states.
- The feature is admin-only on Admiral. Non-admin operators do not see the
section even on Admiral; previous copy implied any Admiral license user
could manage credentials.
- Registry endpoints are not reachable from API tokens; only an admin
browser session can manage credentials. The Security section now states
this without naming internal route paths.
Documents UI behavior the previous page omitted: the inline form (not modal),
the four type-specific form variants, the Docker Hub read-only URL field, the
destructive delete confirmation with its stack-pull warning, the masthead
REGISTRIES count, and the empty-state callout copy.
Screenshots replaced:
- registries-overview.png: section with one configured GHCR card and the
masthead stat at one.
- registries-empty.png: empty state with the Add registry button and callout.
- registries-add-form.png: inline form with the Docker Hub default and the
read-only URL field.
- registries-ecr-form.png: form switched to ECR, showing the AWS Region
field and the relabelled AWS credential inputs.
- registries-card-detail.png: card close-up with the three action icons and
the metadata row.
- registries-delete-confirm.png: destructive ConfirmModal with the kicker,
title, and stack-pull warning body.
- registries-with-entry.png removed (superseded by registries-overview.png
and registries-card-detail.png).
* docs(auto-update): refresh readiness page for v1 redesign
Bring the Auto-Update Readiness doc in line with the shipped UI:
- Replace the hero screenshot with a fresh capture of the redesigned
board (italic-display hero, brand-cyan accent, per-node groups with
local/remote pills, dashed-border changelog separator).
- Rewrite the card-anatomy list. Drop the rollback-target bullet (the
field exists in the backend payload but is not rendered). Add the
"Rebuild available" inline label and the primary-image / multi-service
count line.
- Rewrite the risk-tags table as a risk-badges table using the actual
badge labels and colors emitted by the UI (Safe / Review / Blocked
with the corresponding icons; Digest rebuild for non-semver tags).
- Add an Empty state section and document the per-node group header.
- Tighten the hero subtitle paragraph to match the actual UI string
(only major-bump count is surfaced separately; preview failures are
not).
- Fix workflow step 4: major-bump apply path is the stack lifecycle
Update action, not the Schedules editor (a scheduled task hits the
same block).
- Add the 2-minute manual-refresh cooldown to the Recheck workflow.
- Remove the broken cross-link to the non-existent
/features/image-update-detection page and inline the 6-hour cadence
fact from ImageUpdateService.INTERVAL_MS.
- Convert troubleshooting to AccordionGroup format per the troubleshoot
ing convention used on /features/deploy-progress.
- Sync the Auto-Update entry in /features/overview.mdx to the new
badge labels and the corrected hero-counter description.
* docs(auto-update): fix Auto-Update entry point in Workflow step 1
Workflow step 1 said "Open the Auto-Update view from the sidebar." The
Auto-Update view is opened from the top nav strip (alongside Home,
Fleet, Resources, App Store, Logs, Schedules, Console, Audit). The
sidebar carries the stack list and the per-stack right-click / kebab
context menu that toggles auto-updates on or off; it does not house
the Auto-Update top-level view.
* docs(auto-update): trim enforcement detail from per-stack control note
State the tier requirement once and stop, per Directive 27. The
"The toggle does not appear on Community" sentence enumerates the
enforcement effect of the gate, which the source already reflects;
operator docs do not need to narrate it.
* docs(auto-heal): refresh page for v1 UI and policy hardening
Rewrite Auto-Heal Policies docs against the current Stack Monitor
sheet: corrects the Max restarts / hr field label, documents the
per-policy enable toggle, the consecutive-failures pill, the full
Recent activity action set (including Docker unavailable), the 30s
evaluation cadence, multi-node behavior, notification dispatches,
and the dashboard Configuration status counter.
Replaces the broken /images/auto-heal-policies/policy-sheet.png
reference with three fresh screenshots captured against a live
node: the sheet on the Auto-heal tab, a single policy row, and
the expanded Recent activity panel.
* docs(webhooks): refresh page for v1 UI, correct tier and add Git source sync
- Fix tier note: gate is Skipper or Admiral, management is admin-only.
- Update Settings path to Settings -> Alerts -> Webhooks; document the
read-only Node field and the green secret-reveal callout.
- Add the missing Git source sync action and the git-pull override value.
- Refresh the configured-webhooks card description: action/stack/node
badges, On/Off toggle, copy URL, and the Recent executions disclosure.
- Tighten the trigger section with a constant-time signature check note
and a status/body/meaning response table.
- Add an Accordion troubleshooting block covering common signature
failures, the 404 case, no-op actions on 202, and git-pull prereqs.
- Re-capture all three screenshots from the v1 UI.
* docs(webhooks): wrap troubleshooting accordions in AccordionGroup
* docs(sidebar): refresh page for v1 redesign with filter chips, bulk mode, row anatomy, and troubleshooting
Rewrites the Stack Sidebar page against the live v1 sidebar and the v1
docs refresh template (Frame screenshots, Note tier callouts,
AccordionGroup troubleshooting). Recaptures all four existing
screenshots and adds three new captures: filter chips, row anatomy,
and bulk mode.
Adds coverage for features the previous page omitted entirely: the
ALL / UP / DOWN / UPDATES filter chips with their counts cap and
collapse toggle; bulk mode (B key, sticky toolbar with Start / Stop /
Restart, and Update gated on Skipper or Admiral); stack-row anatomy
(status pill, label dots with +N overflow, image-update dot vs Git
source icon priority, hover kebab); the Auto-update toggle, Schedule
task, and Open App entries in the context menu; the B shortcut for
bulk mode.
Corrects three claims that no longer matched the code or UI:
Auto-Heal is gated on Skipper or Admiral, not universal; the global
Ctrl+K opens the command palette, not the sidebar search box; the
activity footer kicker reads LIVE / IDLE with the verbatim copy from
SidebarActivityTicker. Documents the in-menu ↗ and L › glyphs as
visual hints rather than global keybindings to match
useStackKeyboardShortcuts.ts.
* docs(sidebar): trim enforcement-effect sentence from context-menu tier note
State the Skipper / Admiral requirement once and stop, per Directive 27.
The "They do not appear in the menu on Community" clause described the
enforcement effect alongside the gate, which the directive bans in
operator-facing docs.
* docs(host-console): refresh page for v1 UI and clarify shell metadata
Rewrite the Host Console page to match the current Cockpit layout
(masthead + terminal well + chip strip), replace the legacy PowerShell
screenshot with a fresh bash capture, and document the masthead tone
states, kicker, metadata pills, and session/heartbeat behavior. Trim
the security section to state the tier and role rule once.
* docs(licensing): refresh page for v1 UI, corrected pricing, and trial flow
Rewrites the Licensing & Billing page to match the redesigned v1
Settings layout. The previous draft still described the legacy
Settings Hub: in-app "Upgrade your plan" Skipper/Admiral cards,
"Start monthly trial" / "Start annual trial" buttons, the
"Have a license key?" field, "Manage Subscription" with a capital S,
"Deactivate License" as the button label, and the license-active.png
asset rendering the literal "Sencho Pro" string in the card title.
None of that exists in the current product.
- Refreshes the Plans table to the live pricing on sencho.io/pricing
and adds an Enterprise mention with the floor price ($3,500/year).
Skipper now $11.99 annual / $14.99 monthly / $449 lifetime, Admiral
now $69.99 annual / $89.99 monthly / $2,499 lifetime.
- Rebuilds the Feature breakdown from a code-level audit of every
requirePaid, requireAdmiral, requireScheduledTaskTier, and
requireTierForSsoProvider call site in backend/src/routes, not
from the marketing page. Notable code-grounded items: CVE
suppressions on Community (no requirePaid guard), manual fleet
snapshots on Community (scheduled snapshots on Skipper),
Sencho Mesh under Admiral (entire mesh.ts router is requireAdmiral),
and scheduled-task tiering names update/scan/snapshot as the
Skipper subset with everything else under Admiral.
- Rewrites the Free trial flow end to end. The previous steps told
operators to click in-app "Start monthly trial" or "Start annual
trial" buttons; no such buttons exist. The new flow starts on
sencho.io/pricing, switches to the Annual or Monthly tab, clicks
"Start 14-day trial" on the Admiral card, completes the Lemon
Squeezy checkout (card-required, no charge before day 14), and
pastes the issued key into Settings -> License -> License key.
- Adds a new "The Plan section" anatomy block describing the masthead
SCOPE / PLAN / DURATION (or RENEWS, TRIAL, STATUS) stat pills and
the Plan card fields (Customer, Product, masked License key, status
helper).
- Adds a new "License states" reference table covering
Community / Trial / Active subscription / Active lifetime /
Expired / Disabled, what each surface renders, and which of the
Plan / Activate / Pricing sections is visible in each state.
- Corrects every UI label that drifted: section heading is Activate,
field label is License key (not "Have a license key?"), buttons are
Manage subscription (lowercase s) and Deactivate (not "Deactivate
License"), and the action-row hint reads "Lemon Squeezy manages
billing".
- Documents the redesigned profile dropdown: identity header with
initials chip, role badge, and tier badge, then Settings,
conditional Billing, Documentation, Feedback, an Appearance
segmented control, and Log Out. Billing only appears when the
license is an active non-lifetime subscription.
- Replaces all four screenshots under docs/images/licensing/:
license-admiral-active.png (production Admiral lifetime view),
profile-menu.png (redesigned popover), and two new captures for
the Community-tier surfaces (license-activate-section.png,
license-community.png). Removes the stale license-active.png
(legacy "Sencho Pro" card) and profile-billing.png (legacy
dropdown).
* docs(settings-reference): refresh page for v1 UI with new sections and masthead
Rewrites docs/reference/settings.mdx against the current Settings Hub so a reader
encounters an accurate map of every section. Adds the previously missing **Cloud
Backup** and **Security** sections, restructures **System Limits** into Host
thresholds and Docker hygiene subsections (GiB units, "Global crash capture"
toggle), fixes the Account password minimum to 8 chars and documents the
two-factor subsection, refreshes License/Routing/Webhooks/App Store with the
field labels actually rendered today, and documents the masthead pills
(SCOPE/NODE, EDITED, plus the per-section stats like 2FA, PLAN, CHANNELS, ROUTES,
WEBHOOKS, LABELS, TRIVY, POLICIES, PROVIDER, USED, SNAPSHOTS, DEV MODE).
Replaces five existing screenshots that predated the v1 redesign and adds five
new captures: Account with the 2FA card, License panel, System Limits with both
subsections, Security with the Trivy installer, and Cloud Backup with Sencho
Cloud Backup provisioned. All shots taken against the production node.
* docs(licensing): drop billing-provider name from operator-facing copy
The previous draft named the third-party billing provider in nine
places (checkout, receipt email, error toast verbatim, Customer /
Product field descriptions, the action-row hint, the billing portal,
and the validation API). Operator docs don't need to advertise which
vendor sits behind the checkout, billing portal, and validation
calls. Rewrite each instance to describe what the operator sees and
does without naming the upstream service.
* docs(node-compatibility): refresh page for v1 UI with lock card visuals and current capability list
- Replaces the legacy "dim + blur + pill overlay" description with the
current CapabilityGate behavior: a centered lock card with an Unplug
icon, title "<feature> is not available on this node", and a body line
that names the node and its running version.
- Corrects the tier-interaction section: on the wrong tier the entry
point is hidden entirely, so the lock card only appears for users who
already cleared the license gate.
- Documents the public /api/meta endpoint, the 5-minute success cache,
the 30-second failure cache, and the lazy-fetch behavior visible in
the switcher (the version pill appears once a node has been visited).
- Refreshes the capability table against the current CapabilityRegistry
list, adding container-exec and vulnerability-scanning, with a note
that vulnerability-scanning is only advertised when Trivy is installed.
- Adds three production screenshots captured on the live fleet:
switcher popover with mixed-version pills (one node on v0.76.9, rest
on v0.81.11), a real lock card on an older pilot agent, and the
Connection Details panel from Settings · Nodes.
* docs(security): refresh security architecture page for v1 UI
Add Fleet Secrets and Webhook signatures cards plus tier-matrix rows for
shipped-but-undocumented features. Rename SSO presets from "one-click" to
"preset providers" (presets still require OAuth-app provisioning on the
upstream IdP). Update settings paths to the v1 middle-dot convention:
Settings · Users, Settings · Account, Settings · Developer · Data retention.
Extend the encryption-at-rest list with registry credentials and Fleet
Secrets bundle payloads (both sealed with the same AES-256-GCM data key)
and clarify the password section with bcrypt cost factor 10.
Add a Webhook signature authentication subsection covering the per-webhook
HMAC-SHA256 secret, one-shot display, masked preview thereafter, and
constant-time comparison on inbound triggers.
Replace the API Tokens screenshot with a fresh capture against the v1
Settings · Identity · API Tokens panel.
* docs(security-advisories): retire reference page
The reference/security-advisories page does not survive the v1 docs
refresh:
- Misuses the term "Security Advisories", which industry-wide refers to
published notices for confirmed product CVEs (ID, severity, affected
versions, fix version, remediation). The retired page was a narrative
changelog of internal hardening work between v0.19 and v0.25.2.
- The narrative is also frozen at v0.25.2 (April 2026) while current
release is v0.81.11. Refreshing it would require backfilling ~56
release entries' worth of hardening copy.
- The framing is uniformly "improved from prior behavior" (minimum 8
characters up from 6, 1-year token expiry previously without expiry,
CORS previously allowed all origins, users should upgrade promptly).
Sencho has not shipped publicly; there are no users to address as
upgraders.
All operationally relevant content already lives elsewhere: the
security architecture page covers the current posture, verifying-images
covers the supply-chain attestations, cve-suppressions covers operator
acknowledgment, vulnerability-scanning covers the in-app scanner, and
contact + the security architecture page both surface the disclosure
path. Published Sencho-product advisories, when any exist, will appear
on the GitHub Security tab, which is already linked from those pages.
Inbound-link audit returned a single hit on the nav entry itself; no
other doc, README, or operator artifact deep-links the slug.
* docs: rewrite Pilot Agent page with deep architecture and operations reference
Reframes docs/features/pilot-agent.mdx as the architecture-and-operations
companion to the operator walkthrough in Multi-Node Management. Adds a
mental model section, an explicit security and trust model, a full agent
env-var reference, an honest limitations list, and a 5-item FAQ. Verifies
every constant and label against the current backend source. Refreshes
four production screenshots (admin login, scrubbed) and resolves the
previously-broken /images/pilot-agent/enrollment-dialog.png reference.
Adjacent edits keep the cross-linking coherent:
- multi-node.mdx adds a one-line forward link to the rewritten page
- security.mdx adds a Pilot Agent tunnel credentials subsection
* docs(fleet-federation): deep rewrite with production screenshots
Rewrites the Fleet Federation page against the v1 docs refresh template
following the recent fleet-view, pilot-agent, and multi-node refreshes.
Doubles the page length (92 to 204 lines) while keeping the cut-line v1
MVP scope: operator-driven placement controls (cordon + pin) for
Blueprints, no expansion into mesh/sync/pilot territory.
What changed:
- Adds four production-captured screenshots under docs/images/fleet-federation/:
the Federation tab with a cordoned node populated, the node-card kebab
menu showing the Cordon node entry, the cordon confirmation dialog
with a reason filled in, and a node card displaying the Cordoned pill.
- Expands the page to eleven sections: opening summary, philosophy
(kept), key capabilities, prerequisites, step-by-step usage with
embedded screenshots, behaviour and lifecycle table, security and
audit, limitations and non-goals (expanded), practical workflows (new:
OS patching, host-to-host migration, gateway pinning), troubleshooting
(eight accordions, up from five), and a Where Federation fits
cross-link table.
- Documents the exact production UI strings observed: the cordon
dialog description, the uncordon confirmation copy, the reason field
cap (256 chars), and the audit log action names (node.cordon,
node.uncordon, blueprint.pin).
- Documents the audit visibility surface so operators know how to
filter the Audit view for cordon and pin history.
- Adds eight cross-links to sibling pages (Fleet View, Multi-Node,
Pilot Agent, Mesh, Fleet Actions, Fleet Sync, Blueprints, Licensing)
with one-line scope contrasts so newcomers can place Federation in
the broader fleet picture.
- Tightens lifecycle table to operator-relevant terms (no DB column
names) and audit section to operator-facing wording (no middleware
names), keeping the page operator-focused rather than
implementation-focused.
Validation:
- Captured screenshots against the production node logged in as admin,
using Playwright MCP. Cordoned and pinned actions reverted; audit log
confirmed the matched cordon/uncordon pair.
- Verified every cross-link target exists in the v1-refresh worktree
(/features/fleet-view, /features/multi-node, /features/pilot-agent,
/features/sencho-mesh, /features/fleet-actions, /features/fleet-sync,
/features/blueprint-model, /features/licensing).
- Compliance: no em dashes, no PII, no "previously"/"used to" framing,
no fence-spec language, tier rule stated once in plain language.
* docs(fleet-federation): drop fence-spec phrasing in the open-core context
Sencho is open-core: anyone can clone the repo and read the tier gate.
Operator docs that name exactly where the UI gate sits ("hidden at the
Community and Skipper tiers", "lower-tier users do not see the toggle",
"only the Federation tab is gated") work as a dig-target for a
tech-savvy reader and undercut the open-core posture. Directive 27
already bans enforcement-chain spelling; the open-core threat model
makes the same phrasings risky even when they describe UI surfaces
rather than route guards.
Removes three instances of the pattern on this page:
- Top Note callout: drops "The tab is hidden at the Community and
Skipper tiers." Keeps the one-line requirement: "Federation is an
Admiral feature. Cordon and pin actions require an admin user role."
- Security and audit section: drops the sentence enumerating which UI
affordances are hidden from which tiers. Keeps the customer-visible
behavior (the Cordoned pill stays visible at every tier as a
read-only signal).
- Troubleshooting "Federation tab is not visible" accordion: rewrites
to lead with the requirement and the role check, drops the
"Federation is hidden by design" and "only the toggle and the
Federation tab are gated" phrasings.
Other claims on the page unchanged; rule is still stated once in plain
language at the top of the page.
* docs(fleet-sync): deep rewrite with production screenshots
Replace fleet-sync.mdx with a verified end-to-end reference. The previous
page named two replicated resources but the code syncs three, described a
sync-status panel and a fleet-vs-node scope picker that do not exist in
the shipped UI, and was missing prerequisites and several edge cases.
Highlights of the rewrite:
- Names all three replicated resources (scan policies, CVE suppressions,
misconfig acknowledgements) and treats them uniformly.
- Drops the sync-status-panel and node-scope-picker UI claims; both move
to the Limitations section as honest caveats.
- Adds prerequisites covering the paid-tier requirement on the control,
admin-role requirement, proxy-mode remotes, and reachability.
- Expands lifecycle coverage: per-node serialised pushes, add-node
backfill, monotonic pushedAt, per-resource watermarks, identity-drift
notifications, the 5000-row truncation cap, stale-target warnings,
audit-log entries on the replica.
- New "Where Fleet Sync fits" closing table cross-linking to Fleet View,
Multi-Node Management, Pilot Agent, Vulnerability Scanning, CVE
Suppressions, Fleet Federation, Fleet Actions, and Licensing.
- Two fresh production screenshots: control Security panel and the
"Scanner is per-node" callout shown when proxying to a remote.
* docs(fleet-actions): deep rewrite with production screenshots
Three cards are documented end to end: Stop fleet by label, Bulk label
assign, and Prune Docker resources fleet-wide. Adds the execution-path
distinction (control-orchestrated fan-out vs single-node proxy), per-card
behaviour and partial-failure semantics, prerequisites, limitations,
practical workflows, an Accordion troubleshooting section, and a Where
Fleet Actions fits comparison table linking the surrounding Fleet view
features.
Corrects the prior page's tab-neighborhood claim, confirm-dialog wording,
autocomplete-vs-request scope, and missing batch ceiling. Replaces the
ten-day-old single screenshot with five fresh production captures under
docs/images/fleet-actions/.
* docs(fleet-secrets): deep rewrite with production screenshots
Full rewrite of /features/fleet-secrets matching the fleet-actions
structure. Replaces the sparse v1 page (no Frames, inline Q&A) with a
gold-standard layout: opening Frame, single Note for the tier gate,
'What it covers' table, mental model, prerequisites, create + edit +
versions + push (Target / Preview / Results) sections each with a
production Frame, Import from stack section, behaviour and lifecycle
table, audit-trail mapping with the six exact audit strings,
limitations and non-goals, practical workflows, AccordionGroup
troubleshooting, and a Where-it-fits cross-link table.
Adds six fresh production screenshots under
docs/images/fleet-secrets/ : overview, create, versions, target,
preview, and results.
Documents the Import-from-stack flow (depends on the bundle editor's
new Import action) and uses the post-rename 'Send' wording on the
bundle-row action (depends on the aria-label fix).
Corrects three factual drifts vs the code: env-key regex described as
'letter or underscore, then letters/digits/underscores; case-
sensitive' to match ^[A-Za-z_][A-Za-z0-9_]*$ ; documents only the
'ok' and 'failed' status pills (the 'skipped' enum value is unused);
replaces the bogus 'stack not found' troubleshooting entry with the
real 'env file not declared' cause.
Drops the fence-spec phrasing 'The tab is hidden on Community.' per
Directive 31; the tier requirement is now stated once in plain
language.
* docs(sencho-mesh): deep rewrite with mental model, lifecycle, security, screenshots
Replace the feature-reference page with a deep product + technical guide.
Adds:
- Opening hook framing audience and problem (cross-node service-to-service
without a separate VPN or service-mesh sidecar).
- Mental model: three moving parts (sencho_mesh bridge, alias registry,
cross-node transport) with direction-of-flow described in prose.
- Key capabilities, prerequisites, step-by-step usage with inline screenshots.
- Full lifecycle section covering opt-in, opt-out, sticky stack-stopped state,
peer reconnect, and the proxy-mode bridge with its real default (persistent,
env-override for idle).
- Security and trust boundaries split into authentication, inbound exposure,
encryption, audit, and app-layer caveats.
- Limitations and non-goals: one-alias-per-port, port 1852 reserved,
central-relay for remote-to-remote, shared 1024-stream pool with the Pilot
tunnel, no L7, host-network unsupported, in-memory activity log.
- Three concrete workflow examples and a complete troubleshooting accordion
(every data-plane reason, every probe stage, every unreachable cause) plus
a Common questions FAQ.
- Where Mesh fits CardGroup linking Pilot Agent, Multi-Node, Federation,
Licensing.
Corrections vs prior text:
- Tab is labelled Traffic in the UI (not Routing); all navigation references
updated.
- Proxy-mode bridge default is no idle close (env-overridable to opt into idle
teardown); prior 5-minute-teardown claim removed.
- Audit trail scope tightened: only opt-in / opt-out write durable rows;
tunnel-state and probe events live in the in-memory activity log.
Adds seven production screenshots under docs/images/sencho-mesh covering
Table view, opt-in sheet, graph (Tunnels and Aliases edge modes), Diagnostics,
activity log, and per-stack topology.
* docs(blueprints): add missing detail-state-review screenshot
Captures the Blueprint detail sheet with a deployment row in the
"Awaiting confirmation" status (stateful first-deploy gate), to fix the
broken image referenced at blueprint-model.mdx:132. mint broken-links
now reports zero broken references.
* docs(blueprints): deep rewrite with mental model, lifecycle, security, prerequisites
Restructures the Blueprints page against the v1-refresh template used by the
recently-refreshed mesh, secrets, and atomic-deployments pages. Adds a mental
model, prerequisites table, lifecycle and status-transition map, security and
trust boundaries section, practical workflows, common questions accordion,
and a Where Blueprints fits CardGroup. Removes the internal-style rollout
and watch-plan section. Replaces all nine production screenshots with fresh
captures against the production node signed in as admin, and adds two new
captures (federation pin policy table, stateless eviction dialog). Rewrites
the tier-gate Note to drop the fence-spec phrasing that violated Directive
31. Every retained claim is anchored to current backend or frontend code.
* docs(pilot-agent): recapture enrollment dialog with compose payload
Replaces the pre-0.84 docker-run capture with the current dialog (Compose
file, two-step instructions, "Copy compose file" button) and refines the
alt text to describe the captured content. URL and token redacted to
placeholder values during capture.
@@ -118,7 +118,6 @@
|
||||
"features/global-search",
|
||||
"features/global-observability",
|
||||
"features/alerts-notifications",
|
||||
"features/notification-routing",
|
||||
"features/audit-log"
|
||||
]
|
||||
},
|
||||
@@ -174,7 +173,6 @@
|
||||
"features/licensing",
|
||||
"features/node-compatibility",
|
||||
"security",
|
||||
"reference/security-advisories",
|
||||
"reference/contact"
|
||||
]
|
||||
},
|
||||
|
||||
@@ -1,275 +1,424 @@
|
||||
---
|
||||
title: Alerts & Notifications
|
||||
description: Set threshold-based alerts on container metrics and route them to Discord, Slack, or any webhook.
|
||||
description: Threshold and event alerts for your fleet, dispatched to Discord, Slack, or any webhook, with per-stack rules and Admiral routing.
|
||||
---
|
||||
|
||||
Sencho can watch your containers for resource anomalies and notify you when thresholds are breached. Alerts are defined per stack, and notifications are delivered through external channels you configure.
|
||||
|
||||
## Setting up notification channels
|
||||
|
||||
At least one channel must be enabled before alerts can be delivered externally. Go to **Settings > Notifications**.
|
||||
Sencho watches each node it manages for container crashes, host pressure, scheduled-task results, and update availability, then surfaces every signal in two places: the in-app notification bell at the top of the shell and one of three external channels you configure. This page covers everything from configuring channels to writing per-stack threshold rules, routing alerts to dedicated channels with Admiral routing rules, and tuning retention.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/alerts-notifications/notifications-settings.png" alt="Notifications & Alerts settings showing Discord, Slack, and Webhook tabs" />
|
||||
<img src="/images/alerts-notifications/notifications-settings.png" alt="Settings · Notifications panel showing the Discord, Slack, and Webhook tabs with the masthead breadcrumb, the CHANNELS 3/3 stat, the active Discord tab with its Enabled toggle on, the Webhook URL input, and the Test and Save actions." />
|
||||
</Frame>
|
||||
|
||||
Three channel types are available, each configured with a webhook URL and an enable/disable toggle:
|
||||
## Notification channels
|
||||
|
||||
Open **Settings · Notifications** to configure the three channel types. Each channel is per-node, so switching the active node via the node picker reloads the panel against that node's stored settings. The masthead carries a `CHANNELS` stat showing how many of the three slots are enabled.
|
||||
|
||||
Each tab carries the same controls: an **Enabled** toggle (helper: `Send Sencho events to this <name> channel.`), a **Webhook URL** input (placeholder `https://...`, helper: `Sencho posts JSON payloads here. Use a private channel.`), and the **Test** and **Save** buttons. The kicker on each tab toggles between `enabled` and `off` so you can see at a glance which slots are wired up.
|
||||
|
||||
<Note>
|
||||
Webhook URLs must use HTTPS. Sencho rejects `http://` and any string that does not parse as a URL. The same rule applies to test sends and to routing-rule URLs.
|
||||
</Note>
|
||||
|
||||
### Discord
|
||||
|
||||
1. In Discord, go to your server's **Settings > Integrations > Webhooks**
|
||||
2. Click **New Webhook**, choose a channel, and copy the webhook URL
|
||||
3. In Sencho, open **Settings > Notifications > Discord**, paste the URL, enable the toggle, and click **Save**
|
||||
4. Click **Test** to send a test message and verify the connection
|
||||
Paste an incoming webhook URL from your Discord channel's **Edit Channel · Integrations · Webhooks** view. Sencho posts a single embed per alert, with the title `Sencho Alert [<LEVEL>]`, the message in the description, and the embed color set by severity (info blue, warning yellow, error red).
|
||||
|
||||
### Slack
|
||||
|
||||
1. In Slack, go to **api.slack.com/apps**, create an app, and add the **Incoming Webhooks** feature
|
||||
2. Activate it and copy the generated webhook URL for your chosen channel
|
||||
3. In Sencho, open **Settings > Notifications > Slack**, paste the URL, enable the toggle, and click **Save**
|
||||
Paste an incoming webhook URL from a Slack app installed in your workspace (URL shape `https://hooks.slack.com/services/...`). Sencho posts a single text message: `<emoji> *Sencho Alert [<LEVEL>]*` followed by the message body, with the emoji set per severity (`ℹ️` info, `⚠️` warning, `🚨` error).
|
||||
|
||||
### Generic Webhook
|
||||
### Generic webhook
|
||||
|
||||
Any HTTPS endpoint that accepts a POST with a JSON body can receive Sencho alerts. Go to **Settings > Notifications > Webhook**, enter the URL, enable the toggle, and click **Save**.
|
||||
|
||||
<Note>
|
||||
All webhook URLs (Discord, Slack, and generic) must use HTTPS. HTTP URLs are rejected.
|
||||
</Note>
|
||||
|
||||
The payload format is:
|
||||
Use the Generic Webhook tab when you have your own receiver: a Mattermost or Teams adapter, a serverless function that fans out to email or SMS, or a logging endpoint. Sencho posts JSON of the shape:
|
||||
|
||||
```json
|
||||
{
|
||||
"level": "warning",
|
||||
"message": "cpu_percent exceeded 90% for 5 minutes on stack my-app",
|
||||
"timestamp": "2026-03-22T10:00:00.000Z"
|
||||
"message": "[Node: Local] The cpu_percent for plex has exceeded your threshold of 80% (Currently: 91%).",
|
||||
"timestamp": "2026-05-08T22:14:09.812Z",
|
||||
"source": "sencho"
|
||||
}
|
||||
```
|
||||
|
||||
## Creating stack alerts
|
||||
`level` is one of `info`, `warning`, `error`. `source` is always the literal string `sencho`. `timestamp` is ISO-8601 with millisecond precision.
|
||||
|
||||
Stack alerts are configured per stack. Right-click a stack in the sidebar (or click the three-dot menu) and select **Alerts**. A panel slides open showing existing rules and a form to create new ones.
|
||||
### Test sends and delivery semantics
|
||||
|
||||
The **Test** button on each tab dispatches the literal message `🔌 Test Notification from Sencho!` at level `info` through the same path a real alert would take. Test sends require the admin role; the server returns 403 if an operator or viewer submits one.
|
||||
|
||||
Each dispatch is a single-shot HTTP POST with a 10-second `AbortSignal.timeout`. There is no retry queue. If your endpoint is down at the moment of the dispatch, the alert is recorded in the bell with an internal `dispatch_error` field set and is not redelivered.
|
||||
|
||||
## Notification Routing
|
||||
|
||||
<Note>
|
||||
Notification Routing requires a **Sencho Admiral** license. Admin role is required to create, edit, or delete routes.
|
||||
</Note>
|
||||
|
||||
Routing lets you direct alerts that match specific criteria to dedicated channels. Production crashes can land in `#prod-incidents` on Slack while staging notifications go to a less urgent Discord channel, all without juggling per-channel webhook URLs across teams.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/alerts-notifications/alert-panel.png" alt="Stack Alerts panel showing the notification status banner and alert rule form" />
|
||||
<img src="/images/alerts-notifications/routing-list.png" alt="Settings · Routing card list showing two rules ('Critical to Slack' and 'Production alerts'), each with a Discord badge, an ON pill, the 'Matches all alerts' summary line, a truncated channel URL, a Priority chip, and lightning-test, edit, and delete icon actions on the right." />
|
||||
</Frame>
|
||||
|
||||
The panel includes:
|
||||
### How routing fits into dispatch
|
||||
|
||||
- **Notification status banner** at the top, showing whether channels are configured. If no channels are enabled, a warning explains that rules will be evaluated but no external notifications will be sent.
|
||||
- **Existing Rules** section listing all active rules for this stack, each showing the metric, condition, duration, and cooldown. Hover over a rule to reveal the delete button.
|
||||
- **Add New Rule** form with the fields described below.
|
||||
For every alert Sencho dispatches, the routing engine evaluates every enabled route. A route matches when **all of its non-empty matchers** match the alert: the **Node**, **Stacks**, **Labels**, and **Categories** filters compose with AND. An empty matcher is treated as match-anything.
|
||||
|
||||
If at least one route matches, every matching route fires and the **global channels are skipped** for that alert. If zero routes match, the alert falls back to the global channels configured in the previous section.
|
||||
|
||||
A route with all four matchers left empty matches every alert and intercepts global delivery entirely. Stack-less alerts (host CPU, RAM, disk, fleet-sync warnings, and scheduled-task results without a stack target) almost always miss any populated `Stacks` filter, so they fall back to global by default.
|
||||
|
||||
### Creating a routing rule
|
||||
|
||||
Open **Settings · Routing** and click **+ Add Route**.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/alerts-notifications/routing-modal.png" alt="The 'New routing rule' modal with the 'ROUTING · NEW RULE' kicker, a Name input filled with 'Production alerts', a Node scope select reading 'Any node', empty Stacks, Labels, and Categories combobox pickers, the helper line 'Leave blank to match all categories. All non-empty filters must match (AND).', a Channel tab strip with Discord selected and a webhook URL filled in, and Priority and Enabled fields with Cancel and CREATE actions in the footer." />
|
||||
</Frame>
|
||||
|
||||
| Field | Purpose |
|
||||
|-------|---------|
|
||||
| **Name** | A human label, up to 100 characters. Shown on the rule card. |
|
||||
| **Node scope** | Either `Any node` or a specific node. When set to a node, the rule only matches alerts originating from that node. |
|
||||
| **Stacks** *(optional)* | A combobox of stacks on the active node. Selected stacks appear as removable pills. Empty matches any stack. |
|
||||
| **Labels** *(optional)* | A combobox of stack labels on the active node. Empty matches any label. |
|
||||
| **Categories** *(optional)* | A combobox of notification categories. The helper line reads `Leave blank to match all categories. All non-empty filters must match (AND).` |
|
||||
| **Channel** | Tabs for Discord, Slack, and Webhook with a URL input below. URL must use HTTPS. |
|
||||
| **Priority** | A number used to sort the rule list. Lower numbers appear higher up. Priority does not gate dispatch: when multiple rules match the same alert, every matching rule fires concurrently. |
|
||||
| **Enabled** | Toggle the rule on or off without deleting it. |
|
||||
|
||||
The modal kicker reads `ROUTING · NEW RULE` when adding and `ROUTING · EDIT RULE` when modifying.
|
||||
|
||||
### Managing rules
|
||||
|
||||
Each rule renders as a card on the Routing page with the rule name, channel-type badge, an `ON` / `OFF` pill, then a row of small badges naming each matcher: one mono badge per **Stack**, one outline badge per **Label**, one outline mono badge per **Category** label. When all matchers are empty, a single muted `Matches all alerts` line replaces the badge row. After the badges, a vertical bar separator is followed by the truncated channel URL, then a second separator and a `Priority: N` chip when priority is non-zero.
|
||||
|
||||
Three icon actions appear on the right edge of each card:
|
||||
|
||||
- **Lightning** sends a test message through the rule's channel using the same payload shape as the global Test button.
|
||||
- **Pencil** opens the edit modal pre-filled with the current values.
|
||||
- **Trash** deletes the rule after a confirmation dialog.
|
||||
|
||||
The masthead carries `SCOPE` (`global`), `ROUTES` (total), and `ENABLED` (count).
|
||||
|
||||
## Notification categories
|
||||
|
||||
Every alert Sencho dispatches carries a category that you can filter on in the bell, target with a routing rule's **Categories** matcher, or reason about when reading audit history.
|
||||
|
||||
| Category | Label in the bell | Trigger |
|
||||
|----------|-------------------|---------|
|
||||
| `deploy_success` | Deploy success | Stack deploy completed without error |
|
||||
| `deploy_failure` | Deploy failure | Stack deploy returned a non-zero exit |
|
||||
| `stack_started` | Stack started | Stack started via the dashboard or API |
|
||||
| `stack_stopped` | Stack stopped | Stack stopped via the dashboard or API |
|
||||
| `stack_restarted` | Stack restarted | Stack restarted via the dashboard or API |
|
||||
| `image_update_available` | Update available | Image-update poll found a newer digest |
|
||||
| `image_update_applied` | Update applied | Manual or scheduled auto-update applied new images |
|
||||
| `autoheal_triggered` | Auto-heal | Auto-heal restarted, failed to restart, or auto-disabled a policy |
|
||||
| `monitor_alert` | Monitor alert | Per-stack threshold breach, host CPU/RAM/disk warning, or healthcheck failure |
|
||||
| `scan_finding` | Scan finding | Vulnerability-scan completion, per-violation alert, post-deploy scan failure, or auto-update gate block |
|
||||
| `system` | System | Sencho version, Trivy auto-update, fleet sync, daemon connectivity, scheduled-task lifecycle, cloud-backup upload failure |
|
||||
| `blueprint_deployed` | `blueprint_deployed` | Blueprint provisioned a new deployment (Admiral) |
|
||||
| `blueprint_deployment_failed` | `blueprint_deployment_failed` | Blueprint deployment errored out (Admiral) |
|
||||
| `blueprint_drift_detected` | `blueprint_drift_detected` | Blueprint drift detected in `suggest` or `enforce` mode (Admiral) |
|
||||
| `blueprint_drift_correction_failed` | `blueprint_drift_correction_failed` | Blueprint enforce-mode redeploy failed (Admiral) |
|
||||
|
||||
The four `blueprint_*` categories are accepted by routing rules but render as raw category strings in the bell because the frontend label map omits them.
|
||||
|
||||
## Per-stack alert rules
|
||||
|
||||
Each stack carries its own set of threshold rules that fire when a metric stays above (or below) a value for a configurable window. Rules live on the node where the stack runs and are evaluated locally on a 30-second tick.
|
||||
|
||||
Open the rules editor by right-clicking a stack in the sidebar and choosing **Alerts**, or by pressing `A` while focused on a stack.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/alerts-notifications/alert-panel.png" alt="The Stack › PLEX › MONITOR sheet with the Alerts tab active, a green 'Notifications active via Discord, Slack, Webhook' banner under the NOTIFICATION CHANNELS heading, an ACTIVE RULES section listing 'CPU Usage (%) > 80' with the secondary line 'Trigger after 5m • Cooldown 60m', and an ADD NEW RULE form with Metric, Operator, Threshold, Duration, Cooldown fields and an Add Rule submit button." />
|
||||
</Frame>
|
||||
|
||||
### Alert fields
|
||||
|
||||
| Field | Description |
|
||||
|-------|-------------|
|
||||
| **Metric** | The container metric to watch (see table below) |
|
||||
| **Operator** | Comparison: Greater than, Greater or equal, Less than, Less or equal, Equals |
|
||||
| **Threshold** | The numerical value to compare against |
|
||||
| **Duration (mins)** | How long the condition must hold before firing (default: 5) |
|
||||
| **Cooldown (mins)** | Minimum time between repeated notifications for this rule (default: 60) |
|
||||
| Field | Purpose |
|
||||
|-------|---------|
|
||||
| **Metric** | The system resource or metric to monitor. |
|
||||
| **Operator** | Comparison: `Greater than`, `Greater or eq`, `Less than`, `Less or eq`, `Equals`. |
|
||||
| **Threshold** | A number ≥ 0. The unit follows the chosen metric. |
|
||||
| **Duration (mins)** | How long the breach must persist before firing. Default `5`, range 0 to 1440. |
|
||||
| **Cooldown (mins)** | Silence window between fires after a rule triggers. Default `60`, range 0 to 10080. |
|
||||
|
||||
Sencho tracks the start of each breach in memory; the rule fires only after the breach has lasted for the full **Duration**, and only if the previous fire is older than **Cooldown**.
|
||||
|
||||
### Available metrics
|
||||
|
||||
| Metric | Description |
|
||||
|--------|-------------|
|
||||
| CPU Usage (%) | CPU usage relative to total host cores |
|
||||
| Memory Usage (%) | Memory used as a fraction of the host total |
|
||||
| Memory Usage (MB) | RSS memory used by the container |
|
||||
| Network In (MB) | Cumulative inbound network bytes (in MB) |
|
||||
| Network Out (MB) | Cumulative outbound network bytes (in MB) |
|
||||
| Restart Count | Number of times the container has restarted |
|
||||
| Metric | UI label | Unit |
|
||||
|--------|----------|------|
|
||||
| `cpu_percent` | CPU Usage (%) | percent of all cores |
|
||||
| `memory_percent` | Memory Usage (%) | percent of container memory limit |
|
||||
| `memory_mb` | Memory Usage (MB) | MB of resident set size, minus cache |
|
||||
| `net_rx` | Network In (MB/s) | MB per second, computed as a delta between consecutive samples |
|
||||
| `net_tx` | Network Out (MB/s) | MB per second, computed as a delta between consecutive samples |
|
||||
| `restart_count` | Restart Count | integer count of Docker-reported restarts |
|
||||
|
||||
### Channel banner states
|
||||
|
||||
The **NOTIFICATION CHANNELS** banner above the rules list reflects what dispatch will look like for this stack:
|
||||
|
||||
- **Loading** is a spinner with `Checking notification channels...` while Sencho asks the target node for its agent state.
|
||||
- **Remote node** is a blue banner reading `Remote node: <name>`, with the body `Alert rules are stored and evaluated on this remote instance. Notifications are dispatched using that node's configured channels.` A sub-line reports whether the remote has any channels configured.
|
||||
- **No channels** is an amber banner reading `No notification channels configured`, body `Alert rules will be saved and evaluated, but no notifications will be dispatched. Configure Discord, Slack, or a webhook in Settings → Notifications.`
|
||||
- **Active** is a green banner reading `Notifications active via Discord, Slack, …` with the configured channels listed.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/alerts-notifications/alert-panel-remote-banner.png" alt="The Stack › SAELIX-DB › MONITOR sheet on a remote node, with the blue 'Remote node: node-a' banner explaining that rules are stored and evaluated on the remote, plus a follow-up amber line noting that no notification channels are configured on that remote." />
|
||||
</Frame>
|
||||
|
||||
### Example: alert on high CPU
|
||||
|
||||
To alert when any container in a stack uses more than 80% CPU for over 5 consecutive minutes, with a 60-minute cooldown:
|
||||
To page when CPU on a stack stays above 80% for at least five minutes, with no more than one alert per hour:
|
||||
|
||||
| Field | Value |
|
||||
|-------|-------|
|
||||
| Metric | CPU Usage (%) |
|
||||
| Operator | Greater than |
|
||||
| Threshold | 80 |
|
||||
| Duration | 5 |
|
||||
| Cooldown | 60 |
|
||||
| **Metric** | CPU Usage (%) |
|
||||
| **Operator** | Greater than |
|
||||
| **Threshold** | `80` |
|
||||
| **Duration (mins)** | `5` |
|
||||
| **Cooldown (mins)** | `60` |
|
||||
|
||||
## Notification history
|
||||
### Permissions and validation
|
||||
|
||||
All dispatched notifications appear in the **notification bell** in the top-right corner of the navigation bar. A red dot pulses on the bell when unread notifications exist.
|
||||
Operators and viewers see existing rules read-only; only admins see the add and delete affordances. Submitting an empty threshold surfaces `Please enter a threshold.` Successful saves toast `Alert rule added.` and `Alert rule deleted.` A failed POST surfaces `Network error. Could not reach the node.`
|
||||
|
||||
The delete confirmation dialog reads `Delete Alert Rule` / `This will permanently remove this alert rule. Notifications for this condition will no longer be sent.` with a destructive **Delete** button.
|
||||
|
||||
## Notification history (the bell)
|
||||
|
||||
The bell icon at the top of the shell is the live feed of every alert across the fleet. A pulsing red dot surfaces unread items.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/alerts-notifications/notification-popover.png" alt="Notification popover showing recent alert entries with level badges and timestamps" />
|
||||
<img src="/images/alerts-notifications/notification-popover.png" alt="The notification bell open, with a 'Notifications' italic title, '9 UNREAD' uppercase mono caption, an All / Unread / Alerts segmented control, the filter-toggle, mark-all-read, and clear-all icon actions, a 'TODAY' day-band header, and a stack of severity-tinted rows describing Sencho version updates and a recent stack stop, each with a node-name pill, a relative timestamp, and a per-row dismiss target." />
|
||||
</Frame>
|
||||
|
||||
Click the bell to open the notification popover. Each entry shows:
|
||||
### Anatomy
|
||||
|
||||
- **Level badge** (ERROR in red, WARNING in amber, INFO in default)
|
||||
- **Node name** badge for notifications from remote nodes
|
||||
- **Timestamp** of when the alert was triggered
|
||||
- **Alert message** describing what was breached
|
||||
A title bar with `Notifications` in the italic display face, the unread count to the right (`<n> UNREAD`), and a row of controls below.
|
||||
|
||||
A toolbar above the list provides filtering and bulk actions:
|
||||
The segmented control switches between `All`, `Unread` (with badge), and `Alerts` (rows where level is `warning` or `error`). To the right, three icon buttons:
|
||||
|
||||
| Control | What it does |
|
||||
|---------|--------------|
|
||||
| **All / Unread / Alerts** | Switches the list between every notification, unread only, or alert-level only |
|
||||
| **Filter toggle** | Reveals dropdowns to filter by node and by notification type; an accent dot on the icon indicates an active filter |
|
||||
| **Mark all as read** | Marks every notification as read (removes the red dot) |
|
||||
| **Clear all** | Deletes every notification from the list |
|
||||
- **SlidersHorizontal** toggles a hidden filter row. A small dot on the icon means at least one filter is active.
|
||||
- **CheckCheck** is `Mark all read`.
|
||||
- **Trash2** is `Clear all`. The action issues `DELETE /api/notifications` against every node that contributed a row.
|
||||
|
||||
You can also dismiss individual notifications by hovering over them and clicking the dismiss button.
|
||||
The hidden filter row has two combobox dropdowns: **Filter by node** (only rendered with two or more nodes registered) and **Filter by category**, listing all the user-facing labels from the categories table above.
|
||||
|
||||
<Note>
|
||||
Notifications only reach external channels (Discord, Slack, Webhook) if at least one channel is enabled. Dashboard notifications appear regardless.
|
||||
</Note>
|
||||
<Frame>
|
||||
<img src="/images/alerts-notifications/bell-filter-row.png" alt="The bell with its filter row expanded, the 'ALL TYPES' combobox open over a long popover listing every notification category from 'All types' through 'System', and the underlying notification rows visible beneath." />
|
||||
</Frame>
|
||||
|
||||
## Alerts on remote nodes
|
||||
### Row anatomy
|
||||
|
||||
Alerts work the same way on remote nodes as they do locally. When you switch to a remote node and open a stack's alerts panel, you are managing rules on that remote instance.
|
||||
Each row carries:
|
||||
|
||||
Key details:
|
||||
- A **3 px left rail** colored by severity: brand cyan for `info`, amber for `warning`, destructive rose for `error`. Rail saturation drops to about 30% once the row is read.
|
||||
- A **severity icon** in the same hue: `Info`, `AlertTriangle`, or `AlertOctagon`.
|
||||
- The **message body**, slightly dimmer once read.
|
||||
- A **kicker line** below the message with an optional **node-name pill** (only rendered for multi-node fleets), a `·` separator, and a relative timestamp (`just now`, `Nm ago`, `Nh ago`, `yesterday`, then a short locale date).
|
||||
- A **dismiss** target on hover, top-right.
|
||||
|
||||
- **Alert rules are stored on each node independently.** Rules created while a remote node is selected are saved on that remote instance, not your primary instance.
|
||||
- **Monitoring runs locally on each node.** Each Sencho instance evaluates its own alert rules against its own container metrics.
|
||||
- **Notifications are sent by the node that detects the breach.** Make sure notification channels are configured on each remote node where you want to receive alerts, since channel settings are per-instance.
|
||||
There is no level chip in the row; the rail color and icon do that work. Rows whose payload carries a `stack_name` become click targets that jump to the stack and, when a `container_name` is present, surface that container's logs.
|
||||
|
||||
The alert panel shows a blue info banner when you are configuring alerts on a remote node, including which channels are active on that node.
|
||||
### Day groupings
|
||||
|
||||
### Setup checklist for remote alerts
|
||||
Rows are grouped under uppercase day-band headers: `TODAY`, `YESTERDAY`, `THIS WEEK`, `EARLIER`.
|
||||
|
||||
You can configure a remote node's notification channels directly from the control plane:
|
||||
### Empty states
|
||||
|
||||
1. From your **primary** instance, switch to the remote node using the node picker
|
||||
2. Open **Settings > Notifications** and configure the remote's Discord, Slack, or Webhook channels. Values saved here apply only to the selected node.
|
||||
3. Right-click a stack and select **Alerts** to create rules
|
||||
4. The remote instance handles monitoring and notification delivery independently
|
||||
<Frame>
|
||||
<img src="/images/alerts-notifications/bell-empty.png" alt="The bell open in its empty state with a strikethrough bell icon, the headline 'You're all caught up', and the mono caption 'New notifications appear here in real time.'" />
|
||||
</Frame>
|
||||
|
||||
## Update availability notifications
|
||||
The popover renders three different empty states:
|
||||
|
||||
Sencho can notify you when software updates are available, both for Sencho itself and for your stack images.
|
||||
- All filters, no rows ever: `You're all caught up` over `New notifications appear here in real time.`
|
||||
- Unread filter, all read: `No unread notifications` over `Everything in your feed has been read.`
|
||||
- Alerts filter, no warnings or errors: `No active alerts` over `Warnings and errors will surface here when they occur.`
|
||||
|
||||
### Sencho version updates
|
||||
### What you won't see in the bell
|
||||
|
||||
When a newer version of Sencho is published, an informational notification is dispatched through your configured channels. Each Sencho instance runs its own version check roughly once every 6 hours and notifies exactly once per new version, so remote nodes self-report their own availability alerts. After you update, the cycle resets and you will be notified when the next release becomes available.
|
||||
Sencho deliberately suppresses one class of notification from the popover: rows where the category is one of `deploy_success`, `stack_started`, `stack_stopped`, `stack_restarted`, or `image_update_applied`, AND the row carries an `actor_username` other than `system`. These are confirmations of an action you just clicked, surfaced as a toast at the moment of the action; replaying them in the bell would be noise. The rows are still persisted to `notification_history` and still dispatched to global channels and matching routes; only the bell render hides them.
|
||||
|
||||
The notification message includes the version you are running and the version that was detected, for example:
|
||||
### Limits
|
||||
|
||||
```
|
||||
Sencho 0.47.0 is available (currently running 0.46.16). Visit the Fleet dashboard to update.
|
||||
```
|
||||
The popover holds the latest 50 rows per node per fetch. The backend caps `notification_history` at 100 rows per node, regardless of the user-tunable retention setting; the cap evicts the oldest rows on every insert. There is no `Load more` affordance; older rows roll off the cap.
|
||||
|
||||
### Stack image updates
|
||||
## Cross-node delivery
|
||||
|
||||
When the periodic image check (every 6 hours) detects that a stack has new upstream images available, a notification is dispatched for each affected stack. Notifications are sent only on state transitions: you will be notified once when an update first appears, not on every check cycle. After you update the stack, the status resets so a future release can notify again.
|
||||
The bell aggregates across the entire fleet by hitting `/api/notifications` against every registered node and opening one WebSocket per node: `/ws/notifications` for the local instance and `/ws/notifications?nodeId=<id>` for every remote. A new alert on any remote arrives in the bell within about a second; a 60-second safety-net poll catches anything the WebSocket dropped.
|
||||
|
||||
Both notification types use the same channel routing as alerts: if notification routes are configured for a stack, those channels receive the message; otherwise, global notification channels are used as a fallback.
|
||||
Per-stack alert rules and channel configuration are **stored on the node where the stack runs**. To configure a rule on a remote, switch the active node via the picker, open the stack's **Monitor** sheet, and the form's POST is forwarded to the remote.
|
||||
|
||||
## Scheduled scan completion
|
||||
Crash detection runs only on local Docker; remote nodes run their own copy of `DockerEventService` and emit through the proxy. Each emitted message is prefixed with `[Node: <nodeName>]` so the source is unambiguous in the channel.
|
||||
|
||||
Recurring [vulnerability scans](/features/vulnerability-scanning) dispatch a notification whenever a run finishes so you do not have to check the task history manually.
|
||||
Switching the active node only affects per-stack rule editing and the Settings panels. The bell aggregates every node regardless of which one is active.
|
||||
|
||||
- **Info** when every image scanned successfully.
|
||||
- **Warning** when one or more images failed to scan during the run.
|
||||
- **Error** when the run itself could not start (for example, Trivy is not installed on the target node).
|
||||
## Alerts emitted by the system
|
||||
|
||||
The message includes the scheduled task name, a summary of how many images were scanned, cached, and failed, and a breakdown of findings by severity. A typical clean-run message looks like:
|
||||
Sencho dispatches notifications from many code paths. The list below covers every emitter in current code, organized by source. Each message is prefixed with `[Node: <nodeName>]` when emitted by a node-scoped service.
|
||||
|
||||
```
|
||||
Scheduled scan "nightly-scan" completed: Scanned 12 image(s); 3 skipped (cached). Found 2 critical, 5 high, 10 medium.
|
||||
```
|
||||
### Container crash and health
|
||||
|
||||
If no critical, high, or medium findings are present, the message ends with `No critical, high, or medium findings.` so the outcome is still explicit. When the target node has nothing to scan, the message reads `No images to scan.`, and when every image was already covered by a recent cached scan it reads `All N image(s) already scanned recently (cache hit).` In both cases the notification still fires so you know the run executed.
|
||||
Real-time on the Docker event stream:
|
||||
|
||||
<Note>
|
||||
Severity counts reflect the current security posture of the node, aggregated across both freshly scanned images and cached scan results. They are not a delta of what changed on this run.
|
||||
</Note>
|
||||
- **Crash** at `error`/`monitor_alert`: `Container Crash Detected: <name> exited unexpectedly (Code: <N>).`
|
||||
- **OOM kill** at `error`/`monitor_alert`: `Container OOM Kill: <name> was killed by the OOM killer (out of memory).` When a `die` arrives with exit code 137 but no preceding `oom` event, Sencho inspects the container and reclassifies as OOM if the `OOMKilled` flag is set.
|
||||
- **Healthcheck failure** at `error`/`monitor_alert`: `Healthcheck failed: <name> is unhealthy.`
|
||||
- **Mass exit** kicks in when a daemon-disconnect gap is followed by 20% or more of containers exiting on reconnect, in which case a single `info`/`system` summary `Docker daemon interruption detected: N containers exited during connection gap.` is emitted instead of N crash alerts.
|
||||
- **Rate limit** caps crash dispatches at 20 per 60-second window per node, then a single `warning`/`monitor_alert` `N additional containers crashed in the last minute.`
|
||||
- **Dedup** is 60 minutes per container after a non-rate-suppressed dispatch.
|
||||
|
||||
Failures are typically transient (registry timeouts, missing credentials) and do not stop the rest of the run from completing.
|
||||
### Daemon connectivity
|
||||
|
||||
## Container crash detection
|
||||
- `warning`/`system`: `Lost connection to Docker daemon; monitoring paused.` (one-shot until reconnect)
|
||||
- `info`/`system`: `Reconnected to Docker daemon.`
|
||||
- `warning`/`system`: `Received malformed Docker event payloads. Monitoring continues but some events may be skipped.` (when more than 10 parse errors in a minute)
|
||||
|
||||
Sencho watches every container on each of your nodes in real time and notifies you when something exits unexpectedly. Detection is causal: Sencho distinguishes crashes from intentional stops, so stopping a stack, restarting it, or running `docker compose down` will not produce a false crash alert.
|
||||
### Host thresholds
|
||||
|
||||
### What triggers a crash alert
|
||||
`warning`/`monitor_alert` for host CPU, RAM, and disk when the configured threshold is exceeded. 5-minute cooldown per signal. Example: `Host CPU utilization is critically high: 92% (Threshold: 90%)`. Evaluated on the 30-second monitor tick.
|
||||
|
||||
| Situation | Alert |
|
||||
|---|---|
|
||||
| Container exits with a non-zero exit code without being asked to stop | **Crash** alert (level: error) |
|
||||
| Container is killed by the kernel for exceeding its memory limit | **OOM Kill** alert (level: error) |
|
||||
| Healthcheck reports the container as unhealthy | **Healthcheck failed** alert (level: error) |
|
||||
### Docker janitor
|
||||
|
||||
Alerts arrive within a couple of seconds of the event, not on a polling interval.
|
||||
`info`/`system`: `Node "<name>" has accumulated <N> GB of unused Docker data. Consider using the Janitor tool.` 24-hour cooldown.
|
||||
|
||||
### What does not trigger a crash alert
|
||||
### Sencho version availability
|
||||
|
||||
- Stopping, restarting, updating, or removing a stack from Sencho
|
||||
- Running `docker stop`, `docker restart`, or `docker compose down` from a terminal on the host
|
||||
- A container exiting cleanly with exit code `0`
|
||||
- A container being replaced during an image update
|
||||
`info`/`system`: `Sencho X.Y.Z is available (currently running A.B.C). Visit the Fleet dashboard to update.` Polled every 6 hours; one-shot dedup until the running version reaches or passes the notified version.
|
||||
|
||||
### Global toggle
|
||||
### Image update availability
|
||||
|
||||
Crash and unhealthy alerts are gated by the **Global Crash Detection** toggle in **Settings > System**. When disabled, Sencho stops dispatching crash and unhealthy notifications on all nodes. Stack metric alerts and update notifications are unaffected by this toggle.
|
||||
`info`/`image_update_available`: `Stack "<name>" has image updates available.` Polled every 6 hours, with a 2-minute startup delay and a 2-minute cooldown on manual refresh. Notifies on **state transition only**; pre-existing `has_update` rows are backfilled silently on first run.
|
||||
|
||||
### Docker daemon interruptions
|
||||
### Auto-update execution
|
||||
|
||||
If the Docker daemon becomes unreachable (daemon restart, socket lost, network issue on a remote node), Sencho sends a single **Lost connection to Docker daemon** warning and pauses crash detection on the affected node. When the connection is restored, Sencho reconciles container state against a pre-disconnect snapshot:
|
||||
`info`/`image_update_applied`: `Auto-update: stack "<name>" updated with new images`. If a block-on-deploy policy gates the auto-update: `warning`/`scan_finding` `Policy "<name>" blocked auto-update: N image(s) exceed <severity>`. See [Auto-update policies](/features/auto-update-policies).
|
||||
|
||||
- If most of your containers are still running, individual gap exits are classified and alerts are sent as normal.
|
||||
- If a large share of containers exited during the outage (for example after a daemon restart), Sencho consolidates them into a single **Docker daemon interruption detected** informational notification instead of paging you for every container.
|
||||
### Auto-heal
|
||||
|
||||
A matching **Reconnected to Docker daemon** info notification confirms monitoring has resumed.
|
||||
- `info`/`autoheal_triggered`: `Auto-Heal: Restarted <container> on stack <stack> after being unhealthy for <N> minute(s).`
|
||||
- `warning`/`autoheal_triggered`: `Auto-Heal: Failed to restart <container> on stack <stack>. Error: <err>`
|
||||
- `warning`/`autoheal_triggered`: `Auto-Heal: Policy for <stack>[/<svc>] has been auto-disabled after <N> consecutive failures.`
|
||||
|
||||
### High-churn events
|
||||
See [Auto-heal policies](/features/auto-heal-policies).
|
||||
|
||||
When many crashes land in a short window (for example, a large stack coming down unexpectedly), Sencho dispatches an initial batch and summarizes the remainder in a single **N additional containers crashed in the last minute** entry so your notification channels are not flooded.
|
||||
### Vulnerability scanning
|
||||
|
||||
- **Per-violation, during a scheduled scan** is `warning`/`scan_finding`. One alert per violation: `Policy "<name>" violated by <imageRef>: <severity> exceeds <maxSeverity>`.
|
||||
- **Scan completion** is `info`/`scan_finding` (or `warning` if any image failed): `Scheduled scan "<task name>" completed: <output>` where `<output>` is one of `Scanned N image(s); X skipped (cached). Found A critical, B high, C medium.`, `No critical, high, or medium findings.`, `No images to scan`, or `All N image(s) already scanned recently (cache hit)`.
|
||||
- **Pre-deploy gate when Trivy is missing** is `warning`/`scan_finding`: `Pre-deploy scan for "<stack>" skipped: Trivy not installed on this node`.
|
||||
- **Post-deploy scan finding** is `error`/`scan_finding` if criticals are present, `warning` otherwise: `Vulnerability scan for <imageRef>: <N> critical, <M> high`.
|
||||
- **Post-deploy scan failure** is `warning`/`scan_finding`: `Post-deploy scan failed for <imageRef> (<stack>): <msg>`.
|
||||
- **Trivy auto-update** is `info`/`system`: `Trivy updated from vX to vY` or `Trivy update available: vX (currently vY)`.
|
||||
|
||||
See [Vulnerability scanning](/features/vulnerability-scanning).
|
||||
|
||||
### Scheduled tasks
|
||||
|
||||
- `error`/`system`: `Scheduled task "<name>" (<action>) failed: <err>`
|
||||
- `info`/`system` recovery: `Scheduled task "<name>" (<action>) recovered successfully`
|
||||
|
||||
### Cloud Backup upload failure
|
||||
|
||||
`warning`/`system`: `Cloud backup failed for scheduled snapshot <id>: <message>`. See [Fleet backups](/features/fleet-backups).
|
||||
|
||||
### Blueprints (Admiral)
|
||||
|
||||
- `warning`/`blueprint_drift_detected`: `Blueprint "<n>" drifted on node "<n2>": <reason>` (suggest mode), with stateful-safeguard variants when enforce mode declines to redeploy.
|
||||
- `error`/`blueprint_drift_correction_failed`: `Auto-fix for "<n>" on node "<n2>" failed: <err>`.
|
||||
|
||||
See [Blueprints](/features/blueprint-model).
|
||||
|
||||
### Fleet sync (Admiral)
|
||||
|
||||
- `warning`/`system` identity drift: `Fleet self-identity changed from "X" to "Y". Identity-scoped policies are reapplied on the next sync.`
|
||||
- `warning`/`system` truncation: `Fleet sync truncated <resource> to <N> rows (<dropped> not replicated). Reduce the local set or contact support.`
|
||||
- `warning`/`system` stale: `Fleet sync to node "<n>" has been failing for over an hour for <resource>. Check the node's connectivity and API token.`
|
||||
|
||||
## Retention and limits
|
||||
|
||||
Three retention controls live under **Settings · Developer · Data retention**.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/alerts-notifications/developer-retention.png" alt="Settings · Developer · Data retention card showing three rows: Container metrics with a 24 HRS field, Notification log with a 30 DAYS field, and Audit log (Admiral) with a 90 DAYS field, plus a SAVE SETTINGS button." />
|
||||
</Frame>
|
||||
|
||||
| Control | Range | Default | What it prunes |
|
||||
|---------|-------|---------|----------------|
|
||||
| **Container metrics** | 1 to 8760 hours | 24 hours | Per-container CPU, RAM, and network history used by the dashboard sparklines |
|
||||
| **Notification log** | 1 to 365 days | 30 days | The bell's `notification_history` table |
|
||||
| **Audit log** *(Admiral)* | 1 to 365 days | 90 days | Audit trail entries |
|
||||
|
||||
A hard 100-row per-node cap inside `notification_history` always applies on top of the user-tunable retention; new inserts evict the oldest rows beyond 100 even if your retention window is longer.
|
||||
|
||||
A separate rate limit applies to crash and health alerts only: 20 emits per 60-second window per node, with a single `warning`/`monitor_alert` roll-up `N additional containers crashed in the last minute.` issued at the end of the window.
|
||||
|
||||
## Crash detection toggle
|
||||
|
||||
The global crash-capture switch lives under **Settings · System · Docker hygiene**.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/alerts-notifications/system-crash-toggle.png" alt="Settings · System · System Limits panel showing a HOST THRESHOLDS section with CPU limit, RAM limit, and Disk limit at 100%, then a DOCKER HYGIENE section with a Janitor threshold of 5 GiB and a Global crash capture toggle in the ON state with the helper line 'Watch every managed container for unexpected exits.'" />
|
||||
</Frame>
|
||||
|
||||
The **Global crash capture** toggle controls whether `DockerEventService` raises crash, OOM, and healthcheck alerts on the active node. Helper text: `Watch every managed container for unexpected exits.` Defaults to on; if the database read fails, Sencho falls back to default-deny so the system never leaks alerts you cannot turn off.
|
||||
|
||||
The same panel carries the **Host thresholds** rows (CPU limit, RAM limit, Disk limit, all expressed as percent) that drive the host-level monitor warnings, and the **Janitor threshold** (in GiB) that drives the unused-Docker-data alert.
|
||||
|
||||
## Refresh cadence
|
||||
|
||||
| Surface | Cadence |
|
||||
|---------|---------|
|
||||
| Crash, OOM, and healthcheck events | Real-time over the Docker event stream |
|
||||
| Host CPU / RAM / disk threshold checks | 30 seconds |
|
||||
| Per-stack alert rule evaluation | 30 seconds |
|
||||
| Image update poll | 6 hours, with a 2-minute startup delay and a 2-minute cooldown on manual refresh |
|
||||
| Sencho version check | 6 hours |
|
||||
| Notification fanout to channels | Single shot per dispatch, 10-second timeout, no retries |
|
||||
| Bell live updates | Pushed live over the notifications WebSocket per node |
|
||||
| Bell safety-net reconcile | 60 seconds |
|
||||
| Crash dedup window per container | 60 minutes |
|
||||
| Crash rate-limit window | 20 alerts per 60 seconds per node, then a single roll-up |
|
||||
|
||||
Switching the active node tears down per-stack rule editors and reloads channel state, but does not affect the bell, which keeps every node's WebSocket open in parallel.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### Notifications not being delivered
|
||||
|
||||
- Verify at least one notification channel is enabled in **Settings > Notifications**
|
||||
- Click **Test** on the channel to confirm the webhook URL is reachable
|
||||
- Check that the webhook URL uses HTTPS
|
||||
- If using notification routing (Admiral tier), verify the route pattern matches the stack name
|
||||
|
||||
### Alert not firing
|
||||
|
||||
- Confirm the alert rule exists by opening the stack's alert panel
|
||||
- Check that the **duration** has elapsed; the condition must hold continuously for the configured duration before an alert fires
|
||||
- Check the **cooldown** period; after an alert fires, it will not fire again until the cooldown expires
|
||||
- Verify the metric is being collected; container must be running for stats to be gathered
|
||||
|
||||
### I stopped a stack but got a crash alert
|
||||
|
||||
This should not happen. Sencho distinguishes intentional stops from crashes by correlating each container exit with the action that triggered it. If you do see a crash alert for a stack you stopped, confirm Sencho can reach the Docker daemon on the affected node: the **Lost connection to Docker daemon** warning is dispatched when the socket is unreachable, and exits observed during an outage may be classified as crashes on reconnect.
|
||||
|
||||
### After a Docker daemon restart I only got one summary notification
|
||||
|
||||
Expected. When a large share of containers exits during a Docker daemon interruption, Sencho consolidates them into a single informational notification instead of paging per container. Individual crashes that happen after the reconnect are alerted on as normal.
|
||||
|
||||
### Crash alerts stopped arriving and nothing else looks wrong
|
||||
|
||||
Check, in order:
|
||||
|
||||
1. Open **Settings > System** and confirm **Global Crash Detection** is on. When it is off, crash, OOM, and unhealthy notifications stop on all nodes while stack metric alerts and update notifications continue as normal.
|
||||
2. Look in the in-app notifications panel for a **Lost connection to Docker daemon** warning. If present, Sencho is not receiving events from the affected node and is retrying in the background; a **Reconnected to Docker daemon** info entry will appear when monitoring resumes.
|
||||
3. Confirm the affected node is reachable from the host running Sencho. For remote nodes, the remote Sencho instance is responsible for its own crash detection; check its status from that instance.
|
||||
|
||||
### A burst of crashes happened but I only see about twenty alerts
|
||||
|
||||
Expected, and by design. Sencho caps crash notifications at around 20 per minute per node so a runaway restart loop or a stack-wide failure cannot flood your notification channels. Any crashes beyond the cap in that window are rolled up into a single **N additional containers crashed in the last minute** warning. The complete, unredacted list of events is always visible in the in-app notifications panel; only external channel delivery (Discord, Slack, webhooks) is rate-limited.
|
||||
|
||||
### Delete confirmation dialog
|
||||
|
||||
Deleting an alert rule now requires confirmation. Click the trash icon next to a rule, then confirm in the dialog that appears.
|
||||
|
||||
### Fleet shows an update button but I never received a Sencho version notification
|
||||
|
||||
Fleet and the in-app bell share the same version lookup. Each Sencho instance checks for a new release roughly every 6 hours and notifies exactly once per version, so if a check already fired for the latest release the bell will not notify again for that same version.
|
||||
|
||||
If the update button is visible but no notification is in the bell, a transient network failure during the last successful check can briefly delay delivery; the next evaluation cycle (about every 30 seconds) retries automatically. To confirm external channels (Discord, Slack, webhooks) are reachable, open **Settings > Notifications** and click **Test** on each. The in-app bell receives every notification regardless of external channel configuration.
|
||||
|
||||
### A stack shows the blue update indicator but no notification was received
|
||||
|
||||
Open the notification bell. Dispatch failures (for example, a misconfigured Discord or Slack webhook) are logged as error entries in the bell so they are visible without tailing logs. Confirm at least one channel is enabled in **Settings > Notifications**; the in-app bell receives all notifications regardless of external channel configuration.
|
||||
<AccordionGroup>
|
||||
<Accordion title="Notifications never arrive in Discord, Slack, or my webhook">
|
||||
Check three things in order. First, the channel toggle in **Settings · Notifications** must be on; the kicker on each tab reads `enabled` or `off`. Second, the URL must use HTTPS; the form rejects plain `http://` outright. Third, an Admiral routing rule with empty `Stacks`, `Labels`, and `Categories` matchers will intercept every alert and skip the global channels. Use the per-channel **Test** button to issue a one-shot dispatch and watch your endpoint for the literal message `🔌 Test Notification from Sencho!` Sencho records the failure reason in `notification_history.dispatch_error` when delivery throws, so a row that appears in the bell with no follow-up at the endpoint usually means a 4xx or timeout at the receiver.
|
||||
</Accordion>
|
||||
<Accordion title="An alert rule never fires even when the threshold is breached">
|
||||
Three causes account for almost every case. First, the rule's **Duration** has not elapsed yet: the breach must persist for the full duration before the rule fires. Second, the rule is still in cooldown after a previous fire. Third, the panel's banner is not green: a remote-node banner means the rule was saved on a remote whose channels you may not have configured, and an amber `No notification channels configured` banner means the rule evaluates fine but Sencho has nowhere to send the alert. The evaluator runs on a 30-second tick, so expect up to 30 seconds of latency between the breach starting and the timer engaging.
|
||||
</Accordion>
|
||||
<Accordion title="I stopped a stack but got a crash alert anyway">
|
||||
The Docker event service defers `die` classification by 500 ms to absorb out-of-order `kill` events from the daemon, then asks the lifecycle classifier whether the exit was intentional, clean (exit 0), a crash, or an OOM kill. If your stop happened far enough outside that window, or the daemon emitted the events without the kill marker the classifier looks for, the exit can be classified as a crash. The classifier favors avoiding silent crashes over avoiding noisy false positives. Compare the alert timestamp against your `docker compose down` time; entries within a second of each other are usually the same event seen from two angles.
|
||||
</Accordion>
|
||||
<Accordion title="After a Docker daemon restart, the bell only has one summary line">
|
||||
When the daemon disconnects and reconnects, Sencho snapshots every container at the moment of reconnect. If at least 20% exited during the gap, the service emits a single `info`/`system` summary `Docker daemon interruption detected: N containers exited during connection gap.` instead of one crash alert per container. Below the threshold, every gap exit is classified individually.
|
||||
</Accordion>
|
||||
<Accordion title="Crash alerts stopped arriving and nothing else looks wrong">
|
||||
Check **Settings · System · Docker hygiene · Global crash capture**. The toggle is the master switch for the Docker event service. The panel cache reads the database every 500 ms; if the database read errors, Sencho defaults the toggle to off so the failure mode is silent rather than spammy. Repair the toggle, save, and the next Docker event reaches the dispatcher.
|
||||
</Accordion>
|
||||
<Accordion title="A burst of crashes happened but I only see about twenty alerts">
|
||||
Sencho rate-limits crash and health dispatches to 20 emits per rolling 60-second window per node, then emits a single `warning`/`monitor_alert` roll-up `N additional containers crashed in the last minute.` once the window closes. Every alert is still persisted to `notification_history` and visible in the bell up to the 100-row per-node cap; only the channel fanout is throttled.
|
||||
</Accordion>
|
||||
<Accordion title="My deploy success doesn't appear in the bell, only as a toast">
|
||||
Sencho deliberately hides rows whose category is one of `deploy_success`, `stack_started`, `stack_stopped`, `stack_restarted`, or `image_update_applied` AND whose `actor_username` is set to a real user. The reasoning: those are confirmations of the action you just clicked and are already shown as a toast. The rows are still persisted to `notification_history` and still dispatched to global channels and matching routes; only the bell render hides them.
|
||||
</Accordion>
|
||||
<Accordion title="A routing rule is set up but the global channel still fires">
|
||||
Routing matchers AND together: every non-empty matcher must match the alert. A rule with **Stacks** set to `prod-api` will not match a `monitor_alert` for a different stack, and a rule with a populated **Stacks** matcher will not match host-level alerts (which carry no stack target). When zero rules match, Sencho falls back to global channels. To intercept everything, leave all four matcher fields empty on the rule. Also confirm the rule's **Enabled** pill is `ON`.
|
||||
</Accordion>
|
||||
<Accordion title="Fleet shows an update button but I never received a Sencho version notification">
|
||||
The version poll runs every 6 hours and dedups against the `last_sencho_update_notified_version` system-state key, so each newly-released version produces exactly one alert per node. If you upgraded Sencho between the release and the poll, the dedup self-heals; the alert simply does not fire because the running version already matches. To force a fresh check, restart the Sencho container; the version poll runs once on startup after a 2-minute delay.
|
||||
</Accordion>
|
||||
<Accordion title="A stack shows the blue update indicator but no notification was received">
|
||||
The image-update service emits on **state transitions only**. The first run of the service backfills a flag against rows that already had `has_update: true` so they do not all re-fire on first install. Once the flag is set, only the false-to-true transition triggers an alert. If a stack was already showing the indicator before backfill ran, a notification will only appear the next time it goes from up-to-date back to behind.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
@@ -1,39 +1,81 @@
|
||||
---
|
||||
title: Atomic Deployments
|
||||
description: Zero-downtime deployments with automatic rollback for Skipper and Admiral users.
|
||||
description: Wrap every deploy and update in a backup, a 3-second health probe, and an automatic rollback when a container crashes.
|
||||
---
|
||||
|
||||
<Note>
|
||||
Atomic Deployments require a **Sencho Skipper** or **Admiral** license. Community Edition uses standard deployments without backup or rollback.
|
||||
</Note>
|
||||
Sencho wraps every protected deploy in a four-step safety net: it copies the current `compose.yaml` and `.env` to a writable backup directory, runs the requested compose action, waits 3 seconds for the new containers to settle, then checks them for a non-zero exit code. If any container has crashed, Sencho restores the backed-up files and re-deploys automatically.
|
||||
|
||||
Sencho wraps every deployment in a safety net on Skipper and Admiral tiers. Before applying changes, it backs up your current configuration. If the deployment fails, it automatically rolls back to the previous working state.
|
||||
The same backup also powers the **Rollback** action in the stack editor, so you can roll a stack back to its last good configuration on demand.
|
||||
|
||||
<Note>
|
||||
Atomic Deployments require a Sencho **Skipper** or **Admiral** license. Community Edition runs the same compose actions without a backup or automatic rollback.
|
||||
</Note>
|
||||
|
||||
## How it works
|
||||
|
||||
1. **Backup** - Before a deploy or update, Sencho copies your `compose.yaml` and `.env` files to a safe internal location
|
||||
2. **Deploy** - Sencho runs the requested compose operation (up, pull + recreate, etc.)
|
||||
3. **Health probe** - After deployment, Sencho waits briefly, then checks whether any containers exited with a non-zero exit code
|
||||
4. **Auto-rollback** - If a crash is detected, Sencho restores the backed-up files and re-deploys automatically
|
||||
1. **Backup.** Before the action runs, Sencho copies `compose.yaml` (or `compose.yml` / `docker-compose.yaml` / `docker-compose.yml`) and `.env`, if present, into the backup directory. The deploy progress modal streams `=== Backup created for atomic deployment ===` once the copy completes, before any `docker compose` output.
|
||||
2. **Run the action.** Sencho executes the requested compose action: `up -d` for a deploy, or a pull-then-`up -d` recreate for an update.
|
||||
3. **Health probe.** Sencho waits 3 seconds, then lists every container with the `com.docker.compose.project=<stack>` label and checks each one for a non-zero exit code. Any container that has exited with a non-zero status counts as a crash.
|
||||
4. **Auto-rollback on failure.** When a crash is detected, Sencho streams `=== Deployment failed - rolling back to previous version ===`, restores the backed-up files, and re-runs `docker compose up -d` with the restored configuration. On success it streams `=== Rolled back successfully ===`. The original deploy error is preserved and reported as the deploy result, so a failed-then-rolled-back deploy still registers as a failure in the deploy progress modal.
|
||||
|
||||
This entire sequence happens transparently. You see a single deploy action; Sencho handles the safety logic behind the scenes.
|
||||
If the rollback itself fails (for example, the re-deploy step cannot pull a previously available image, or the file restore is blocked by filesystem permissions), Sencho streams `=== Rollback failed - manual intervention may be required ===`. The backup files remain at `<DATA_DIR>/backups/<stack>/` so you can copy them back manually.
|
||||
|
||||
## Which operations are protected
|
||||
|
||||
Atomic deployments apply to:
|
||||
Atomic deployments wrap:
|
||||
|
||||
- **Deploy** from the stack editor (compose down + up)
|
||||
- **Update** from the stack editor (pull latest images + recreate)
|
||||
- **Webhook triggers** for deploy and pull actions
|
||||
- **Scheduled tasks** that perform deploys or updates
|
||||
- **App Store installs** when deploying a new stack
|
||||
- **Deploy** and **Update** from the stack editor's action bar.
|
||||
- **App Store** installs of a new stack.
|
||||
- **Webhook** triggers for deploy and pull actions.
|
||||
- **Image auto-updates** triggered by an auto-update policy.
|
||||
|
||||
Scheduled tasks invoke `docker compose` without the atomic wrapper, so a deploy or update launched from a schedule runs without a backup or auto-rollback. If you need atomic safety on a recurring deploy, trigger it through a webhook on a cron rather than through Scheduled Tasks.
|
||||
|
||||
## Manual rollback
|
||||
|
||||
You can manually roll back to the previous deployment at any time by clicking the **Rollback** button in the stack editor's action bar. The button only appears when a backup exists from a prior deployment.
|
||||
The stack editor's action bar carries a **More actions** overflow menu (the three-dot icon next to **Update**). Open it on a Skipper or Admiral instance and you'll see **Rollback** at the top, with the timestamp of the most recent backup rendered beneath the label. Selecting it restores the backed-up files and re-runs `docker compose up -d` non-atomically, to avoid nesting a rollback inside another atomic wrapper and overwriting the good backup with the broken state from the just-failed deploy.
|
||||
|
||||
Hover over the Rollback button to see a tooltip with the timestamp of the last backup. Clicking it restores the backed-up `compose.yaml` and `.env` files, then re-deploys the stack with the restored configuration.
|
||||
<Frame>
|
||||
<img src="/images/atomic-deployments/rollback-menu.png" alt="Stack editor action bar with the More actions overflow menu open, showing the Rollback entry at the top with the backup timestamp rendered beneath the label, followed by Scan config and Delete entries" />
|
||||
</Frame>
|
||||
|
||||
The menu entry is hidden when no backup exists for the stack, for example on a freshly created stack that has never been deployed atomically, and on a Community Edition instance. The endpoint additionally requires the `stack:deploy` permission, so a user without it will see the menu entry but receive a permission error if they invoke it.
|
||||
|
||||
## Where backups are stored
|
||||
|
||||
Backups live under `<DATA_DIR>/backups/<stack>/`, in the same writable volume Sencho uses for its database and other persisted state. They are intentionally kept outside the user's compose folder, so the operation works even when a container has chowned its bind-mounted stack directory to root.
|
||||
|
||||
Each backup is a flat copy of the compose file Sencho found, plus `.env` if it exists, plus a `.timestamp` marker that records when the backup was taken. There is one backup slot per stack: every protected deploy or update overwrites the previous backup, so the **Rollback** menu always reverts to the configuration that was on disk immediately before the most recent run.
|
||||
|
||||
## Community Edition behavior
|
||||
|
||||
Community users continue to use the standard deploy flow: no backup is created and no rollback is available. Upgrading to Skipper or Admiral enables atomic deployments immediately with no configuration required.
|
||||
On Community Edition, Sencho runs the same `docker compose` commands without the atomic wrapper. There is no backup, no health probe, and no automatic rollback, and the **Rollback** menu entry is hidden. On a Skipper or Admiral license, atomic deployments are active immediately for every protected action; no configuration is required.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="The Rollback option is not in the More actions menu">
|
||||
Sencho hides the entry whenever a rollback is not possible. The most common reasons are:
|
||||
|
||||
- The stack has never been deployed atomically, so no backup file exists yet. Run **Deploy** or **Update** once and the entry will appear.
|
||||
- The instance is on Community Edition. Atomic Deployments require Skipper or Admiral.
|
||||
|
||||
A user without the `stack:deploy` permission will still see the menu entry; the rejection comes from the backend with a permission error after they click. Ask an admin to grant `stack:deploy` through **Settings · Roles & Access** if that happens.
|
||||
</Accordion>
|
||||
<Accordion title="The deploy succeeded but a service crashed seconds later">
|
||||
The health probe is a 3-second window after `docker compose up -d` returns. Crashes that happen after that window are out of scope for atomic rollback by design, because Sencho cannot tell whether a late exit is a real failure or a normal restart. For ongoing health, use **Auto-Heal Policies** to restart unhealthy containers automatically and **Alert Rules** to page you when a container exits unexpectedly.
|
||||
</Accordion>
|
||||
<Accordion title="The deploy progress modal showed 'Rollback failed - manual intervention may be required'">
|
||||
This message means the auto-rollback attempted to restore the backup and re-deploy, but the restore step or the re-deploy itself errored out. The backup files are still at `<DATA_DIR>/backups/<stack>/`. To recover:
|
||||
|
||||
1. Copy `compose.yaml` (or the variant Sencho backed up) and `.env` from `<DATA_DIR>/backups/<stack>/` back into the stack directory.
|
||||
2. Open the stack in the editor and click **Deploy** to re-run with the restored configuration.
|
||||
|
||||
The most common causes are filesystem permissions on the stack directory and a missing image in a private registry that the original deploy could not pull.
|
||||
</Accordion>
|
||||
<Accordion title="Rollback ran but the stack still has the broken configuration">
|
||||
Sencho keeps a single backup per stack. If you ran two atomic deploys back to back, the second deploy overwrote the first backup with the broken configuration before it failed. **Rollback** then restores that broken configuration, because as far as Sencho is concerned it is the most recent known state.
|
||||
|
||||
To recover, edit the compose file or `.env` directly in the editor, fix the bad change, and click **Deploy**. The next protected deploy will write a fresh backup of the now-good configuration.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
@@ -1,10 +1,10 @@
|
||||
---
|
||||
title: Audit Log
|
||||
description: Track all mutating actions across your Sencho instance with a searchable, exportable audit trail for team accountability.
|
||||
description: Track every mutating action on your Sencho instance with a searchable, exportable trail for team accountability.
|
||||
---
|
||||
|
||||
<Note>
|
||||
The Audit Log requires a Sencho **Admiral** license. Skipper and Community Edition do not include this feature.
|
||||
The Audit Log requires a Sencho **Admiral** license. Skipper and Community do not include this feature.
|
||||
</Note>
|
||||
|
||||
<Note>
|
||||
@@ -23,86 +23,85 @@ Every `POST`, `PUT`, `DELETE`, and `PATCH` request to the Sencho API is automati
|
||||
| **User** | The authenticated username that performed the action |
|
||||
| **Method** | HTTP method (`POST`, `PUT`, `DELETE`, `PATCH`), shown as a color-coded badge |
|
||||
| **Action** | Human-readable summary (e.g., "Deployed stack: nginx-proxy") |
|
||||
| **Status** | HTTP response status code, color-coded by result (green for success, yellow for client errors, red for server errors) |
|
||||
| **Node** | Which node the action targeted, or `-` for local-only actions |
|
||||
| **Status** | HTTP response status code, color-coded by result (green for success, amber for client errors, rose for server errors) |
|
||||
| **Node** | Numeric node ID the action targeted, or `-` for local-only actions |
|
||||
|
||||
Expanding a row reveals additional detail:
|
||||
Expanding a row in the Table view reveals additional detail:
|
||||
|
||||
| Field | Description |
|
||||
|-------|-------------|
|
||||
| **Request Path** | Full API path of the request |
|
||||
| **IP Address** | Client IP address |
|
||||
| **IP Address** | Client IP address as recorded by the gateway |
|
||||
| **Node ID** | Numeric node ID, or "Local" for local actions |
|
||||
| **Entry ID** | Unique identifier for the audit entry |
|
||||
|
||||
### Example actions tracked
|
||||
|
||||
- Stack lifecycle: deploy, stop, start, restart, update, rollback, delete
|
||||
- Stack creation, file edits, and env file changes
|
||||
- Stack creation, file edits, env file changes, label changes
|
||||
- Per-service lifecycle (start, stop, restart of an individual compose service)
|
||||
- Container operations: start, stop, restart
|
||||
- Node management: add, update, delete
|
||||
- User management: create, delete, role assignment and removal
|
||||
- Node management: add, update, delete, cordon, uncordon
|
||||
- Fleet operations: backup creation, restore, deletion, single-node and fleet-wide updates
|
||||
- Fleet replica role changes (re-anchor, demote to control)
|
||||
- Fleet Secrets: create, update, delete, import-from-stack, push (and push preview)
|
||||
- Blueprint federation pin updates
|
||||
- Sencho Cloud Backup: config update, connection test, provisioning, snapshot upload, snapshot deletion
|
||||
- User management: create, update, delete, role assignment and removal
|
||||
- Password changes and node token generation
|
||||
- Settings changes
|
||||
- System prune operations (general, orphans, system)
|
||||
- Image, volume, and network deletion
|
||||
- Network creation
|
||||
- License activation/deactivation
|
||||
- License activation and deactivation
|
||||
- Webhook and notification agent configuration
|
||||
- Notification route management and testing
|
||||
- Fleet backup creation, restoration, and deletion
|
||||
- Fleet node updates (individual and fleet-wide)
|
||||
- Scheduled task management
|
||||
- Registry credential management
|
||||
- SSO configuration changes and connection tests
|
||||
- API token creation and revocation
|
||||
- Label management and bulk actions
|
||||
- Template deployments
|
||||
- Registry credential management
|
||||
- Scheduled task management and manual triggers
|
||||
- Settings changes
|
||||
- Auto-update execution
|
||||
- Template deployments
|
||||
- Console token generation
|
||||
- System prune (general, orphans, system), image / volume / network deletion, network creation
|
||||
|
||||
## Viewing the audit log
|
||||
|
||||
Navigate to the **Audit** tab in the sidebar. This tab is visible to users with the **Admin** or **Auditor** role on an Admiral license.
|
||||
Navigate to the **Audit** tab in the sidebar. The tab is visible on Admiral to users with the **Admin** or **Auditor** role.
|
||||
|
||||
The Audit Log has two views, toggled from the header: **Stream** (default) and **Table**.
|
||||
The page has two views, toggled from the segmented control in the card header: **Stream** (default) and **Table**. The card subtitle reports the total number of entries that match the current filters.
|
||||
|
||||
### Stream view
|
||||
|
||||
Stream gives you an at-a-glance read on activity. A signal rail at the top summarizes the last 24 hours across four tiles, and the feed below groups entries by day with severity dots, relative times, and inline anomaly callouts.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/audit-log/audit-stream.png" alt="Audit Log Stream view showing the signal rail with events, actors, failure rate, and peak hour tiles above a day-banded chronological feed" />
|
||||
<img src="/images/audit-log/audit-stream.png" alt="Audit Log Stream view with the four-tile signal rail (Events 111 +91% vs 7d avg, Actors 1, Failure rate 9% with sparkline, Peak hour 06:00) above a day-banded chronological feed of admin POST and DELETE entries." />
|
||||
</Frame>
|
||||
|
||||
**Signal rail tiles:**
|
||||
|
||||
| Tile | What it shows |
|
||||
|------|---------------|
|
||||
| **Events · 24h** | Count of audit entries in the last 24 hours, with a percent change versus the prior 7-day average |
|
||||
| **Actors** | Unique users active in the last 24 hours; notes the count of new IP addresses when present |
|
||||
| **Failure rate** | Share of 24-hour requests with 4xx or 5xx responses, rendered with an inline sparkline of the hourly failure trend |
|
||||
| **Peak hour** | The hour with the highest activity, flagged in amber when it falls outside 08:00-18:00 |
|
||||
| **Events · 24h** | Count of audit entries in the last 24 hours, with a percent change versus the prior 7-day daily average |
|
||||
| **Actors** | Unique users active in the last 24 hours; the detail line names a sample actor and the count of new IP addresses when any new IPs are seen for an existing user |
|
||||
| **Failure rate** | Share of 24-hour requests with 4xx or 5xx responses, rendered with an inline sparkline of the hourly failure trend; tints amber at 5% and above, rose at 20% and above |
|
||||
| **Peak hour** | The hour with the highest activity. Renders blank when the peak sits inside working hours (08:00 to 17:59); flips to amber and shows the hour when the peak falls outside that window |
|
||||
|
||||
**Feed entries:** each row shows the relative time (e.g. `58m ago`), a severity dot (green for 2xx, amber for 3xx, rose for 4xx/5xx), the actor and action summary, a meta line with exact timestamp, node, status code, IP, and any anomaly flags, and the method and path on the right.
|
||||
**Feed entries:** each row shows the relative time (e.g. `58m ago`), a severity dot (green for 2xx, amber for 3xx, rose for 4xx and 5xx), the actor and action summary, a meta line with the exact timestamp, the node, the status code, the IP, and any anomaly flags, plus the method and path on the right.
|
||||
|
||||
Rows colored in rose or amber indicate failures; the tinting makes spikes of errors visible at a glance.
|
||||
|
||||
### Table view
|
||||
|
||||
Table keeps the full-featured detail grid for power users: exact timestamps, method badges, action summaries, and status codes in sortable columns. Clicking any row expands it to show the full request path, IP address, node ID, and entry ID.
|
||||
Table keeps the full-featured detail grid for power users: exact timestamps, method badges, action summaries, and status codes in fixed columns. Clicking any row expands it to show the full request path, IP address, node ID, and entry ID.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/audit-log/audit-log-overview.png" alt="Audit Log Table view showing the action table with search filters, color-coded method badges, and export controls" />
|
||||
<img src="/images/audit-log/audit-log-overview.png" alt="Audit Log Table view with a filter strip (search box, All Methods dropdown, From and To date pickers) above a six-column table of admin POST, PUT, and DELETE entries against /api/blueprints, /api/notifications, and /api/stacks." />
|
||||
</Frame>
|
||||
|
||||
<Frame>
|
||||
<img src="/images/audit-log/audit-log-expanded.png" alt="Audit Log Table view with an expanded row showing request path, IP address, node ID, and entry ID" />
|
||||
<img src="/images/audit-log/audit-log-expanded.png" alt="Audit Log Table view with the second row expanded to reveal Request Path /api/stacks/plex/restart, IP Address ::ffff:192.0.2.10, Node ID 1, and Entry ID #820 in a four-cell detail strip." />
|
||||
</Frame>
|
||||
|
||||
The header displays the total number of matching entries and provides **Refresh** and **Export** controls in both views.
|
||||
|
||||
Results are paginated at 50 entries per page. Navigation controls appear at the bottom of the feed or table when there are multiple pages.
|
||||
Both views share the **Refresh** button and the **Export** dropdown in the card header, and both paginate at 50 entries per page with chevron controls at the bottom of the feed or table.
|
||||
|
||||
## Anomaly detection
|
||||
|
||||
@@ -110,53 +109,55 @@ In Stream view, Sencho annotates individual entries with lightweight anomaly fla
|
||||
|
||||
| Flag | When it fires |
|
||||
|------|---------------|
|
||||
| **unusual hour** | The entry's hour sits outside the actor's typical activity window, computed from the central 90% of their last 7 days. Requires a baseline of at least 5 prior entries to avoid flagging new users. |
|
||||
| **new ip** | The IP address on this entry has not been seen for this actor in the last 30 days, and the actor already has prior history. |
|
||||
| **unusual hour** | The entry's hour sits outside the central 90% of the actor's typical activity window, computed from the last 7 days. Requires a baseline of at least 5 prior entries to avoid flagging new users. |
|
||||
| **new ip** | The IP address on this entry has not been seen for this actor in the last 30 days, and the actor already has at least one prior IP on record. |
|
||||
| **first seen** | The actor has no prior entries in the 30-day history window. Useful for spotting brand-new service accounts, CLI tools, or compromised sessions. |
|
||||
|
||||
Flags are computed at read time against your existing audit history. No new tables, no per-entry storage overhead, and the logic is cache-friendly, so enabling Stream view does not slow down the audit log endpoint.
|
||||
|
||||
## Filtering and search
|
||||
|
||||
The audit log provides several ways to find specific entries:
|
||||
Filters live in the **Table view only**. Switching to Stream hides the filter strip; the Stream feed always shows the most recent unfiltered entries grouped by day.
|
||||
|
||||
- **Full-text search** - Search across action summaries, API paths, and usernames
|
||||
- **Method filter** - Filter by HTTP method (POST, PUT, DELETE, PATCH) using the dropdown
|
||||
- **Date range** - Set a start and/or end date to narrow results to a specific time window
|
||||
In Table view you can combine:
|
||||
|
||||
All filters work together and are applied server-side with pagination.
|
||||
- **Full-text search** across action summaries, API paths, and usernames
|
||||
- **Method** filter (POST, PUT, DELETE, PATCH) via the dropdown
|
||||
- **Date range** with **From** and **To** pickers, narrowing results to a specific time window
|
||||
|
||||
All filters AND together and are applied server-side with pagination.
|
||||
|
||||
## Export
|
||||
|
||||
Export the currently filtered audit log dataset as **CSV** or **JSON** using the **Export** dropdown in the header. The export respects all active filters, so you can narrow down to a date range or specific user before exporting.
|
||||
Export the currently filtered audit log as **CSV** or **JSON** using the **Export** dropdown in the card header. The export respects every active filter, so you can narrow down to a date range or specific user before exporting.
|
||||
|
||||
Exports are capped at 10,000 entries per download.
|
||||
Exports are capped at 10,000 entries per download. To export a larger window, narrow the date range and download in chunks.
|
||||
|
||||
## Auditor role
|
||||
|
||||
The **Auditor** role provides read-only access to the audit log without granting any administrative privileges. Auditors can:
|
||||
|
||||
- View the full audit log
|
||||
- Search and filter entries
|
||||
- Search and filter entries (Table view)
|
||||
- Export audit data as CSV or JSON
|
||||
- View stacks and nodes (read-only)
|
||||
|
||||
Auditors **cannot** modify settings, manage users, deploy stacks, or perform any other administrative actions. This role is ideal for compliance officers, security reviewers, or team leads who need visibility into system activity without operational access.
|
||||
Auditors **cannot** modify settings, manage users, deploy stacks, or perform any other administrative action. The role is ideal for compliance officers, security reviewers, or team leads who need visibility into system activity without operational access.
|
||||
|
||||
To create an Auditor user, go to **Settings > Users** and select the **Auditor** role when creating a new user.
|
||||
To create an Auditor user, go to **Settings · Users** and select the **Auditor** role when creating a new user.
|
||||
|
||||
## Configurable data retention
|
||||
|
||||
Audit log entries are automatically cleaned up based on your configured retention period. The default is **90 days**.
|
||||
Audit log entries are automatically pruned based on your configured retention period. The default is **90 days**.
|
||||
|
||||
To change the retention period:
|
||||
|
||||
1. Go to **Settings > Developer > Data Retention**
|
||||
2. Set the **Audit Log Retention** value (1-365 days)
|
||||
3. Click **Save Developer Settings**
|
||||
1. Open **Settings · Developer**
|
||||
2. In the **Data retention** card, set the **Audit log** value (1 to 365 days)
|
||||
3. Click **Save settings**
|
||||
|
||||
<Frame>
|
||||
<img src="/images/audit-log/data-retention.png" alt="Data Retention settings showing Audit Log Retention set to 90 days" />
|
||||
<img src="/images/audit-log/data-retention.png" alt="Data retention card with three rows: Container metrics 24 hrs, Notification log 30 days, and Audit log 90 days." />
|
||||
</Frame>
|
||||
|
||||
Cleanup runs automatically as part of Sencho's periodic maintenance cycle.
|
||||
@@ -164,3 +165,23 @@ Cleanup runs automatically as part of Sencho's periodic maintenance cycle.
|
||||
## Security at rest
|
||||
|
||||
Sensitive database values (such as remote node API tokens) are encrypted at rest. The encryption key is stored separately from the database, ensuring that database file exposure alone does not compromise secrets.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="The Audit tab is missing from the sidebar">
|
||||
The tab is visible only on **Admiral**, and only to users whose role grants the `system:audit` permission. By default that means **Admin** or **Auditor**. If your license is Community or Skipper, the tab is gated by the audit-log capability and will not render. If you are signed in as a Deployer or Viewer on an Admiral instance, ask an admin to assign you the Auditor role from **Settings · Users**.
|
||||
</Accordion>
|
||||
<Accordion title="Stream view shows everything but I want to filter to a specific user, action, or date">
|
||||
Filters live in **Table view only**. Toggle the segmented control in the card header from **Stream** to **Table** and the search box, method dropdown, and From / To date pickers will appear above the grid. Switching back to Stream clears the filter strip but does not remember the last filter.
|
||||
</Accordion>
|
||||
<Accordion title="An anomaly flag did not fire on a login I expected to be flagged">
|
||||
The three flags have intentional baseline thresholds so they do not fire on incomplete data. **first seen** requires the actor to have zero entries in the prior 30 days; if you have any history at all, you will not be flagged again. **new ip** requires the actor to already have at least one stored IP in the prior 30 days; the very first IP for an actor is implicit in **first seen**, not surfaced as **new ip**. **unusual hour** requires at least 5 prior entries in the last 7 days to build a baseline; sparse actors will not flag at all until they have built up history.
|
||||
</Accordion>
|
||||
<Accordion title="Export came back smaller than the row count in the table">
|
||||
Each export is capped at 10,000 entries. If your filter selects more than that, narrow the date range using the **From** and **To** pickers and download in chunks. The cap protects the API from generating very large CSVs in a single response; for full archives, schedule periodic exports from your own tooling.
|
||||
</Accordion>
|
||||
<Accordion title="Old entries vanished even though I never deleted anything">
|
||||
Cleanup runs automatically against the **Audit log** retention value in **Settings · Developer · Data retention** (default 90 days). Entries older than the configured window are pruned on the next maintenance tick. Increase the value (up to 365 days) before the next cleanup runs to retain a longer history; the change applies forward only and cannot bring back already-pruned entries.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: "Auto-Heal Policies"
|
||||
description: "Automatically restart containers that fail Docker healthchecks."
|
||||
description: "Restart containers that fail their Docker healthcheck and stay unhealthy, with per-policy thresholds and a built-in safety rail set."
|
||||
---
|
||||
|
||||
<Note>
|
||||
@@ -9,69 +9,128 @@ description: "Automatically restart containers that fail Docker healthchecks."
|
||||
|
||||
## Overview
|
||||
|
||||
Auto-Heal Policies let you define rules that restart containers when they have been in an `unhealthy` Docker healthcheck state for longer than a specified threshold. This keeps long-running services recoverable without manual intervention.
|
||||
Auto-Heal Policies watch each container's Docker healthcheck and restart it when it has been continuously unhealthy for longer than you allow. Policies are scoped to a stack and can target every container in the stack or a single Compose service. Each policy runs with its own thresholds and four built-in safety rails so a persistently broken container cannot be restarted in a tight loop.
|
||||
|
||||
Policies live next to your stack-level alert rules in the stack's **Monitor** sheet, which has two tabs: **Alerts** and **Auto-heal**.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
- Your containers must define a `HEALTHCHECK` instruction in their `Dockerfile` or in the `healthcheck` section of your `docker-compose.yml`.
|
||||
- You must be an admin user.
|
||||
- Containers must declare a `HEALTHCHECK` in the Dockerfile or a `healthcheck` block in `docker-compose.yml`. Auto-Heal only acts on containers that report a Docker health status; a container that fails to start or exits with a non-zero code without ever reaching `unhealthy` is not in scope.
|
||||
- You must be signed in as an admin.
|
||||
- A Skipper or Admiral license.
|
||||
|
||||
## Creating a Policy
|
||||
## Workflow
|
||||
|
||||
1. In the sidebar, right-click the stack you want to protect.
|
||||
2. Select **Auto-Heal** from the context menu.
|
||||
3. In the sheet that opens, fill in the form:
|
||||
- **Service** — Select a specific service from your stack, or choose **All services** to apply the policy to every container.
|
||||
- **Unhealthy for (minutes)** — How long a container must be continuously unhealthy before it is restarted.
|
||||
- **Cooldown (minutes)** — How long to wait after a restart before evaluating the container again.
|
||||
- **Max restarts per hour** — The maximum number of times this container can be restarted within a rolling hour window.
|
||||
- **Auto-disable after (failures)** — How many consecutive failed restart attempts disable the policy automatically.
|
||||
4. Click **Add Policy**.
|
||||
1. In the sidebar, right-click the stack you want to protect (or focus the stack and press **H**).
|
||||
2. Click **Auto-Heal** in the **inspect** group. The stack's **Monitor** sheet opens directly on the **Auto-heal** tab.
|
||||
3. In **Add new policy**, pick a **Service** (or leave the combobox on **All services** to cover every container in the stack) and tune the four thresholds.
|
||||
4. Click **Add Policy**. The new policy appears in **Active policies** above the form, with its enable toggle already on.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/auto-heal-policies/policy-sheet.png" alt="Auto-Heal Policies sheet showing a policy for the web service" />
|
||||
<img src="/images/auto-heal-policies/sheet.png" alt="The Stack › PLEX › MONITOR sheet on the Auto-heal tab, showing one All-services policy in the ACTIVE POLICIES section with its ON toggle, and the ADD NEW POLICY form below with Service, Unhealthy for, Cooldown, Max restarts / hr, and Auto-disable after fields, plus the Add Policy button" />
|
||||
</Frame>
|
||||
|
||||
## Stack vs Service Scope
|
||||
You can add as many policies to a stack as you need. Each policy is evaluated independently.
|
||||
|
||||
- **All services** — The policy applies to every container in the stack that reports `unhealthy` status.
|
||||
- **Named service** — The policy targets only the containers for that specific Compose service (matched by the `com.docker.compose.service` label).
|
||||
## Form fields
|
||||
|
||||
Multiple policies can coexist on the same stack. Each policy is evaluated independently.
|
||||
| Field | Meaning |
|
||||
|-------|---------|
|
||||
| **Service** | The Compose service this policy applies to. Leave on **All services** for a stack-wide policy, or pick a specific service to limit the policy to that service's containers. |
|
||||
| **Unhealthy for (minutes)** | How long a container must be continuously reporting `unhealthy` before Auto-Heal restarts it. |
|
||||
| **Cooldown (minutes)** | After a restart fires, the policy pauses evaluation for this many minutes so the container has time to come back up. |
|
||||
| **Max restarts / hr** | The most times this policy will restart a given container in a rolling one-hour window. Once the cap is reached, further restarts are skipped until the window clears. |
|
||||
| **Auto-disable after (failures)** | If the restart call itself fails this many times in a row, the policy disables itself so it stops looping on a broken setup. The counter resets after any successful restart. |
|
||||
|
||||
## Safety Rails
|
||||
## Stack vs service scope
|
||||
|
||||
Each policy includes four built-in safety mechanisms:
|
||||
- **All services** applies the policy to every container in the stack that reports `unhealthy`. Containers are matched by the `com.docker.compose.service` label.
|
||||
- A **named service** scopes the policy to the containers for that Compose service only. Multiple policies can coexist on the same stack, so you can keep an aggressive policy on one service while running a more permissive default for the rest.
|
||||
|
||||
**Cooldown period** — After a restart is triggered, the policy pauses evaluation for the configured number of minutes. This gives the container time to recover before being evaluated again.
|
||||
## Safety rails
|
||||
|
||||
**Hourly restart cap** — If a container has been restarted the configured maximum number of times within the last hour, further restarts are skipped until the window clears. This prevents a persistently broken container from being restarted in a tight loop.
|
||||
Each policy has four safety rails built in. They run before any restart.
|
||||
|
||||
**Recent user action suppression** — If you or another operator has manually stopped or restarted the container in the last 60 seconds, the policy skips evaluation for that container. This avoids interfering with in-progress manual interventions.
|
||||
**Cooldown.** After a restart, the policy ignores the container for the configured number of minutes. This gives the container time to start up and pass its healthcheck before the next evaluation. The cooldown is per container, so other services in the same stack are not blocked.
|
||||
|
||||
**Auto-disable on repeated failures** — If the restart attempt itself fails (for example, because the Docker daemon is temporarily unavailable) the configured number of times in a row, the policy is automatically disabled. A notification is sent, and you can re-enable the policy from the sheet once the underlying issue is resolved.
|
||||
**Hourly cap.** The **Max restarts / hr** value is enforced as a rolling one-hour window per container. When the cap is reached, evaluations continue but restarts are skipped until the window clears.
|
||||
|
||||
## Policy History
|
||||
**Recent operator action.** If you or another operator manually stops or restarts the container, the policy suppresses evaluation for that container for 60 seconds. This avoids fighting an operator who is intentionally power-cycling a service.
|
||||
|
||||
Each policy row in the sheet can be expanded to show recent activity: restarts, skipped evaluations, and any auto-disable events. The history shows the container name, action taken, and the reason.
|
||||
**Auto-disable.** If the restart call itself fails (for example, the Docker daemon is not reachable, or the container is in a state that cannot be restarted), the policy increments a consecutive-failure counter. When the counter reaches the configured **Auto-disable after (failures)** value, the policy disables itself and emits a notification. A successful restart resets the counter to zero.
|
||||
|
||||
## Managing policies
|
||||
|
||||
Each row in **Active policies** shows the service it targets, a compact summary of the thresholds, and three controls on the right: an **ON/OFF** toggle, a chevron that expands the history panel, and a trash icon that deletes the policy.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/auto-heal-policies/policy-row.png" alt="Single Auto-heal policy row showing the threshold summary, ON toggle, history chevron, and delete button" />
|
||||
</Frame>
|
||||
|
||||
- **Toggle a policy off** to pause evaluation without losing its configuration. This is useful when you need to take a service down for maintenance without auto-heal interfering.
|
||||
- **Delete a policy** to remove it and its history. The policy stops evaluating immediately.
|
||||
- When a policy has accumulated restart failures but has not yet hit the auto-disable threshold, a red **N failures** pill appears next to its summary so you can see trouble building before the policy disables itself.
|
||||
|
||||
## Recent activity
|
||||
|
||||
Expand the chevron on any policy to reveal its history. Each entry shows the timestamp, the container that was acted on, a colored action label, and a one-line reason.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/auto-heal-policies/recent-activity.png" alt="Active policy row with the Recent activity panel expanded, showing No history yet" />
|
||||
</Frame>
|
||||
|
||||
The action label tells you what Auto-Heal did or, if it chose not to act, why it skipped.
|
||||
|
||||
| Action | When you see it |
|
||||
|--------|-----------------|
|
||||
| **Restarted** | A restart fired successfully. |
|
||||
| **Skipped (user action)** | The container was manually stopped or restarted in the last 60 seconds. |
|
||||
| **Skipped (cooldown)** | The policy fired recently and is still inside its cooldown window. |
|
||||
| **Skipped (rate limit)** | The hourly cap for this container has been hit. |
|
||||
| **Failed** | The restart call was attempted and failed. Counts toward the auto-disable threshold. |
|
||||
| **Auto-disabled** | The consecutive-failure threshold was reached and the policy disabled itself. |
|
||||
| **Docker unavailable** | The Docker daemon was not reachable at evaluation time. The policy is not penalized and will retry on the next 30-second tick. |
|
||||
|
||||
## Multi-node
|
||||
|
||||
Auto-heal is per node. Policies you create against a remote node are stored on and evaluated by that remote Sencho instance, so the policy keeps running even if the central node is offline. The **Monitor** sheet operates against whichever node is currently selected in the sidebar; switch nodes to see and edit that node's policies.
|
||||
|
||||
## Notifications
|
||||
|
||||
Auto-Heal dispatches a notification through your configured channels when:
|
||||
|
||||
- A restart fires successfully (info severity).
|
||||
- A restart call fails (warning).
|
||||
- A policy auto-disables itself (warning).
|
||||
|
||||
Each notification includes the stack and container names. Configure delivery channels under **Settings → Notifications**. See [Alerts & Notifications](/features/alerts-notifications) for setup details.
|
||||
|
||||
## Dashboard visibility
|
||||
|
||||
The dashboard's **Configuration status** card surfaces an **Auto-heal policies** entry that reads `N / M active` across the current node, so you can see at a glance whether stacks are covered. The entry shows **None** when no policies exist on the active node.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="The policy was auto-disabled">
|
||||
The policy disabled itself after the configured number of consecutive restart failures. Check the container logs to understand why the restart is failing. Common causes include the Docker daemon being unavailable, insufficient system resources, or a misconfigured compose file. Once the issue is resolved, re-enable the policy from the Auto-Heal sheet.
|
||||
The policy hit its **Auto-disable after (failures)** threshold. Open the policy's **Recent activity** to see which container the restart calls were failing on and what reason was reported. Common causes are a corrupted container that cannot be restarted, missing image layers after a registry change, or a stack that needs to be recreated rather than restarted. Resolve the underlying issue, then flip the policy back to **ON**.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="The policy did not fire when expected">
|
||||
Verify that the container's Docker healthcheck actually reports `unhealthy`. A container can fail to start (exit code non-zero) without ever reaching the `unhealthy` state. Use `docker inspect <container>` and check `State.Health.Status`.
|
||||
Auto-Heal only acts on Docker's `unhealthy` state. A container that crashes (non-zero exit) or never starts will not trigger a policy. Use `docker inspect <container>` and check `State.Health.Status` to confirm the container actually reaches `unhealthy`.
|
||||
|
||||
If a manual restart was performed recently, the policy suppresses evaluation for 60 seconds after the restart to avoid conflicting with operator actions.
|
||||
Evaluation runs every 30 seconds, so there is up to a 30-second delay between the **Unhealthy for (minutes)** threshold being crossed and the restart firing. The 60-second operator-action window also suppresses evaluation right after a manual stop or restart.
|
||||
</Accordion>
|
||||
|
||||
Also check that the **Unhealthy for (minutes)** threshold has elapsed. The policy evaluates every 30 seconds, so there may be up to a 30-second delay between the threshold being crossed and the restart firing.
|
||||
<Accordion title="History shows 'Docker unavailable'">
|
||||
The Docker daemon was not reachable when Auto-Heal tried to read or restart a container. This is treated as transient: the policy is not penalized, the consecutive-failures counter is not incremented, and the next 30-second tick will retry. If you see this entry repeatedly, check that the Docker socket is mounted and the daemon is running on that node.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="The container keeps getting restarted in a loop">
|
||||
If the container becomes unhealthy immediately after each restart, the hourly restart cap and the auto-disable-after-failures setting limit how many times this can occur. Review the container logs to address the root cause. You can lower the **Max restarts per hour** value or increase the **Unhealthy for (minutes)** threshold to reduce restart frequency while you investigate.
|
||||
Two safety rails are designed to stop this: the **Max restarts / hr** cap and the **Auto-disable after (failures)** threshold. If a container is unhealthy immediately after each restart, lower the cap and the failure threshold while you investigate the root cause in the container's logs. You can also raise **Unhealthy for (minutes)** so transient blips do not count against the policy.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="I cannot see the Auto-heal tab on a stack">
|
||||
The tab is hidden for Community-tier installs. Upgrade to Skipper to unlock auto-heal alongside the rest of Sencho's automation surface.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: "Auto-Update Readiness"
|
||||
description: "Review pending container updates across your fleet, with risk tags, changelogs, and rollback targets, before applying."
|
||||
description: "Review pending container updates across your fleet, with risk badges, changelogs, and scheduled run times, before applying."
|
||||
---
|
||||
|
||||
<Note>
|
||||
@@ -17,39 +17,44 @@ Auto-Update Readiness is the launchpad for every pending update across your stac
|
||||
|
||||
Each card shows:
|
||||
|
||||
- The current and next tag, with the new version highlighted in cyan.
|
||||
- A **risk tag** derived from the version delta: `patch` (safe), `minor`, `major`, or `digest rebuild`.
|
||||
- A one-line changelog preview when the registry publishes one.
|
||||
- The **rollback target** (the tag Sencho will fall back to if an update is reverted).
|
||||
- The next scheduled run for the matching auto-update task, if one exists.
|
||||
- The current tag and (for semver updates) the next tag, with the new version highlighted in cyan. For digest-only updates, the card shows the current tag with a **Rebuild available** marker instead of a version diff.
|
||||
- A **risk badge** derived from the version delta: `Safe · patch` (green), `Review · minor` (amber), `Blocked · major` (red), or `Digest rebuild` (gray) for non-semver tags.
|
||||
- The primary image reference, plus a count of additional services if more than one image in the stack has an update.
|
||||
- A one-line changelog preview when the registry publishes one. Registries that omit changelog metadata render the card with "No changelog available from the registry yet."
|
||||
- The next scheduled run for the matching auto-update task, if one exists, or "No schedule".
|
||||
|
||||
<Frame>
|
||||
<img src="/images/auto-update/readiness-board.png" alt="Readiness board with hero count, per-stack cards, and risk tags" />
|
||||
<img src="/images/auto-update/readiness-board.png" alt="Readiness board with hero counter, per-node groups, and risk badges" />
|
||||
</Frame>
|
||||
|
||||
The hero at the top counts pending updates across every node in your fleet and tells you how many of them are ready to apply without human review. Major version jumps and stacks with blocked registries are counted separately so they can be reviewed before the scheduler runs.
|
||||
The hero at the top counts pending updates across every node in your fleet and tells you how many are ready to apply without human review. Its subtitle reads `X of Y ready to apply automatically across N nodes`. Stacks with a major version bump are surfaced as a separate count (`· Z blocked by major bump`) so they can be reviewed before the scheduler runs.
|
||||
|
||||
Cards are grouped by node, with a section header for each node that has at least one pending update. The local node is listed first, followed by remote nodes alphabetically. If any of your online nodes is unreachable when the page loads, a small line under the hero shows how many of your online nodes responded.
|
||||
Cards are grouped by node, with a section header for each node that has at least one pending update. The header shows the node name, a `local` or `remote` pill, and the stack count. The local node is listed first, followed by remote nodes alphabetically. If any of your online nodes is unreachable when the page loads, a small line under the hero reads `X of Y nodes reachable. Unreachable nodes are not shown.`
|
||||
|
||||
## Empty state
|
||||
|
||||
When nothing is pending, the board renders a single Shield-icon panel with the headline "All stacks on current builds" and the sub-line "Sencho will recheck registries on the scheduler interval." Image update detection runs every six hours on each node; the readiness board reflects that cached status until the next cycle (or until you press **Recheck**).
|
||||
|
||||
## Workflow
|
||||
|
||||
1. Open the **Auto-Update** view from the sidebar.
|
||||
2. Skim the card grid. Patch and minor bumps render with a green or amber tag; major bumps render in red and are marked as blocked.
|
||||
1. Open **Auto-Update** from the top nav strip.
|
||||
2. Skim the card grid. The badge tells you the risk at a glance: `Safe · patch` is green, `Review · minor` is amber, `Blocked · major` is red, and a digest-only rebuild on a non-semver tag shows the gray `Digest rebuild` badge.
|
||||
3. For a safe update, click **Apply now** on the card to pull and recreate the stack immediately.
|
||||
4. For a major bump, review the changelog preview and the rollback target before deciding. If you still want to apply it, switch to the **Schedules** view and create or edit an auto-update task for that stack.
|
||||
5. Use **Recheck** in the hero to force an immediate registry poll across every reachable node. Per-node cooldowns still apply, and the toast tells you how many nodes were triggered, rate-limited, or failed.
|
||||
4. For a major bump, review the changelog preview and the upstream release notes. **Apply now** is disabled on the readiness board for blocked cards; to apply a major bump after review, use the stack's lifecycle **Update** action (right-click the stack in the sidebar, or open the kebab menu and choose **Update**, or click **Deploy** in the stack editor).
|
||||
5. Use **Recheck** in the hero to force an immediate registry poll across every reachable node. A 2-minute per-node cooldown applies, and the toast tells you how many nodes were triggered, rate-limited, or failed.
|
||||
|
||||
## Risk tags
|
||||
## Risk badges
|
||||
|
||||
| Tag | Meaning | Source |
|
||||
|-----|---------|--------|
|
||||
| **patch** | Safe automated update (e.g. `1.2.3` → `1.2.4`) | Semver comparison of tags |
|
||||
| **minor** | Backwards-compatible update (e.g. `1.2.3` → `1.3.0`) | Semver comparison of tags |
|
||||
| **major** | Potentially breaking change (e.g. `1.2.3` → `2.0.0`). Marked **blocked** by default | Semver comparison of tags |
|
||||
| **digest rebuild** | Same tag, new image digest (e.g. `latest` pushed again) | Local vs remote digest diff |
|
||||
| **unknown** | Non-semver tag (e.g. `main`, `stable`) | Fallback when tags cannot be compared |
|
||||
| Badge | Color | When it appears |
|
||||
|-------|-------|-----------------|
|
||||
| `Safe · patch` | Green (Shield icon) | Patch-level semver bump (e.g. `1.2.3` to `1.2.4`) |
|
||||
| `Review · minor` | Amber (AlertTriangle icon) | Minor semver bump (e.g. `1.2.3` to `1.3.0`) |
|
||||
| `Blocked · major` | Red (ShieldAlert icon) | Major semver bump (e.g. `1.2.3` to `2.0.0`). **Apply now** is disabled; the card surfaces the reason "Major version jumps require human review before applying." |
|
||||
| `Digest rebuild` | Gray | Non-semver tag (e.g. `main`, `stable`) with an updated digest |
|
||||
|
||||
Blocked updates still schedule check runs, but the apply button is disabled until you review them manually.
|
||||
A separate inline `Rebuild available` label replaces the version diff when only the digest changed (same tag, new image). The risk badge on those cards still reflects the underlying semver classification reported by the registry.
|
||||
|
||||
Blocked updates still surface in scheduled check runs so you stay informed, but the apply button is disabled until you review them manually.
|
||||
|
||||
## Per-stack control
|
||||
|
||||
@@ -62,7 +67,7 @@ Auto-updates can be disabled on a per-stack basis from the stack's context menu.
|
||||
3. The setting persists across restarts. The readiness board shows an **Auto: Off** pill on that card, and the **Apply now** button is disabled with the tooltip "Auto-updates are disabled for this stack. Update it from its actions menu."
|
||||
|
||||
<Note>
|
||||
Per-stack auto-update control requires a **Skipper** or **Admiral** license. The toggle does not appear on Community.
|
||||
Per-stack auto-update control requires a **Skipper** or **Admiral** license.
|
||||
</Note>
|
||||
|
||||
### What disabling means
|
||||
@@ -89,7 +94,7 @@ The task lives alongside restart, prune, snapshot, and scan tasks in the same ti
|
||||
|
||||
The readiness board shows pending updates from every node in your fleet in a single view, regardless of which node is selected in the sidebar. You do not need to switch nodes to inspect what is pending elsewhere.
|
||||
|
||||
Each node group renders its own card grid. **Apply now** runs on the node that owns the stack, and **Recheck** fans out to every reachable node so registries get polled in parallel. Sencho handles the routing through the Distributed API; no additional configuration is needed.
|
||||
Each node group renders its own card grid under a header showing the node name, a `local` or `remote` pill, and a stack count. **Apply now** runs on the node that owns the stack, and **Recheck** fans out to every reachable node so registries get polled in parallel. Sencho handles the routing through the Distributed API; no additional configuration is needed.
|
||||
|
||||
## How readiness is computed
|
||||
|
||||
@@ -97,34 +102,31 @@ For each stack with a pending image update, Sencho computes a preview by:
|
||||
|
||||
1. Parsing the compose file to enumerate every pullable image reference.
|
||||
2. Calling the registry with your configured credentials to fetch the current tag list and remote digest.
|
||||
3. Picking the highest semver tag greater than the current tag (keeping the same prefix and suffix) or, if tags match but digests differ, treating it as a digest rebuild.
|
||||
3. Picking the highest semver tag greater than the current tag (keeping the same prefix and suffix). If the highest available tag matches the current one but the remote digest has changed, the card surfaces as a `Rebuild available` update.
|
||||
4. Scoring the overall stack by the most severe image bump. Any major bump marks the stack as blocked.
|
||||
5. Deriving the rollback target from the running tag and normalizing Docker Hub library paths.
|
||||
5. Normalizing Docker Hub library paths so credentials and changelog lookups resolve correctly.
|
||||
|
||||
The preview is recomputed each time the readiness board loads, so it reflects the live state of your registries and local images.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### Card shows "No changelog available"
|
||||
|
||||
Sencho reads changelog metadata from the registry's manifest and OCI annotations. Registries that do not publish this metadata (most private registries and many self-hosted ones) will simply render the card without a changelog. The risk tag is still accurate because it is computed from the tag itself.
|
||||
|
||||
### Apply button is disabled with a "blocked" tooltip
|
||||
|
||||
The stack has a major version bump. Open the cron schedule for that stack and apply manually after reviewing the upstream release notes. The block is a policy decision: major updates never auto-apply without human review.
|
||||
|
||||
### Card stays stuck on "Checking"
|
||||
|
||||
The registry call is either still pending or failed. Click **Recheck** in the hero to retry. If the stack has private-registry credentials, confirm they are still valid in **Settings > Registries**.
|
||||
|
||||
### "Nothing to update" but I see an update on another view
|
||||
|
||||
Image update detection runs every six hours on each node. The readiness board uses the same cached status. Trigger **Recheck** to force a fresh check across every reachable node, or see [Image update detection](/features/image-update-detection) for details on the refresh cycle.
|
||||
|
||||
### Scheduled auto-update runs are not applying to a specific stack
|
||||
|
||||
The stack likely has auto-updates disabled. Open the stack's kebab menu or right-click context menu and check the **Inspect** group. If the item reads **Auto-update: Disabled**, click it to re-enable. Once re-enabled, the next scheduled run will include the stack, or you can trigger an immediate run from the Auto-Update view.
|
||||
|
||||
### Banner says "X of Y nodes reachable"
|
||||
|
||||
One or more nodes that are marked online in your fleet did not respond within the request timeout. Pending updates from those nodes are not shown until they come back. Check the node's status from the Fleet view and the network path between this Sencho instance and the unreachable node.
|
||||
<AccordionGroup>
|
||||
<Accordion title='Card shows "No changelog available"'>
|
||||
Sencho reads changelog metadata from the registry's manifest and OCI annotations. Registries that do not publish this metadata (most private registries and many self-hosted ones) render the card without a changelog. The risk badge is still accurate because it is computed from the tag itself.
|
||||
</Accordion>
|
||||
<Accordion title='Apply now is disabled with a "Blocked · major" tooltip'>
|
||||
The stack has a major version bump and is blocked on the readiness board by policy: major updates never auto-apply without human review. To apply after reviewing the upstream release notes, use the stack's lifecycle **Update** action from the sidebar kebab or right-click menu, or open the stack editor and click **Deploy**.
|
||||
</Accordion>
|
||||
<Accordion title='Card stays stuck on "Checking registry..."'>
|
||||
The registry call is either still pending or it failed. Click **Recheck** in the hero to retry. If the stack uses private-registry credentials, confirm they are still valid in **Settings > Registries**.
|
||||
</Accordion>
|
||||
<Accordion title='"Nothing to update" but I see an update on another view'>
|
||||
Image update detection runs every six hours on each node and the readiness board uses the same cached status. Trigger **Recheck** to force a fresh check across every reachable node.
|
||||
</Accordion>
|
||||
<Accordion title='Scheduled auto-update runs are not applying to a specific stack'>
|
||||
The stack likely has auto-updates disabled. Open the stack's kebab menu or right-click context menu and check the **Inspect** group. If the item reads **Auto-update: Disabled**, click it to re-enable. Once re-enabled, the next scheduled run will include the stack, or you can trigger an immediate run from the Auto-Update view.
|
||||
</Accordion>
|
||||
<Accordion title='Banner says "X of Y nodes reachable"'>
|
||||
One or more nodes that are marked online in your fleet did not respond within the request timeout. Pending updates from those nodes are not shown until they come back. Check the node's status from the Fleet view and the network path between this Sencho instance and the unreachable node.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
@@ -1,217 +1,400 @@
|
||||
---
|
||||
title: "Blueprints"
|
||||
description: "Fleet-wide compose templates that Sencho keeps in sync across the nodes you choose."
|
||||
description: "Fleet-wide compose templates that Sencho keeps in sync across the nodes you choose, with stateless or stateful classification, drift detection, and three response modes."
|
||||
---
|
||||
|
||||
A **Blueprint** is a docker-compose.yml plus a node selector. Sencho ensures that every node matching the selector runs that stack at the latest revision and reports back when reality drifts from the plan. You declare a stack once; Sencho handles the distribution.
|
||||
A **Blueprint** bundles a `docker-compose.yml` with a node selector and a drift policy. Every minute Sencho compares the live state on each targeted node against the blueprint's desired revision, and either reports the drift, dispatches a notification, or auto-corrects, depending on the mode you picked. You declare a stack once; Sencho handles the distribution and the reconciliation loop.
|
||||
|
||||
Blueprints live under **Fleet → Deployments**.
|
||||
Blueprints live under **Fleet · Deployments**.
|
||||
|
||||
<Note>
|
||||
Blueprints are a Skipper feature. Configuring blueprints requires an admin role; viewers see the catalog read-only.
|
||||
Blueprints require a Sencho **Skipper** or **Admiral** license. Creating, editing, and withdrawing blueprints requires an admin role; operators and viewers can read the catalog and the detail sheet. Pinning a blueprint to a single node requires Admiral.
|
||||
</Note>
|
||||
|
||||
<Frame>
|
||||
<img src="/images/blueprint-model/catalog.png" alt="Blueprint catalog in Fleet Deployments showing blueprint cards, deployment counts, drift policy, and the New Blueprint action" />
|
||||
<Frame caption="Fleet · Deployments catalog with blueprint tiles, the All / Drifted / Observe / Suggest / Enforce filter chips, and the New Blueprint action in the top-right.">
|
||||
<img src="/images/blueprint-model/catalog.png" alt="Blueprint catalog in Fleet Deployments" />
|
||||
</Frame>
|
||||
|
||||
## What problem this solves
|
||||
## Mental model
|
||||
|
||||
Without Blueprints, running the same stack on multiple nodes means SSHing or clicking through each node's stack manager and keeping them in sync by hand. When something drifts (someone restarts a container, edits a compose, or a node forgets to pull a new image) you find out when it breaks.
|
||||
Three moving parts cooperate per blueprint.
|
||||
|
||||
With Blueprints you get:
|
||||
1. **The declared spec.** A blueprint is a row in Sencho's database. It carries the compose YAML, the selector (labels or explicit node IDs), the drift policy, and a monotonic revision number. The revision auto-increments every time the compose changes.
|
||||
2. **The reconciler.** A background loop on the controlling Sencho ticks every 60 seconds (with a 5-second initial delay after startup) and on demand via **Apply now**. Each tick: resolve the selector, compare every desired-vs-live node, and queue one of five actions per node (deploy, withdraw, drift-check, request operator confirmation, or block on operator confirmation).
|
||||
3. **The executor.** Per-node deploy and withdraw run against the local Docker socket on local nodes, and through the standard authenticated proxy to `/api/stacks` on remote nodes. Every blueprint deployment writes a `.blueprint.json` marker into the stack directory; the marker carries the blueprint ID, the revision, and the last-applied timestamp.
|
||||
|
||||
- **One declaration covers many nodes.** Pick nodes by label (`production`) or by ID. The set is recomputed every reconciliation tick, and adding a node with the right label deploys the stack automatically.
|
||||
- **Drift detection always on.** Every tick, Sencho compares each target node's actual state to the desired one. You choose what happens when drift is found.
|
||||
- **Safety rails for stateful workloads.** Blueprints are classified as stateless or stateful at author time; stateful blueprints get explicit confirmation prompts before first deploy and before eviction.
|
||||
The marker is the trust root. If a directory by the blueprint's name already exists on a node and does not carry a matching marker, the reconciler refuses to touch it and surfaces a **Name conflict** on the deployment row. A Blueprint named `nginx` will never overwrite an existing user-authored `nginx` stack on any node.
|
||||
|
||||
Statelessness vs statefulness is decided at author time by parsing the compose file. Stateless blueprints (no persistent volumes, or only `tmpfs` mounts) deploy and evict freely. Stateful blueprints (named volumes or bind mounts) get explicit operator-confirmation prompts on the first deploy to a fresh node and on eviction from any node. Blueprints with `external: true` volumes are classified as **unknown** and treated as stateful for safety.
|
||||
|
||||
Drift detection runs on every tick for every Active deployment regardless of policy. The policy only governs what Sencho does next: surface the drift silently, notify, or auto-redeploy.
|
||||
|
||||
## Key capabilities
|
||||
|
||||
**One declaration covers many nodes.** Pick nodes by label or by node ID. The selector set is re-resolved on every reconciliation tick, so adding a node with a matching label deploys the stack within one minute. Removing a label, removing a node, or changing the selector withdraws the deployment on the same cadence (subject to the stateful-eviction safety rail).
|
||||
|
||||
**Drift detection always on.** Each tick compares the marker's revision against the live containers and the blueprint's current revision, then checks that all containers labeled with the compose project name are running. Drift is recorded on the deployment row whatever the policy is. **Observe** records it silently, **Suggest** also dispatches a notification, **Enforce** also redeploys.
|
||||
|
||||
**Stateful safety rails.** Stateful and state-unknown blueprints enter **Awaiting confirmation** on every fresh node before the first deploy, and **Evict blocked** when a node falls out of the selector while still hosting a stateful deployment. The reconciler never deploys empty volumes or destroys named volumes without a human acknowledging the action.
|
||||
|
||||
**Pin override and cordon respect.** Admiral users can pin a blueprint to a single node from **Fleet · Federation**. A pin replaces the selector entirely, deploys only to the pinned node, and overrides the cordon flag on that node. Cordoning a node otherwise prevents the reconciler from picking it for new placements; existing deployments on a cordoned node keep running and stay drift-checked.
|
||||
|
||||
**Vulnerability-policy participation.** Local blueprint deploys evaluate against the same pre-deploy policy gate that the per-stack deploy lane uses. If an enabled policy blocks one of the blueprint's image references, the deployment row moves to **Failed** and the stack is never written to disk. Remote blueprint deploys are routed through the remote node's stack deploy endpoint, so policy enforcement runs on the remote instance with that node's credentials and scanner state.
|
||||
|
||||
**Per-node, per-revision marker ownership.** Every blueprint deployment writes a `.blueprint.json` marker carrying the blueprint ID, revision, and last-applied timestamp. The reconciler refuses to touch any directory whose marker does not match its blueprint ID, so user-authored stacks and other blueprints can coexist on the same node without collision.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
| Requirement | Detail |
|
||||
|---|---|
|
||||
| License tier | **Skipper** or **Admiral** to read, create, edit, and withdraw blueprints. **Admiral** to pin a blueprint to a node from the Federation tab. |
|
||||
| User role | **Admin** to create, edit, withdraw, accept, and pin. Operators and viewers can read the catalog and the detail sheet. |
|
||||
| Nodes | At least one node that the selector resolves to. Remote nodes need a healthy proxy connection; see [Multi-node management](/features/multi-node) and [Pilot Agent](/features/pilot-agent) for enrollment. |
|
||||
| Compose YAML | Valid `docker-compose.yml`, 96 KiB or fewer. |
|
||||
| Blueprint name | 1 to 64 characters matching `^[a-z0-9][a-z0-9_-]*$`. The name doubles as the stack directory on every targeted node and is immutable after creation. |
|
||||
| Selector | A `labels` or `nodes` selector with up to 200 entries per side. An empty resolved set is allowed but produces no deployments. |
|
||||
| Compose directory | Per-node compose directory must be writable. Remote nodes must accept the controlling instance's bearer token; this is the same channel the rest of the fleet management already uses. |
|
||||
|
||||
## The catalog
|
||||
|
||||
The Deployments tab opens on a catalog of every blueprint configured on this instance. Each tile shows the name, an optional description, the classification chip (stateless, stateful, or unknown), the active-vs-targeted node count, and the current drift mode. Filter chips at the top break the catalog down by drift mode and let you isolate any tile that has drift.
|
||||
|
||||
<Frame caption="Catalog tiles. Each tile carries the classification chip, the active-over-targeted count, the selector summary, and the drift mode. The filter chips above the grid pivot the view.">
|
||||
<img src="/images/blueprint-model/catalog.png" alt="Deployments tab catalog with classification chips and filter chips" />
|
||||
</Frame>
|
||||
|
||||
The first time you visit the tab, the catalog is empty and a three-step explainer walks you through Author, Target, and Reconcile, with a **Create your first Blueprint** button at the bottom.
|
||||
|
||||
<Frame caption="Empty Deployments tab. Three numbered cards (01 Author, 02 Target, 03 Reconcile) explain the model before you create the first blueprint.">
|
||||
<img src="/images/blueprint-model/empty-state.png" alt="Empty Deployments tab with three-step explainer" />
|
||||
</Frame>
|
||||
|
||||
## Anatomy of a Blueprint
|
||||
|
||||
| Field | Purpose |
|
||||
|---|---|
|
||||
| **Name** | Used as the stack directory on every targeted node (`<COMPOSE_DIR>/<blueprint-name>/`). Lowercase letters, digits, hyphens, and underscores. |
|
||||
| **Description** | Short prose for the catalog tile and detail header. |
|
||||
| **Compose** | Standard `docker-compose.yml`. The same file ships to every targeted node. The YAML must parse successfully and stay under 96 KiB. |
|
||||
| **Selector** | Either `labels` (any/all expressions) or a list of node IDs. |
|
||||
| **Drift policy** | Observe, Suggest, or Enforce. See below. |
|
||||
| **Reconciler enabled** | Toggle the reconciliation loop without deleting the blueprint. |
|
||||
| **Name** | Used as the stack directory on every targeted node (`<COMPOSE_DIR>/<blueprint-name>/`). Lowercase letters, digits, hyphens, and underscores only. Fixed once the blueprint exists. |
|
||||
| **Description** | Short prose for the catalog tile and the detail header. |
|
||||
| **Compose** | Standard `docker-compose.yml`. The same file ships to every targeted node. Sencho parses it on save and classifies the blueprint as stateless, stateful, or unknown. The YAML must parse successfully and stay under 96 KiB. |
|
||||
| **Selector** | Either label expressions (any-of plus all-of) or a list of node IDs picked by hand. |
|
||||
| **Drift policy** | Observe, Suggest, or Enforce. Drift detection always runs; only the response differs. |
|
||||
| **Reconciler enabled** | Toggle the reconciliation loop without deleting the blueprint. The **Disable** button on the detail sheet flips this toggle off; **Enable** turns it back on. Useful when you want to pause auto-deploys without losing the configuration. |
|
||||
|
||||
Sencho writes a `.blueprint.json` marker into each targeted node's stack directory. The marker carries the blueprint ID, revision, and the timestamp of the last apply. The reconciler refuses to touch any directory that does not carry a matching marker, so a Blueprint named `nginx` will never overwrite an existing user-authored `nginx` stack on any node.
|
||||
Sencho writes a `.blueprint.json` marker into each targeted node's stack directory. The marker carries the blueprint ID, revision, and the timestamp of the last apply. The reconciler refuses to touch any directory that does not carry a matching marker, so a Blueprint named `nginx` will never overwrite an existing user-authored `nginx` stack on any node. When the directory exists without a matching marker, the deployment row enters the **Name conflict** status until you rename one of the two.
|
||||
|
||||
Before a local Blueprint deploy starts, Sencho applies the same pre-deploy vulnerability policy gate used by standard stack deploys. If an enabled policy blocks one of the Blueprint's image references, the deployment row moves to failed and the stack is not written to disk. Successful local Blueprint deploys also trigger the normal post-deploy scan. Remote Blueprint deploys are routed through the remote node's stack deploy endpoint, so policy enforcement runs on the remote instance with that node's credentials and scanner state.
|
||||
Before a local Blueprint deploy starts, Sencho applies the same pre-deploy [vulnerability scan](/features/vulnerability-scanning) policy gate used by standard stack deploys. If an enabled policy blocks one of the Blueprint's image references, the deployment row moves to **Failed** and the stack is not written to disk. Successful local Blueprint deploys also trigger the normal post-deploy scan. Remote Blueprint deploys are routed through the remote node's stack deploy endpoint, so policy enforcement runs on the remote instance with that node's credentials and scanner state.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/blueprint-model/editor-dialog.png" alt="New Blueprint editor dialog with name, compose YAML, selector, drift policy, and classification banner" />
|
||||
## The editor
|
||||
|
||||
**New Blueprint** opens a single editor dialog with everything inline: name and description fields, a Monaco YAML editor with the classification banner sitting above it, the selector picker (Labels or Specific nodes), the drift policy cards, and the Reconciler enabled toggle.
|
||||
|
||||
<Frame caption="The Blueprint editor. Compose YAML on the left with the classification banner pinned above it; selector mode tabs, drift policy cards, and the Reconciler enabled toggle below.">
|
||||
<img src="/images/blueprint-model/editor-dialog.png" alt="Blueprint editor dialog with classification banner, selector, and drift policy" />
|
||||
</Frame>
|
||||
|
||||
## Selectors
|
||||
### Classification banner
|
||||
|
||||
A **labels** selector matches any node whose labels satisfy the expression:
|
||||
The banner above the YAML editor shows what Sencho found when it parsed your compose file:
|
||||
|
||||
```
|
||||
all = [docker]
|
||||
any = [production, staging]
|
||||
```
|
||||
- **Stateless · portable** when no persistent volumes were detected, or only `tmpfs` mounts. Sencho can deploy and evict the blueprint freely on any node.
|
||||
- **Stateful · pins to data** when named volumes or bind mounts were detected. Each node holds its own data; Sencho does not replicate volumes between nodes.
|
||||
- **State unknown** when `external: true` volumes were detected. Sencho cannot prove portability and treats the blueprint as stateful for safety.
|
||||
|
||||
This resolves to nodes that have *every* label in `all` AND *at least one* label in `any`. Either side may be empty. An entirely empty labels selector matches nothing, so choose at least one label.
|
||||
Click the signals toggle on the right of the banner (`1 SIGNAL`, `2 SIGNALS`, etc.) to expand the list of reasons Sencho classified the way it did, line by line.
|
||||
|
||||
A **nodes** selector picks specific node IDs by hand. Useful when you want a one-off blueprint that runs only on a known node.
|
||||
<Frame caption="The classification banner expanded. The signals list shows the exact rule that fired for each detected mount, volume, or external reference.">
|
||||
<img src="/images/blueprint-model/classification-banner.png" alt="Classification banner in the Stateless · portable variant with one signal expanded" />
|
||||
</Frame>
|
||||
|
||||
Add labels to nodes from **Settings → Nodes**. Each node row has a Labels column with a `+` button.
|
||||
### Selector
|
||||
|
||||
## Drift policy
|
||||
A **Labels** selector matches any node whose labels satisfy the expression. The picker is split into two pill rows: **Match nodes with ANY of these labels** (the OR set), and **AND ALSO require ALL of these** (the AND set). Either side may be empty. The dialog renders a one-line summary just below the pills, so you can see exactly which nodes the expression resolves to before you save.
|
||||
|
||||
Drift detection runs every minute regardless of the policy. Only the response differs:
|
||||
A **Specific nodes** selector picks nodes by ID with a checkbox list. Use it when you want a one-off blueprint that runs only on a known node, or when no labels exist yet.
|
||||
|
||||
| Mode | What happens on drift |
|
||||
Add labels to nodes from **Settings · System · Nodes**. Each node row carries a Labels column with a `+` button.
|
||||
|
||||
<Frame caption="Settings · System · Nodes. The Labels column carries pills per node; the `+` button opens an inline add-label popover.">
|
||||
<img src="/images/blueprint-model/node-labels.png" alt="Settings System Nodes page with the Labels column and add-label popover" />
|
||||
</Frame>
|
||||
|
||||
### Drift policy
|
||||
|
||||
The three policy cards in the editor map exactly to the three modes the reconciler runs against every tick:
|
||||
|
||||
| Card label | Reconciler behavior |
|
||||
|---|---|
|
||||
| **Observe** | Drift surfaces in the deployment table; no notification, no auto-fix. |
|
||||
| **Suggest** (default) | Sencho dispatches a `blueprint_drift_detected` notification through your notification routes, if any. |
|
||||
| **Enforce** | Sencho re-deploys the blueprint silently when drift is detected. A notification fires only when an auto-fix attempt fails. |
|
||||
| **Observe** · Detect & display, no notifications | Drift surfaces in the deployment table; no notification, no auto-fix. |
|
||||
| **Suggest** · Detect & notify, operator decides | Sencho dispatches a `blueprint_drift_detected` notification through your notification routes, if any. |
|
||||
| **Enforce** · Detect & auto-fix, silent on success | Sencho re-deploys the blueprint silently when drift is detected. A `blueprint_drift_correction_failed` notification fires only when the auto-fix attempt fails. |
|
||||
|
||||
Even **Observe** keeps Sencho honest about what it found. The deployment row shows "drifted 3h ago: service caddy exited code 1". Silence would forfeit Sencho's authority over your fleet.
|
||||
Even **Observe** keeps Sencho honest about what it found: the deployment row shows "drifted 3h ago: service caddy exited code 1". Silence would forfeit Sencho's authority over your fleet.
|
||||
|
||||
For **stateful** blueprints under Enforce, Sencho declines auto-fixes that would destroy named volumes (for example, when you rename a volume in the compose). The drift downgrades to Suggest semantics for that event with the reason `auto-fix declined: would destroy volume data`.
|
||||
|
||||
## Stateless vs Stateful Blueprints
|
||||
## The detail sheet
|
||||
|
||||
Sencho classifies your compose at author time:
|
||||
Click any tile in the catalog to open the blueprint detail sheet on the right. The sheet header shows the name, the meta line (`<selector> · <drift mode> · rev <n>`), the action bar (Apply now, Edit, Disable, Delete), the description, the deployment table, and a collapsible compose source.
|
||||
|
||||
- **Stateless**: no persistent volumes detected, or only `tmpfs`. Sencho can deploy and evict freely.
|
||||
- **Stateful**: named volumes or bind mounts detected. Each node holds its own data; Sencho does not replicate volumes between nodes.
|
||||
- **State unknown**: `external: true` volumes detected. Sencho cannot prove portability and treats the blueprint as stateful for safety.
|
||||
|
||||
The classification appears as a chip on the catalog tile and as a banner above the YAML editor. Click the banner to see exactly what made Sencho classify the way it did.
|
||||
|
||||
### Safety rails on stateful blueprints
|
||||
|
||||
| Trigger | What Sencho does |
|
||||
|---|---|
|
||||
| Selector matches a node that has never run this blueprint | Deployment enters `pending_state_review`. The reconciler refuses to deploy until you click **Confirm deploy** in the deployment table. |
|
||||
| A stateful or unknown blueprint revision changes on a node that already runs it | Deployment enters `pending_state_review`. Confirm the redeploy before Sencho writes the new compose revision. |
|
||||
| A node leaves the selector while a deployment is active | Deployment enters `evict_blocked`. The reconciler refuses to evict until you choose **Snapshot, then evict** or **Evict and destroy data**. |
|
||||
| You target more than one node | The editor warns: "Each node will hold its own data. Sencho does not replicate volumes between nodes." |
|
||||
|
||||
Stateless blueprints flow through these states automatically.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/blueprint-model/detail-state-review.png" alt="Blueprint detail sheet with a deployment row waiting for state review confirmation" />
|
||||
<Frame caption="The detail sheet. Header carries the selector and revision; the action bar offers Apply now, Edit, Disable, and Delete; the deployment table lists one row per resolved node.">
|
||||
<img src="/images/blueprint-model/detail-sheet.png" alt="Blueprint detail sheet showing header, action bar, and deployment table" />
|
||||
</Frame>
|
||||
|
||||
The deployment table lists every node currently in the blueprint's resolved selector set. Each row carries the node name, the deployment status, the last-activity timestamp, a notes column (drift summary or the last error message), and the action available for that row.
|
||||
|
||||
Possible status values:
|
||||
|
||||
| Status | Meaning |
|
||||
|---|---|
|
||||
| **Pending** | The reconciler has acquired the deployment lock and is about to start. |
|
||||
| **Awaiting confirmation** | A stateful blueprint reached a new node. The reconciler refuses to deploy until you click **Confirm deploy** in the action column. |
|
||||
| **Deploying** | The compose action is in flight on this node. |
|
||||
| **Active** | The container set matches the blueprint's revision. |
|
||||
| **Drifted** | The reconciler detected a difference between the live state and the desired one. |
|
||||
| **Correcting** | An Enforce-mode auto-fix is in flight. |
|
||||
| **Failed** | The deployment or withdrawal errored. The notes column carries the message. |
|
||||
| **Withdrawing** | A withdraw is in flight. |
|
||||
| **Withdrawn** | The deployment was successfully removed; the row is about to be cleared. |
|
||||
| **Evict blocked** | A node left the selector while a stateful deployment was active. The reconciler refuses to evict until you choose an explicit mode. |
|
||||
| **Name conflict** | A directory by the blueprint's name already exists on the node and does not carry our marker file. The reconciler will not touch it. |
|
||||
|
||||
<Frame caption="A deployment row in Awaiting confirmation. The Confirm deploy action sits at the right; the notes column explains why the reconciler is waiting on the operator.">
|
||||
<img src="/images/blueprint-model/detail-state-review.png" alt="Blueprint detail sheet with a deployment row in Awaiting confirmation state" />
|
||||
</Frame>
|
||||
|
||||
## Lifecycle: status transitions
|
||||
|
||||
The common paths through the status enum.
|
||||
|
||||
- **Stateless first deploy.** No row → **Pending** → **Deploying** → **Active**. Drift detected later on Enforce flips to **Drifted** → **Correcting** → **Active**.
|
||||
- **Stateful first deploy.** No row → **Awaiting confirmation**. Operator clicks **Confirm deploy** and picks **Deploy fresh** → **Deploying** → **Active**.
|
||||
- **Compose change on a stateless deployment.** **Active** → **Deploying** on the next tick → **Active**.
|
||||
- **Compose change on a stateful or unknown deployment.** **Active** → **Awaiting confirmation** on every targeted node so the operator can decide whether the new revision is safe for that node's local data. Operator confirms → **Deploying** → **Active**.
|
||||
- **Node leaves a stateless selector.** **Active** → **Withdrawing** → **Withdrawn** and the row clears.
|
||||
- **Node leaves a stateful selector.** **Active** → **Evict blocked**. Operator chooses **Snapshot, then evict** or **Evict and destroy data** → **Withdrawing** → **Withdrawn**.
|
||||
- **Pre-existing stack collision.** New row → **Name conflict**. No further action without operator rename of the user-authored stack or of the blueprint itself.
|
||||
- **Deploy or withdraw error.** Any in-flight status → **Failed** with the message in the notes column. The row stays put until the next tick or until the operator clicks **Apply now** to retry.
|
||||
|
||||
## Working with Blueprints
|
||||
|
||||
### Create
|
||||
|
||||
1. Go to **Fleet → Deployments**.
|
||||
1. Go to **Fleet · Deployments**.
|
||||
2. Click **New Blueprint**.
|
||||
3. Fill in the name, description, compose YAML, selector, and drift policy.
|
||||
4. Watch the classification banner update as you type. It tells you whether the blueprint is portable or pinned.
|
||||
4. Watch the classification banner update as you type. It tells you whether the blueprint is portable or pinned to data.
|
||||
5. Click **Create blueprint**. Sencho immediately runs one reconciliation tick.
|
||||
|
||||
If the YAML is malformed or larger than 96 KiB, Sencho rejects the save before creating a Blueprint row.
|
||||
|
||||
### Apply on demand
|
||||
|
||||
The reconciler runs every minute. To trigger it now (for example, after editing the selector or compose), click **Apply now** on the detail sheet.
|
||||
The reconciler runs every minute. To trigger it now, for example after editing the selector or the compose file, click **Apply now** on the detail sheet. The action also resurfaces a deployment that is in **Failed** status by retrying it.
|
||||
|
||||
### Edit
|
||||
|
||||
Click **Edit** on the detail sheet. Editing the compose bumps the revision; the reconciler will redeploy on every targeted node on the next tick. Stateful blueprints follow the volume-destroying drift rule under Enforce.
|
||||
Click **Edit** on the detail sheet. Editing the compose bumps the revision. Stateless blueprints redeploy on the next reconciliation tick. Stateful and state-unknown blueprints re-enter **Awaiting confirmation** on every targeted node so the operator can decide whether the new revision is safe for each node's local data. Stateful blueprints also follow the volume-destroying-drift rule under Enforce: a compose change that would destroy named volumes downgrades to Suggest for that event.
|
||||
|
||||
For stateful or state-unknown Blueprints, a compose edit does not redeploy automatically. Each existing deployment enters **Awaiting confirmation** so you can decide whether the new revision is safe for that node's local data.
|
||||
### Confirm a stateful first deploy
|
||||
|
||||
<Frame>
|
||||
<img src="/images/blueprint-model/detail-sheet.png" alt="Blueprint detail sheet showing deployment rows, Apply now, Edit, Delete, and the current compose content" />
|
||||
The first time a stateful blueprint reaches a node, the deployment row enters **Awaiting confirmation**. Click **Confirm deploy** in the action column to open the state review dialog. **Deploy fresh** creates empty named volumes on the target node and starts the stack with the default state from each container image. **Restore from snapshot** is reserved for the future Volume Migration feature and is currently disabled.
|
||||
|
||||
<Frame caption="State review dialog. Deploy fresh creates empty named volumes on the target node; Restore from snapshot is reserved for the future Volume Migration feature.">
|
||||
<img src="/images/blueprint-model/state-review-dialog.png" alt="State review dialog with Deploy fresh enabled and Restore from snapshot disabled" />
|
||||
</Frame>
|
||||
|
||||
### Withdraw a single deployment
|
||||
|
||||
In the deployment table, click **Withdraw** on the node's row. For stateless blueprints, Sencho runs `docker compose down` and removes the directory. For stateful blueprints, you choose between **Snapshot, then evict** (records the compose definition to Fleet → Snapshots, then evicts) and **Evict and destroy data** (typed-confirm, destroys named volumes).
|
||||
In the deployment table, click **Withdraw** on the node's row. For stateless blueprints, Sencho confirms once and runs `docker compose down` plus a directory removal. For stateful blueprints, the eviction dialog asks how to handle the data on this node:
|
||||
|
||||
- **Snapshot, then evict (recommended)** captures the blueprint's compose definition into [Fleet · Snapshots](/features/fleet-backups) before running `docker compose down`. The volume bytes still leave the node when compose tears down the named volumes; the snapshot only preserves the YAML so you can redeploy it elsewhere.
|
||||
- **Evict and destroy data** runs the eviction without a snapshot. Type the blueprint name to confirm.
|
||||
|
||||
<Frame caption="Stateful eviction dialog. Snapshot, then evict captures the compose YAML to Fleet · Snapshots; Evict and destroy data requires typing the blueprint name to confirm.">
|
||||
<img src="/images/blueprint-model/eviction-dialog.png" alt="Stateful eviction dialog with Snapshot then evict and Evict and destroy data options" />
|
||||
</Frame>
|
||||
|
||||
<Frame caption="Stateless eviction dialog. A single Withdraw deployment action; no data prompt because nothing persistent was detected.">
|
||||
<img src="/images/blueprint-model/eviction-stateless.png" alt="Stateless eviction dialog with a single Withdraw deployment action" />
|
||||
</Frame>
|
||||
|
||||
<Note>
|
||||
**Snapshot, then evict** captures the compose definition only. Volume bytes are not shipped. The named volumes managed by this stack on the target node are removed by `docker compose down` just as with **Evict and destroy data**. To preserve data, capture volumes manually before withdrawing (see *Migrating stateful data between nodes* below).
|
||||
**Snapshot, then evict** captures the compose definition only. Volume bytes are not shipped. The named volumes managed by this stack on the target node are removed by `docker compose down` just as with **Evict and destroy data**. To preserve data, capture volumes manually before withdrawing (see *Migrating stateful data between nodes* below).
|
||||
</Note>
|
||||
|
||||
### Delete the blueprint
|
||||
|
||||
Stateless blueprints withdraw all deployments and then delete. Stateful blueprints with active deployments refuse to delete. Withdraw each deployment explicitly first.
|
||||
Stateless blueprints withdraw all deployments and then delete in a single click. Stateful blueprints with active or pending deployments refuse to delete to avoid silent orphans. Withdraw each deployment explicitly first, then delete.
|
||||
|
||||
## Federation: pin a blueprint to a single node
|
||||
|
||||
Admiral users can pin a blueprint to a specific node from **Fleet · Federation**. A pinned blueprint deploys only to its pinned node, regardless of the configured selector, and overrides the cordon flag on that node. The Blueprint detail sheet shows a read-only `Pin` section when a pin is in place; pin management itself lives in the Federation tab.
|
||||
|
||||
<Frame caption="Fleet · Federation, Blueprints subsection. Each row shows the blueprint, its configured selector, the Pinned to dropdown, and the effective placement that the reconciler will use.">
|
||||
<img src="/images/blueprint-model/federation-pin.png" alt="Federation tab pin policy table with one blueprint pinned and one unpinned" />
|
||||
</Frame>
|
||||
|
||||
See [Fleet Federation](/features/fleet-federation) for the cordon and pin model.
|
||||
|
||||
## Cordoned nodes
|
||||
|
||||
Cordoning a node prevents the reconciler from selecting it for **new** placements and skips state-review provisioning on it. Existing deployments on a cordoned node keep running, and drift checks continue normally. Pin policy overrides cordon, so a blueprint pinned to a cordoned node still deploys.
|
||||
|
||||
## Notifications
|
||||
|
||||
Two notification events fire from the reconciler:
|
||||
|
||||
- `blueprint_drift_detected` (warning severity), in Observe and Suggest modes when drift is detected. In Observe the row updates silently in the table; in Suggest the notification is also dispatched through your notification routes.
|
||||
- `blueprint_drift_correction_failed` (error severity), in Enforce mode when an auto-fix attempt fails.
|
||||
|
||||
Both events route through the standard alert pipeline. Configure delivery channels under [Notification routing](/features/alerts-notifications#notification-routing).
|
||||
|
||||
## Security and trust boundaries
|
||||
|
||||
**Who can do what.** The license tier and the user role together determine the available actions. Reading the catalog, the detail sheet, and the deployment status requires Skipper or Admiral. Creating, editing, withdrawing, accepting a stateful deploy, and applying on demand require the admin role on top of the tier. Pinning a blueprint requires Admiral plus the admin role.
|
||||
|
||||
**The marker file is the trust root.** The reconciler will only deploy into, modify, or withdraw a directory that carries a `.blueprint.json` marker whose blueprint ID matches. A pre-existing directory with no marker, or a marker referencing a different blueprint, surfaces as **Name conflict** and is never modified.
|
||||
|
||||
**Local vs remote policy enforcement.** Local blueprint deploys evaluate the pre-deploy vulnerability policy gate against the local scanner state and credentials before the compose file is written to disk. Remote blueprint deploys are routed through the remote node's standard `/api/stacks` and `/api/stacks/<name>/deploy` endpoints over the proxy, so the remote node enforces its own policy with its own scanner state. The controlling instance does not bypass remote policy.
|
||||
|
||||
**Remote-node call path.** Sencho's controlling instance reaches remote nodes through the same authenticated proxy used by the rest of the fleet management surfaces, carrying the remote node's bearer token. Remote nodes do not need to accept any inbound connection beyond the one they already accept for fleet operations.
|
||||
|
||||
**Audit.** Blueprint create, edit, delete, pin, withdraw, and accept are administrative actions and are recorded through the standard [Audit log](/features/audit-log) pipeline along with the rest of the admin surface.
|
||||
|
||||
## Limitations and non-goals
|
||||
|
||||
By design, Blueprints do not include:
|
||||
|
||||
- A distributed storage layer (no CSI, no Longhorn-style replication).
|
||||
- Automatic volume migration between nodes.
|
||||
- Per-node parameter overrides or templating (one compose, all nodes).
|
||||
- Staged or canary rollouts.
|
||||
- Versioning history with one-click rollback. To revert, paste the prior compose into the editor and save.
|
||||
|
||||
Concrete operational constraints:
|
||||
|
||||
- **Reconciler cadence.** The tick interval is 60 seconds (5-second initial delay after startup). Use **Apply now** to force an immediate tick after a change.
|
||||
- **Compose size.** YAML must be 96 KiB or fewer. Split very large compose files into smaller blueprints, or move generated content out of the compose body.
|
||||
- **Selector size.** A selector accepts up to 200 entries per side (200 `nodes.ids`, or up to 200 each in `labels.any` and `labels.all`).
|
||||
- **Name.** 1 to 64 characters matching `^[a-z0-9][a-z0-9_-]*$`. Names are immutable after creation; to rename, recreate the blueprint and withdraw the old one.
|
||||
- **Snapshot semantics.** **Snapshot, then evict** captures the compose YAML only. Volume bytes are removed by `docker compose down` just as with **Evict and destroy data**.
|
||||
- **Restore from snapshot.** Reserved for the future Volume Migration feature and currently disabled in the state review dialog.
|
||||
|
||||
These omissions keep Blueprints honest: a compose-native fleet primitive that distributes the file you already have to the nodes you choose.
|
||||
|
||||
## Migrating stateful data between nodes (manual)
|
||||
|
||||
Sencho's compose-native lane does not include automatic volume shipping. **Snapshot, then evict** is a compose-only safety net: it preserves the YAML so you can redeploy elsewhere, but it does not move data. To relocate a stateful Blueprint's data from node A to node B, do it by hand before withdrawing:
|
||||
|
||||
1. Stop the Blueprint deployment on node A from the deployment table. *Tip:* use **Snapshot, then evict** so the compose YAML is parked in Fleet → Snapshots while you handle volumes.
|
||||
2. Use your host tooling (`docker run --rm -v <volume>:/data busybox tar -czf - /data > snapshot.tar.gz`, or app-aware tooling such as `pg_basebackup` / `mysqldump` / `mongodump`) to capture the volume on node A.
|
||||
1. Stop the Blueprint deployment on node A from the deployment table. Use **Snapshot, then evict** so the compose YAML is parked in [Fleet · Snapshots](/features/fleet-backups) while you handle volumes.
|
||||
2. Use your host tooling (`docker run --rm -v <volume>:/data busybox tar -czf - /data > snapshot.tar.gz`, or app-aware tooling such as `pg_basebackup`, `mysqldump`, or `mongodump`) to capture the volume on node A.
|
||||
3. Transfer the artifact to node B and restore it into the named volume there.
|
||||
4. Update the Blueprint's selector to include node B; click **Apply now**.
|
||||
|
||||
A future Volume Migration feature will automate this with app-aware backup tooling.
|
||||
|
||||
## Practical workflows
|
||||
|
||||
**Multi-node identical reverse proxy.** A stateless `caddy-edge` blueprint with a labels selector matching `any=[edge]` and drift mode **Enforce**. Adding a new edge node deploys the proxy automatically within one tick. Editing the compose to bump the Caddy image redeploys every edge node on the next tick without operator intervention. Drift caused by a manual `docker compose down` on one edge node is corrected silently within a minute.
|
||||
|
||||
**Single-node managed Postgres.** A stateful `pg-fleet` blueprint with a `nodes` selector pointing at one database node and drift mode **Suggest**. The first deploy enters **Awaiting confirmation** so the operator chooses **Deploy fresh**. Subsequent compose changes (image bump, config change) re-enter **Awaiting confirmation** on the same node so the operator can decide whether the new revision is safe for the existing volume. Drift on the running container fires a `blueprint_drift_detected` notification but never auto-redeploys.
|
||||
|
||||
**Pin a blueprint to a specific node despite the selector.** Admiral users open **Fleet · Federation**, find the blueprint in the pin policy table, and pick the target node from the **Pinned to** dropdown. The pin overrides the selector for that blueprint, deploys only to the pinned node, and also overrides cordon on that node. Useful for relocating a stateful service to a specific host without rewriting the selector. Clear the pin to restore selector-driven placement.
|
||||
|
||||
**Observe-only audit blueprint.** A stateless monitoring stack (Vector, Promtail, a Prometheus exporter) with drift mode **Observe**. Drift is recorded silently in the deployment table; no notification fires and no auto-fix runs. Useful when you want Sencho to track placement and detect divergence on a low-signal stack without paging anyone.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### "Name conflict" on a deployment row
|
||||
<AccordionGroup>
|
||||
<Accordion title="A deployment row shows 'Name conflict' and refuses to deploy">
|
||||
A directory by the blueprint's name already exists on that node and does not carry the `.blueprint.json` marker. The most likely cause is a manually created stack with the same name. Resolution: rename either the existing stack or the blueprint, then click **Apply now** on the detail sheet to retry.
|
||||
</Accordion>
|
||||
<Accordion title="A stateful deployment is stuck in 'Awaiting confirmation'">
|
||||
The reconciler will not auto-deploy a stateful blueprint to a node it has never run on. Click **Confirm deploy** on the row, then choose **Deploy fresh** in the dialog. Sencho will create empty named volumes and start the stack.
|
||||
</Accordion>
|
||||
<Accordion title="A deploy was blocked by vulnerability policy">
|
||||
Open **Settings · Security** and review the active scan policies for the target node. The Blueprint deployment row records the blocking policy and affected image count. Either fix the image, relax the policy, or deploy through an explicit admin-approved bypass on the stack surface. Blueprints do not silently bypass deploy enforcement.
|
||||
</Accordion>
|
||||
<Accordion title="Compose content was rejected on save">
|
||||
Sencho accepts valid YAML up to 96 KiB. Split very large compose files into smaller blueprints or move generated content out of the Blueprint. Do not paste secrets into the compose body; use environment files or [Fleet Secrets](/features/fleet-secrets) where appropriate.
|
||||
</Accordion>
|
||||
<Accordion title="A remote node is offline or disconnected during apply">
|
||||
The row moves to **Failed** with the remote error. Reconnect the node (see [Pilot Agent](/features/pilot-agent) and [Multi-node management](/features/multi-node)), verify whether the stack directory contains `docker-compose.yml` and `.blueprint.json`, then click **Apply now**. If the directory exists without a matching marker, Sencho treats it as a name conflict until you rename or remove the remote stack manually.
|
||||
</Accordion>
|
||||
<Accordion title="A Docker daemon or registry failure surfaced during apply">
|
||||
The deployment row moves to **Failed** and records the Docker or registry error. Resolve the daemon, socket, registry credentials, rate limit, disk, or volume-permission issue on the affected node, then click **Apply now**. Watch that node's stack activity and security scan status after retry.
|
||||
</Accordion>
|
||||
<Accordion title="Drift is reported but never gets corrected">
|
||||
Confirm the drift policy is **Enforce** and that **Reconciler enabled** is on. Open the detail sheet to check the row's status and the most recent drift summary. If the drift was caused by a compose change that would destroy named volumes, Enforce intentionally downgrades to Suggest semantics for that event. Either change the compose to one that preserves the volumes, or withdraw the deployment with explicit operator confirmation and let the new revision deploy fresh.
|
||||
</Accordion>
|
||||
<Accordion title="'Apply now' is greyed out">
|
||||
**Apply now** requires the blueprint to be enabled. Open the detail sheet, click **Enable** to flip the reconciler back on, then click **Apply now**. The button is also disabled while the detail sheet is in edit mode; save or cancel the edit first.
|
||||
</Accordion>
|
||||
<Accordion title="Disabling a blueprint is rejected">
|
||||
Blueprints with active or drifted deployments refuse to disable, because doing so would orphan the running containers without a follow-up plan. Withdraw each affected deployment first, then disable the blueprint.
|
||||
</Accordion>
|
||||
<Accordion title="The data is gone after 'Snapshot, then evict'">
|
||||
Both eviction modes run `docker compose down`, which removes the named volumes managed by the stack on the target node. The snapshot stored in **Fleet · Snapshots** captures the compose definition only; the volume bytes are not included. To preserve data, capture the volume by hand before withdrawing (see *Migrating stateful data between nodes*). Bind mounts on the host filesystem are left in place by both eviction modes.
|
||||
</Accordion>
|
||||
<Accordion title="Eviction was aborted with 'Failed to capture compose snapshot before eviction'">
|
||||
Sencho aborted the eviction because the pre-eviction compose snapshot could not be written to the database. The deployment is still in place; nothing was destroyed. Check that the database is reachable, then retry the eviction. If you accept data loss and want to evict regardless, use **Evict and destroy data** instead.
|
||||
</Accordion>
|
||||
<Accordion title="A pinned blueprint did not deploy to nodes in its selector">
|
||||
By design. A pin replaces the selector entirely with the single pinned node, even if the selector resolves to other nodes. Either clear the pin in the Federation tab to restore selector-driven placement, or move the pin to a different node.
|
||||
</Accordion>
|
||||
<Accordion title="The reconciler took a minute to react to a change I just made">
|
||||
The reconciler tick interval is one minute. To force an immediate evaluation after editing the selector, the compose, or the drift policy, click **Apply now** on the detail sheet. **Apply now** is also the recovery action for a row in **Failed** status.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
A directory by the blueprint's name already exists on that node and does not carry our `.blueprint.json` marker. Most likely cause: a manually created stack with the same name. Resolution: rename either the existing stack or the blueprint, then click **Apply now**.
|
||||
## Common questions
|
||||
|
||||
### Stateful blueprint stuck in "Awaiting confirmation"
|
||||
<AccordionGroup>
|
||||
<Accordion title="Do Blueprints replace my existing stacks?">
|
||||
No. Stacks created in the per-stack lane (the **Stacks** sidebar) and Blueprints (the **Deployments** tab) coexist on every node. A Blueprint is identified by its `.blueprint.json` marker; the reconciler refuses to touch any directory that does not carry one with a matching blueprint ID, so user-authored stacks are safe from blueprint actions.
|
||||
</Accordion>
|
||||
<Accordion title="Is a Blueprint the same as a Stack?">
|
||||
A Blueprint is a fleet-wide declaration: one compose YAML targeting a set of nodes. Each node where the blueprint resolves materializes as a Stack on that node, in the same compose directory layout the rest of the per-stack lane uses, with a `.blueprint.json` marker added. You can browse the materialized stack in the per-node Stacks view; the Blueprint is the source of truth that drives it.
|
||||
</Accordion>
|
||||
<Accordion title="How fast does drift get noticed?">
|
||||
Within one minute. The reconciler ticks every 60 seconds with a 5-second initial delay after startup; **Apply now** on the detail sheet forces an immediate tick. Within Enforce mode, drift correction begins on the same tick that detects the drift.
|
||||
</Accordion>
|
||||
<Accordion title="Can I roll back to a previous revision?">
|
||||
Not through a one-click history. Blueprints intentionally do not keep a versioned revision history. To revert, paste the prior compose into the editor and save; the reconciler treats the change as a new revision and redeploys.
|
||||
</Accordion>
|
||||
<Accordion title="Does the snapshot in 'Snapshot, then evict' contain my database?">
|
||||
No. The snapshot captures the compose YAML only, so you can redeploy the same definition elsewhere. Volume bytes (the database files, the cached state, the uploads directory) are removed by `docker compose down` along with the named volumes. To preserve data across an eviction, capture the volume by hand before withdrawing.
|
||||
</Accordion>
|
||||
<Accordion title="Why doesn't Sencho move volumes between nodes for me?">
|
||||
Volume migration is app-aware. A Postgres datadir is not the same kind of artifact as a Redis RDB file or a MinIO bucket, and a generic `tar` of the volume is rarely safe while the container is running. A future Volume Migration feature will integrate app-aware backup tooling. Until then, the manual workflow in *Migrating stateful data between nodes* is the supported path.
|
||||
</Accordion>
|
||||
<Accordion title="If I cordon a node, does that uninstall the Blueprints on it?">
|
||||
No. Cordoning only blocks new placements and skips state-review provisioning. Existing deployments on the cordoned node keep running and continue drift-checking. To uninstall a deployment from a cordoned node, withdraw it explicitly from the deployment table.
|
||||
</Accordion>
|
||||
<Accordion title="Can a Pilot-attached node host a Blueprint deployment?">
|
||||
Yes. The reconciler dispatches deploys to remote nodes through the standard proxy regardless of whether the remote runs in Pilot Agent or Distributed API Proxy mode. The remote node enforces its own vulnerability policy and writes its own marker file, the same way a local deploy does.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
Click **Confirm deploy** on the row, then choose **Deploy fresh**. Sencho will create empty named volumes and start the stack. If the row appeared after a compose edit, review the revision first; confirming writes the new compose file and deploys it on that node.
|
||||
## Where Blueprints fits
|
||||
|
||||
### Deploy blocked by vulnerability policy
|
||||
|
||||
Open **Settings → Security** and review the active scan policies for the target node. The Blueprint deployment row records the blocking policy and affected image count. Either fix the image, relax the policy, or deploy through an explicit admin-approved bypass on the stack surface. Blueprints do not silently bypass deploy enforcement.
|
||||
|
||||
### Compose content rejected
|
||||
|
||||
Sencho accepts valid YAML up to 96 KiB. Split very large compose files into smaller stacks or move generated content out of the Blueprint. Do not paste secrets into the compose body; use environment files or fleet secrets where appropriate.
|
||||
|
||||
### Remote node disconnected during apply
|
||||
|
||||
The row moves to failed with the remote error. Reconnect the node, verify whether the stack directory contains `docker-compose.yml` and `.blueprint.json`, then click **Apply now**. If the directory exists without a matching marker, Sencho treats it as a name conflict until you rename or remove the remote stack manually.
|
||||
|
||||
### Docker daemon or registry failure during apply
|
||||
|
||||
The deployment row moves to failed and records the Docker or registry error. Resolve the daemon, socket, registry credentials, rate limit, disk, or volume-permission issue on the affected node, then click **Apply now**. Watch that node's stack activity and security scan status after retry.
|
||||
|
||||
### Drift never gets corrected
|
||||
|
||||
Confirm the drift policy is `enforce` and the blueprint is enabled. Open the detail sheet to see the deployment row's status and the most recent drift summary. If the drift was caused by a compose change that would destroy named volumes, Enforce intentionally downgrades. Change the compose to one that preserves volumes, or withdraw and re-deploy with explicit operator confirmation.
|
||||
|
||||
### Cannot disable a blueprint
|
||||
|
||||
Blueprints with active or drifted deployments refuse to disable; you would orphan them silently. Withdraw the deployments first, then disable.
|
||||
|
||||
### Where is my data after "Snapshot, then evict"?
|
||||
|
||||
The named volumes managed by the stack on the target node are removed when the eviction runs `docker compose down`. The snapshot in Fleet → Snapshots holds the compose definition only; volume bytes are not included. To preserve data, capture the volume by hand before withdrawing (see *Migrating stateful data between nodes*). Bind mounts on the host filesystem are left in place by both eviction modes.
|
||||
|
||||
### "Failed to capture compose snapshot before eviction"
|
||||
|
||||
Sencho aborted the eviction because the pre-eviction compose snapshot could not be written. The deployment is still in place. Check the database is reachable (the snapshot lives in `fleet_snapshots`), then retry. If you accept data loss and want to evict regardless, use **Evict and destroy data** instead.
|
||||
|
||||
## What's not in scope
|
||||
|
||||
By design, Blueprints do not include:
|
||||
|
||||
- A distributed storage layer (no CSI, no Longhorn-style replication)
|
||||
- Automatic volume migration between nodes
|
||||
- Per-node parameter overrides or templating (one compose, all nodes)
|
||||
- Staged or canary rollouts
|
||||
- Versioning history with one-click rollback (re-paste the prior compose to revert)
|
||||
|
||||
These omissions keep Blueprints honest: a compose-native fleet primitive that distributes the file you already have to the nodes you choose.
|
||||
|
||||
## Rollback and watch plan
|
||||
|
||||
To roll back a Blueprint hardening change, revert the application build, restart the Sencho backend and frontend, and re-paste the prior compose revision into any affected Blueprint. No database migration rollback is required for the validation and reconciliation guards described here.
|
||||
|
||||
For the first 24 to 48 hours after deployment, watch:
|
||||
|
||||
- Backend logs tagged `[BlueprintReconciler]`, `[BlueprintService]`, and `[BlueprintService:diag]` when Developer Mode is on.
|
||||
- Deployment-row counts by node: `failed`, `pending_state_review`, `name_conflict`, and `drifted`.
|
||||
- Vulnerability policy blocks and post-deploy scan failures on each node.
|
||||
- Docker daemon errors, registry timeouts, image pull failures, volume permission errors, and out-of-disk errors on the affected node.
|
||||
|
||||
Rollback if failed Blueprint deploys exceed 5 percent of apply attempts for 30 minutes, if a single node repeatedly fails all Blueprint deploys after Docker recovers, or if stateful deployments leave `pending_state_review` without an operator action path.
|
||||
|
||||
In a fleet, roll back the controlling instance first so it stops issuing new Blueprint actions. Then roll back remote nodes. During version skew, older remote nodes can still receive stack deploy requests, but they may lack matching validation or diagnostic logs. Inspect per-node logs rather than relying only on aggregate fleet counts.
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Fleet Federation" icon="network-wired" href="/features/fleet-federation">
|
||||
Cordon nodes to block new placements and pin blueprints to a specific node to override the selector.
|
||||
</Card>
|
||||
<Card title="Fleet Snapshots" icon="camera" href="/features/fleet-backups">
|
||||
Where Snapshot-then-evict parks the compose YAML and where you redeploy it from.
|
||||
</Card>
|
||||
<Card title="Fleet Actions" icon="bolt" href="/features/fleet-actions">
|
||||
Bulk operations on the running fleet. Blueprints sit on the desired-state lane; Fleet Actions sit on the imperative lane.
|
||||
</Card>
|
||||
<Card title="Atomic Deployments" icon="shield-check" href="/features/atomic-deployments">
|
||||
Deploy safety on the per-stack lane. The same vulnerability policy gate runs before each Blueprint deploy.
|
||||
</Card>
|
||||
<Card title="Pilot Agent" icon="link" href="/features/pilot-agent">
|
||||
Outbound-only remote mode. A Pilot-attached node is eligible as a Blueprint target the same way a Distributed API Proxy node is.
|
||||
</Card>
|
||||
<Card title="Alerts & Notifications" icon="bell" href="/features/alerts-notifications">
|
||||
Where `blueprint_drift_detected` and `blueprint_drift_correction_failed` route. Configure channels per notification.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
|
||||
@@ -1,109 +1,113 @@
|
||||
---
|
||||
title: "CVE Suppressions"
|
||||
description: "Accept known-benign vulnerabilities fleet-wide so your scan results stay focused on findings that actually need action."
|
||||
description: "Accept known-benign CVEs once and have them dimmed across the fleet at read time, without ever modifying stored scan results."
|
||||
---
|
||||
|
||||
Not every CVE that Trivy reports requires a response. Some are false positives on your base image, some have been accepted by your security review, and some are waiting on an upstream patch. CVE suppressions let you annotate these findings once so they stop competing for attention in every scan, comparison, and alert.
|
||||
Not every CVE that Trivy reports needs a response. Some are false positives on your base image, some have been accepted after security review, and some are waiting on an upstream patch. CVE suppressions let you annotate those findings once so they stop competing for attention in every scan, comparison, and alert.
|
||||
|
||||
<Note>
|
||||
CVE suppressions are available on every tier. Suppressions written on a control node replicate to its replicas at fleet scope.
|
||||
CVE suppressions are available on every tier. The **admin** role is required to add, edit, or remove them. Suppressions created on the control Sencho instance replicate to every remote at fleet scope.
|
||||
</Note>
|
||||
|
||||
## What suppressions do
|
||||
|
||||
A suppression is a rule that says "this CVE is acknowledged." When a scan's findings are read back for display or comparison, Sencho checks each finding against the active suppression list:
|
||||
A suppression is a rule that says "this CVE is acknowledged." When a scan's findings are read back for display, export, or comparison, Sencho checks each finding against the active suppression list:
|
||||
|
||||
- Suppressed findings remain in the database and in the scan totals. Nothing is deleted.
|
||||
- In the scan drawer and the comparison sheet, suppressed rows are dimmed and marked with a shield-off icon.
|
||||
- The reason you recorded is visible in the row so reviewers understand why it was accepted.
|
||||
- Suppressed findings remain in the database and in the stored severity counts. Nothing is deleted.
|
||||
- In the scan drawer and the comparison sheet, suppressed rows render dimmed and carry a small shield-off icon next to the CVE ID.
|
||||
- Hovering the package column on a suppressed row surfaces the recorded reason so reviewers can see why it was accepted.
|
||||
|
||||
Counts on badges and summary ribbons continue to reflect the raw findings. Suppressions are a visual filter, not an accounting trick. If you suppress a CVE and then remove the suppression, the finding resurfaces on the next read, without rescanning.
|
||||
Badge counts in the Resources Hub continue to reflect the raw findings. Suppressions are a visual filter, not an accounting change. Remove a suppression and the underlying finding resurfaces on the next read, without rescanning.
|
||||
|
||||
## Creating a suppression
|
||||
## Creating a suppression from Settings
|
||||
|
||||
Go to **Settings → Security** and scroll to **CVE Suppressions**, then click **Add Suppression**.
|
||||
Open **Settings → Security** and scroll to **CVE Suppressions**, then click **Add Suppression**.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/cve-suppressions/settings-panel.png" alt="Settings Security section showing the CVE Suppressions panel with a list of accepted CVEs" />
|
||||
<img src="/images/cve-suppressions/settings-panel.png" alt="Settings Security page with the CVE Suppressions panel listing two accepted CVEs, each showing the CVE ID, an outlined package badge, a clamped reason line, the author and expiry metadata, and a trash icon for removal" />
|
||||
</Frame>
|
||||
|
||||
The dialog has the following fields:
|
||||
|
||||
| Field | Description |
|
||||
|-------|-------------|
|
||||
| **CVE ID** | The identifier of the finding. Accepts both `CVE-YYYY-NNNN` and GitHub advisory IDs like `GHSA-xxxx-xxxx-xxxx`. |
|
||||
| **Package** | Optional. Leave empty to suppress every occurrence of this CVE regardless of package, or set it to a specific package name (e.g. `openssl`) to narrow the scope. |
|
||||
| **Image pattern** | Optional glob against image references (e.g. `registry.example.com/app*`). Leave empty to apply fleet-wide. |
|
||||
| **Reason** | Required. A short note explaining why this CVE is accepted. Surfaced on every suppressed row. Do not paste credentials, tokens, or vendor secrets here — reasons replicate fleet-wide and surface on every node's suppressions panel. |
|
||||
| **Expires in** | Optional. Number of days after which the suppression stops applying. Useful for "patched in the next release" acknowledgements. Leave empty for an indefinite suppression. |
|
||||
| **CVE or advisory ID** | Required. Accepts both `CVE-YYYY-NNNN` and `GHSA-xxxx-xxxx-xxxx`. |
|
||||
| **Package (optional)** | Leave blank to suppress every occurrence of this CVE across every package, or pin a specific package name (e.g. `openssl`) to narrow the scope. |
|
||||
| **Image pattern (optional)** | Glob applied to image references (`*` matches any sequence, case-sensitive). For example, `lscr.io/linuxserver/*` matches every LinuxServer image, and `*alpine*` matches anything containing `alpine`. Leave blank to apply fleet-wide. |
|
||||
| **Reason** | Required. A short note explaining why the CVE is accepted. Surfaced on every suppressed row and on the hover title in scan results. Do not paste credentials, tokens, or vendor secrets, since reasons replicate fleet-wide. |
|
||||
| **Expires in (days, optional)** | Number of days after which the suppression stops applying. Useful for "patched in the next release" entries. Leave blank for an indefinite suppression. |
|
||||
|
||||
<Frame>
|
||||
<img src="/images/cve-suppressions/create-dialog.png" alt="Dialog for adding a new CVE suppression with fields for CVE ID, package, image pattern, reason, and expiry" />
|
||||
<img src="/images/cve-suppressions/create-dialog.png" alt="New suppression dialog with kicker SUPPRESSIONS · NEW, title New suppression, and all five fields populated with example values for a Git CVE scoped to LinuxServer images" />
|
||||
</Frame>
|
||||
|
||||
### Suppressing directly from a scan result
|
||||
|
||||
The panel's empty state hints at the faster path: from any vulnerability scan, click the small shield icon at the right edge of a finding's row. The dialog opens pre-filled with the CVE ID and the package name from that row (both read-only in this flow), leaving you to add a Reason, an optional Image pattern, and an optional Expiry. This is the recommended workflow for everyday triage, because it keeps the scope as narrow as the originating finding. To broaden the scope (for example, to suppress across every package), create the rule from **Settings → Security** instead.
|
||||
|
||||
### How specificity is resolved
|
||||
|
||||
When multiple suppressions match the same finding, the most specific one wins:
|
||||
When more than one suppression matches the same finding, the most specific one wins. Specificity is scored as:
|
||||
|
||||
1. A suppression that pins both a package name and an image pattern is the most specific.
|
||||
2. A suppression with only a package name beats a wildcard pattern.
|
||||
3. A suppression with only an image pattern beats a fully-wildcard rule.
|
||||
1. Both **Package** and **Image pattern** set: score 3 (most specific).
|
||||
2. **Package** only: score 2.
|
||||
3. **Image pattern** only: score 1.
|
||||
4. Neither (the broadest fleet-wide rule): score 0.
|
||||
|
||||
The reason field of the winning suppression is the one displayed on the row.
|
||||
The winning rule's Reason field is the one displayed on the row.
|
||||
|
||||
## Viewing suppressed findings
|
||||
|
||||
Open any scan drawer and scroll to the vulnerability table. Suppressed rows look like this:
|
||||
Open any scan drawer and look at the Vulnerabilities table. Suppressed rows are dimmed and carry a shield-off icon next to the CVE ID:
|
||||
|
||||
<Frame>
|
||||
<img src="/images/cve-suppressions/suppressed-row.png" alt="Vulnerability scan drawer with a suppressed row dimmed and labeled with a shield-off icon" />
|
||||
<img src="/images/cve-suppressions/suppressed-row.png" alt="Vulnerabilities tab of a Sencho image scan showing two dimmed suppressed rows for Docker CVEs at the top with the shield-off icon, and one un-suppressed GHSA row below at full brightness with an inline Suppress button on the right" />
|
||||
</Frame>
|
||||
|
||||
Suppressed rows also carry through to the **Compare scans** view. Both the added and removed columns show the suppression state, so a finding that you've already accepted will not look like a new regression when comparing an older baseline.
|
||||
The same dim-and-icon treatment carries through to the **Compare** sheet, so a CVE you've already accepted does not look like a new regression when comparing against an older baseline. Hovering the package column of a suppressed row reveals the recorded Reason without expanding the row.
|
||||
|
||||
## Fleet-wide replication
|
||||
|
||||
Suppressions are managed on the **control** Sencho instance and replicate automatically to every remote Sencho you've registered. There is nothing extra to configure:
|
||||
Suppressions are managed on the **control** Sencho instance and replicate automatically to every remote you've registered:
|
||||
|
||||
- Creating or editing a suppression on the control pushes the full list to every remote.
|
||||
- Remote instances show the suppression list in a read-only state. The **Add Suppression** and **Delete** buttons are hidden, and a banner explains that rules are managed upstream.
|
||||
- Incoming scan results on the control and every remote apply the same suppression set.
|
||||
- Creating, editing, or removing a suppression on the control pushes the full list to every remote.
|
||||
- A Sencho instance that has received at least one push from a control is a **replica**. On a replica, **Settings → Security** shows the suppression list read-only, and replicated rows carry a small `replicated` badge so they are easy to tell apart from any locally-created entries.
|
||||
- When you're signed into a control and have a **remote node selected** from the node switcher, the CVE Suppressions panel itself is hidden and a "Scanner is per-node" banner explains that scanning runs on the remote while rules live on the control.
|
||||
|
||||
If a push to a remote fails (for example because the remote is temporarily offline), Sencho records the failure and retries on the next fleet sync tick. See [Fleet Sync](/features/fleet-sync) for the details of how replication works and how to inspect push status.
|
||||
The full replication, retry, and reanchor flow (including the API call to re-bind a replica to a new control) is documented in [Fleet Sync](/features/fleet-sync).
|
||||
|
||||
## Removing a suppression
|
||||
|
||||
Click the trash icon on any row in the suppressions panel. A confirmation dialog calls out that removing the rule will cause matching findings to reappear in scan results.
|
||||
Click the trash icon on any row in the panel. A confirmation dialog ("Remove suppression", kicker `SUPPRESSIONS · REMOVE · IRREVERSIBLE`) warns that future scan results will surface the CVE again wherever it applies.
|
||||
|
||||
To change a suppression's scope (for example, to narrow an image pattern or extend an expiry), delete the existing rule and create a new one with the updated fields.
|
||||
To change a suppression's scope (for example, to narrow an image pattern or extend the expiry), remove the existing rule and create a fresh one with the updated fields.
|
||||
|
||||
## Export awareness
|
||||
|
||||
Suppressed findings carry through to the [SARIF export](/features/vulnerability-scanning#sarif-export) with a SARIF `suppressions` entry of `kind: external` and `status: accepted`. Code-scanning dashboards that respect SARIF suppressions dismiss those findings with the Reason you recorded.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
<Accordion title="I suppressed a CVE but the count on the badge is unchanged">
|
||||
Badge counts reflect the raw findings so they remain meaningful for alerting and policy evaluation. The filter is applied in the scan drawer, the comparison sheet, and every read surface, but the stored totals do not change. Open the scan drawer to confirm the row is dimmed with a shield-off icon.
|
||||
</Accordion>
|
||||
<AccordionGroup>
|
||||
<Accordion title="I suppressed a CVE but the count on the badge is unchanged">
|
||||
Badge counts reflect the raw findings so they remain meaningful for alerting and policy evaluation. Suppressions apply when results are read for display (the scan drawer, the Compare sheet, the SARIF export), not when they are stored. Open the scan drawer to confirm the row is dimmed with a shield-off icon.
|
||||
</Accordion>
|
||||
<Accordion title="A suppression I added on the control is not visible on a replica">
|
||||
Replication runs on every write. If the push failed (network blip, replica restart), the control retries every 5 minutes for 24 hours and the replica picks up the latest state on the next successful push. See [Fleet Sync](/features/fleet-sync) for how to investigate persistent push failures.
|
||||
</Accordion>
|
||||
<Accordion title="I see suppressions on a replica but cannot edit them">
|
||||
Replicas are read-only for security rules. Sign in to the control instance to add, edit, or delete suppressions; changes sync automatically. Replicated rows show a `replicated` badge and no trash icon.
|
||||
</Accordion>
|
||||
<Accordion title="A suppression does not match a finding I expect it to">
|
||||
Two common causes:
|
||||
|
||||
<Accordion title="A suppression I added on the control is not visible on a remote">
|
||||
Replication runs on every write. If the push failed (network blip, remote restart), check the fleet sync status on the control under **Fleet → Sync status**. The remote picks up the latest state on the next successful push.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="I see suppressions on a remote but cannot edit them">
|
||||
Remote Sencho instances are read-only for security rules. Sign in to the control instance to add, edit, or delete suppressions. Changes sync automatically.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="A suppression does not match a finding I expect it to">
|
||||
Two common causes:
|
||||
|
||||
- **Image pattern mismatch.** The pattern uses glob syntax where `*` matches any sequence. `nginx*` matches `nginx:1.25` but not `docker.io/library/nginx:1.25`; use `*nginx*` for a broader match.
|
||||
- **Expired rule.** If **Expires in** was set, the suppression stops applying after the deadline. The row shows an "expired" badge; edit it or create a fresh rule.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="A remote refuses pushes from my control with a 'control identity mismatch' error">
|
||||
Each replica anchors to the first control fingerprint it sees. If the same replica is later pushed from a different control (for example, after rebuilding the control instance from a snapshot or migrating to a new server), the replica rejects the new control until it is reanchored. On the replica, open **Fleet → Sync status** and use **Reanchor** to clear the cached fingerprint so the next push from the new control establishes a fresh anchor. See [Fleet Sync](/features/fleet-sync) for the full reanchor flow.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="I demoted the control and the old replica still shows the mirrored suppressions">
|
||||
Demoting a control to a standalone instance only affects that one node. Replicas that were following it keep the last set of replicated rows until either a new control pushes to them or you reanchor the replica. From the replica's **Fleet → Sync status**, use **Reanchor** to drop the mirrored rules; the suppressions panel returns to a clean local-only state.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="My suppression list is enormous and the fleet sync warns about truncation">
|
||||
The fleet sync wire protocol caps a single push at 10,000 rows so a misconfigured control cannot wedge a replica with an unbounded payload. If your local list exceeds that cap, the warning appears in the control logs and only the first 10,000 rows are pushed. Trim expired or unused entries from the suppressions panel, or split scopes across distinct rules so each remains meaningful, before relying on fleet replication for the full set.
|
||||
</Accordion>
|
||||
- **Image pattern mismatch.** The pattern uses glob syntax (case-sensitive) where `*` matches any sequence. `nginx*` matches `nginx:1.25` but not `docker.io/library/nginx:1.25`; use `*nginx*` for a broader match.
|
||||
- **Expired rule.** If **Expires in** was set, the suppression stops applying after the deadline. The panel shows an `expired` badge on that row; remove it and create a fresh rule.
|
||||
</Accordion>
|
||||
<Accordion title="A replica refuses pushes from my control with 'control identity mismatch'">
|
||||
Each replica anchors to the first control fingerprint it receives. If a different control later tries to push (for example, after rebuilding the control from a snapshot or migrating it to a new server), the replica rejects the push until it is reanchored. The reanchor procedure (admin API call on the replica) is documented in [Fleet Sync](/features/fleet-sync#control-anchor); the next push from any control after a reanchor becomes the new anchor.
|
||||
</Accordion>
|
||||
<Accordion title="My suppression list is large and the fleet sync warns about truncation">
|
||||
The fleet sync wire protocol caps a single push at 5,000 rows so a misconfigured control cannot wedge a replica with an unbounded payload. If your local list exceeds that cap, the warning appears in the control logs and only the first 5,000 rows are pushed. Trim expired or unused entries from the suppressions panel, or split scopes across distinct rules, so each remains meaningful before relying on fleet replication for the full set.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
@@ -1,119 +1,231 @@
|
||||
---
|
||||
title: Dashboard
|
||||
description: Real-time system stats, stack health, configuration overview, and recent activity for your active node.
|
||||
description: Real-time system health, stack load, configuration overview, fleet activity, and recent alerts for the active node.
|
||||
---
|
||||
|
||||
The **Home** tab is the first thing you see after logging in. It provides a live overview of your node's health, resource usage, stack status, active configuration, and recent activity.
|
||||
The **Home** tab is the first thing you see after logging in. It surfaces the active node's overall health, live system metrics, the load on every stack, the on/off state of every automation and security feature, a fleet- or restart-activity panel, and the running alert tape, all on a single scrollable page.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/dashboard/dashboard-overview.png" alt="Sencho dashboard showing status masthead, unified gauge strip, and stack health table" />
|
||||
<img src="/images/dashboard/dashboard-overview.png" alt="Sencho Home dashboard showing the Critical state masthead, four-tile resource gauge strip, paginated stack health table, Configuration Status card paired with Fleet Heartbeat, and the Recent Alerts feed." />
|
||||
</Frame>
|
||||
|
||||
## Status masthead
|
||||
|
||||
The masthead at the top of the dashboard is the single place to read the node's current condition. It carries:
|
||||
The masthead at the top of the dashboard is the single place to read the node's current condition.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/dashboard/status-masthead.png" alt="Status masthead in the Critical state with the editorial state word, the meta line 'LOCAL · 4 NODES · LAST SYNC 1S', the reasons line 'RAM 94% · 18 unread errors', the RUNNING / CPU / MEM stat tiles, and a 92-alerts counter." />
|
||||
</Frame>
|
||||
|
||||
It carries:
|
||||
|
||||
- A **state word** (Healthy, Degraded, or Critical) set in the editorial display face so the reader sees it first.
|
||||
- A **pulsing dot** that mirrors the state color: green when nominal, amber when degraded, rose when critical.
|
||||
- A **meta line** with the active node, the number of nodes in the fleet, and the time since the last successful sync.
|
||||
- A **reasons line** that names exactly which signals moved the state away from Healthy (for example, "CPU 84% · 2 exited · 3 unread errors"). No hovering required.
|
||||
- Three quick stat tiles on the right: containers running, aggregate CPU, and memory in use. Hover the **Running** tile to see the managed / external / exited breakdown.
|
||||
- An alerts counter pinned to the far right.
|
||||
- A **pulsing dot** that mirrors the state color: green when nominal, amber when degraded, rose when critical. The dot is solid (no pulse) when Healthy.
|
||||
- A **meta line** in uppercase mono tracking with the active node's name, the number of nodes registered to this Sencho instance, and the time since the last successful poll, for example `LOCAL · 4 NODES · LAST SYNC 1S`.
|
||||
- A **reasons line** that names exactly which signals moved the state away from Healthy, for example `RAM 95% · 18 unread errors`. When the state is Healthy the reasons line reads `All systems nominal`.
|
||||
- Three quick stat tiles on the right edge of the bar: **RUNNING** (`active/total`), **CPU**, and **MEM**. Hover the **RUNNING** tile to expand a managed / external (and exited, when present) breakdown in a cursor-following tooltip.
|
||||
- An alerts counter pinned to the far right, showing the number of unread notifications next to a bell icon. The icon and count tint amber while at least one alert is unread.
|
||||
|
||||
Sencho derives the state from CPU, RAM, disk usage, exited containers, and unread error alerts:
|
||||
<Frame>
|
||||
<img src="/images/dashboard/running-tile-hover.png" alt="Status masthead with the RUNNING tile hovered, the cursor-following tooltip showing '15 managed · 1 external' under the masthead values." />
|
||||
</Frame>
|
||||
|
||||
| State | Meaning |
|
||||
|-------|---------|
|
||||
| **Healthy** | All systems nominal. No resources above warning thresholds, no unread errors. |
|
||||
| **Degraded** | At least one resource is above 80%, there are exited containers, or there are unread error alerts. |
|
||||
| **Critical** | At least one resource is above 90%, or there are exited containers combined with unread errors. |
|
||||
The masthead's CPU stat tile tints amber once host CPU crosses 80% and stays amber even at 90% or higher. The MEM and RUNNING values stay neutral; the reasons line and the gauge strip below are where you read severity.
|
||||
|
||||
## Unified gauge strip
|
||||
## Resource gauges
|
||||
|
||||
A single rail of four tiles shows the numbers that change minute-to-minute:
|
||||
A single rail of four tiles shows the numbers that change minute-to-minute.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/dashboard/resource-gauges.png" alt="Four-tile resource gauge strip: a CPU hero tile at 0.3% with the 'avg 0% last 10m · peak 0% @ 03:05 PM' caption and a 10-minute sparkline; a MEMORY tile at 94% with the gauge bar in destructive red and '14.7 GB / 15.6 GB' below; a DISK tile at 57% in brand cyan with '93.6 GB / 173.6 GB'; and a NETWORK tile reading '8.8 KB/s' with the rx/tx split and a rhythm sparkline." />
|
||||
</Frame>
|
||||
|
||||
| Tile | What it shows |
|
||||
|------|---------------|
|
||||
| **CPU (hero)** | Current usage with a 10-minute sparkline, average for the window, and the peak value with its time offset |
|
||||
| **Memory** | RAM usage with a compact bar and the exact used/total split |
|
||||
| **Disk** | Mount usage with a compact bar and the exact used/total split |
|
||||
| **Network** | Total throughput per second with the received/transmitted split and a live rhythm spark |
|
||||
| **CPU** (hero) | Current usage with the host's core count in the kicker, a 10-minute sparkline, and a caption with `avg X% last 10m · peak Y% @ HH:MM` |
|
||||
| **MEMORY** | RAM usage as a percentage with a gauge bar and the exact `used / total` split below |
|
||||
| **DISK** | Mount usage as a percentage with a gauge bar and the exact `used / total` split below |
|
||||
| **NETWORK** | Total throughput per second with a `↓ <rx>/s · ↑ <tx>/s` breakdown and a live rhythm sparkline |
|
||||
|
||||
The CPU tile's sparkline uses cyan as the data color and marks the peak in amber. Memory and disk bars turn amber at 80% and rose at 90%.
|
||||
The gauge bars (and the corresponding numeric values) pick up amber at 80% and rose at 90%. The CPU sparkline uses brand cyan as the data color and does not mark the peak, since the caption already names the peak value and time. The network strip falls back to a dashed baseline when there is no traffic to plot.
|
||||
|
||||
While the dashboard is loading the CPU tile reads `--` and the caption shows `collecting metrics…`; bars and sparklines render once the first sample arrives.
|
||||
|
||||
## Stack health
|
||||
|
||||
A mono table of every stack discovered in your `COMPOSE_DIR`, sorted by load so the stacks demanding attention sit at the top:
|
||||
A mono table of every stack discovered in the active node's `COMPOSE_DIR`, sorted so the stacks demanding attention sit at the top.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/dashboard/stack-health.png" alt="Stack health table titled 'Stack health · 15 STACKS · SORTED BY LOAD' with pagination chevrons reading 1 / 2, listing eight stacks (profilarr, cloudflared, tautulli, swag, plex, bazarr, radarr, prowlarr) with a green status dot, host 'Local', uptime, current CPU and MEM, and a 10-minute CPU sparkline per row." />
|
||||
</Frame>
|
||||
|
||||
| Column | Description |
|
||||
|--------|-------------|
|
||||
| **State dot** | Green when healthy, amber when any container in the stack is pushing CPU above 80%, rose when one has exited or is above 90% |
|
||||
| **Stack** | Stack name (derived from the directory name) |
|
||||
| **Host** | Active node this stack belongs to |
|
||||
| **Up** | How long the oldest running container has been up, in compact units (s/m/h/d) |
|
||||
| **CPU** | Latest aggregate CPU for the stack's containers |
|
||||
| **Memory** | Total memory allocated by the stack's containers |
|
||||
| **CPU · 10m** | Per-stack sparkline of the last 10 minutes, tinted to match the row state |
|
||||
| **Status dot** | Green when the stack is running and its 10-minute peak CPU is under 80%, amber when peak CPU is at or above 80%, rose when any container has exited or peak CPU is at or above 90% |
|
||||
| **STACK** | Stack name, derived from the compose file (extension stripped) |
|
||||
| **HOST** | Active node this stack belongs to |
|
||||
| **UP** | How long the oldest running container has been up, in compact units (`s` / `m` / `h` / `d`); a stopped or never-started stack reads `--` |
|
||||
| **CPU** | Latest aggregate CPU across the stack's containers |
|
||||
| **MEM** | Latest aggregate memory across the stack's containers, formatted in MB or GB |
|
||||
| **CPU · 10m** | Per-stack 10-minute sparkline tinted to match the row state; warn and error rows mark the peak with a contrasting accent color |
|
||||
|
||||
Warning rows take on a subtle amber wash; critical rows take on a rose wash. Click any row to jump to that stack's editor. If you have more than 8 stacks, the list paginates automatically.
|
||||
Rows sort first by state (errors → warnings → healthy) and then by 10-minute peak CPU descending, so an exited stack always rises to the top and the noisiest healthy stacks float above the quiet ones. Warning rows pick up a subtle amber wash; critical rows pick up a rose wash. Click any row, or focus it and press <kbd>Enter</kbd> or <kbd>Space</kbd>, to jump to that stack's editor.
|
||||
|
||||
The table paginates at eight rows; chevrons appear in the header along with an `N / M` indicator when there is more than one page. When the active node has no stacks at all the card renders an empty state with a layered-disks glyph and the message `No stacks found. Create one from the sidebar.`
|
||||
|
||||
## Configuration Status
|
||||
|
||||
The Configuration Status card gives you an at-a-glance view of every toggleable automation and security feature on the active node, so nothing is silently off when you expect it to be on.
|
||||
The Configuration Status card is the at-a-glance audit of every toggleable automation and security feature on the active node, so nothing is silently off when you expect it to be on.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/dashboard/configuration-status.png" alt="Dashboard showing the Configuration Status card on the left and the Recent Activity feed on the right" />
|
||||
<img src="/images/dashboard/configuration-status.png" alt="Configuration Status card with four sections: Notifications (Notification agents, Alert rules, Notification routing all reading None), Automation (Auto-heal None, Auto-update '1 / 1', Webhooks None, Scheduled tasks '2 actives'), Security (MFA Off, SSO Off, Vulnerability scanning None), and Backups & Thresholds (Cloud Backup 'Sencho Cloud', Alert thresholds 'CPU 100% · RAM 100% · Disk 100%', Crash detection On)." />
|
||||
</Frame>
|
||||
|
||||
The card is divided into four sections:
|
||||
The card is divided into four sections.
|
||||
|
||||
### Notifications & Alerts
|
||||
### Notifications
|
||||
|
||||
| Row | What it shows |
|
||||
|-----|---------------|
|
||||
| **Notification agents** | Which delivery agents (Discord, Slack, custom webhook) are configured and enabled |
|
||||
| **Alert rules** | Number of per-stack alert rules in effect |
|
||||
| **Notification routing** | Number of enabled routing rules that direct specific categories to specific agents |
|
||||
| **Notification agents** | The list of enabled delivery agents from `Discord`, `Slack`, `Webhook`, joined by commas; reads `None` when no agent is enabled |
|
||||
| **Alert rules** | Total per-stack alert rules in effect, formatted `<n> rules` |
|
||||
| **Notification routing** (Admiral) | Number of enabled routing rules that direct categories to specific agents, formatted `<n> routes` |
|
||||
|
||||
### Automation
|
||||
|
||||
The Automation block only renders on Skipper or Admiral.
|
||||
|
||||
| Row | What it shows |
|
||||
|-----|---------------|
|
||||
| **Auto-heal policies** | Enabled and total crash-recovery policies across all stacks |
|
||||
| **Auto-update stacks** | Stacks enrolled in automated image-update checks |
|
||||
| **Webhooks** | Outbound webhook triggers that fire on stack events |
|
||||
| **Scheduled tasks** | Active scheduled operations (backups, restarts, scripts) |
|
||||
| **Auto-heal policies** | `<enabled> / <total> active` for crash-recovery policies across all stacks; reads `None` when no policies exist |
|
||||
| **Auto-update stacks** | `<enabled> / <total>` count of stacks enrolled in automated image-update checks; reads `None` when none are enrolled |
|
||||
| **Webhooks** (Skipper) | Active inbound deploy webhooks tied to Git Sources or stacks, formatted `<n> actives` |
|
||||
| **Scheduled tasks** (Admiral) | Active scheduled operations (backups, restarts, scripts), formatted `<n> actives` |
|
||||
|
||||
### Security
|
||||
|
||||
| Row | What it shows |
|
||||
|-----|---------------|
|
||||
| **MFA** | Whether multi-factor authentication is set up for your account |
|
||||
| **SSO** | Whether single sign-on is enabled and the active provider |
|
||||
| **Vulnerability scanning** | Number of enabled scan policies |
|
||||
| **MFA** | `On` when TOTP is configured for the signed-in operator, `Off` when configured but disabled, `Not set up` when there is no MFA secret on file |
|
||||
| **SSO** | The active SSO provider name (`OIDC`, `Google`, `GitHub`, `Okta`, `LDAP`); reads `Off` when SSO is not enabled |
|
||||
| **Vulnerability scanning** (Skipper) | Count of enabled scan policies on the active node; reads `None` when no policy is enabled |
|
||||
|
||||
### Backups & Thresholds
|
||||
|
||||
| Row | What it shows |
|
||||
|-----|---------------|
|
||||
| **Cloud Backup** | Active cloud backup provider (or Disabled) |
|
||||
| **Alert thresholds** | Current CPU, RAM, and disk alert thresholds |
|
||||
| **Crash detection** | Whether global crash detection is active |
|
||||
| **Cloud Backup** (Admiral) | The active cloud backup target: `Sencho Cloud`, `Custom S3` (with ` (auto)` appended when auto-upload is enabled), or `Disabled` |
|
||||
| **Alert thresholds** | The current host thresholds, formatted `CPU x% · RAM y% · Disk z%` |
|
||||
| **Crash detection** | `On` when global container-crash notifications are enabled, `Off` otherwise |
|
||||
|
||||
**Click any row** to jump directly to the settings section that manages it.
|
||||
Click any row to jump directly to the settings section that manages it. Rows that require a higher tier than the active license are not rendered at all; you do not see a locked placeholder. The section headers (Notifications, Automation, Security, Backups & Thresholds) only render when at least one of their rows is visible, so a Community node sees a tighter card with no Automation block and no Cloud Backup row.
|
||||
|
||||
Rows for features locked to a higher tier are shown in a muted state with an upgrade indicator. The data refreshes automatically every 60 seconds and updates immediately when any setting changes.
|
||||
The data refreshes automatically every 60 seconds and immediately on any container start/stop/restart event broadcast over the live notification stream, so the card stays in lockstep with what the rest of the UI shows.
|
||||
|
||||
## Recent Activity
|
||||
## Activity panel
|
||||
|
||||
The Recent Activity feed, to the right of Configuration Status, shows the ten most recent events on the active node: deployments, image updates, auto-heal actions, vulnerability scan results, cloud backups, and system notices. Each entry shows a category icon, the event message, and a relative timestamp.
|
||||
To the right of Configuration Status, the dashboard shows one of two activity panels depending on whether you have remote nodes registered:
|
||||
|
||||
If no activity has been recorded yet, the feed shows an empty state.
|
||||
- **Fleet Heartbeat** when at least one remote node is configured.
|
||||
- **Stack Restarts (7d)** when this Sencho instance manages only the local node.
|
||||
|
||||
## Recent alerts
|
||||
### Fleet Heartbeat
|
||||
|
||||
The bottom section displays triggered alert notifications sorted by recency. Each entry shows the severity level (info, warning, or error), the alert message, and a relative timestamp.
|
||||
<Frame>
|
||||
<img src="/images/dashboard/fleet-heartbeat.png" alt="Fleet Heartbeat card with the radio-tower kicker '4 NODES' on the right, listing four entries with status dots: Local with the brand 'LOCAL' pill and 16 containers; Opsix with 6 containers and 91 ms latency; Pitt-Moba with 5 containers and 52 ms latency; sencho-pilot-test with no container count and 'n/a' latency." />
|
||||
</Frame>
|
||||
|
||||
Use the **Clear All Notifications** button to dismiss all alerts. If you have more than 8 alerts, the list paginates.
|
||||
The header carries a radio-tower icon, the total node count, and a `<n> unreachable` callout in destructive red whenever any node is offline. Each row shows:
|
||||
|
||||
Alert thresholds are configured per stack in **Settings > Notifications**. See [Alerts & Notifications](/features/alerts-notifications) for setup details.
|
||||
- A **status dot**: green when the node has answered a recent heartbeat, amber when its state is unknown, rose when offline or unreachable.
|
||||
- The **node name** in mono.
|
||||
- A **LOCAL** brand pill on the local node.
|
||||
- The **active container count** for that node, returned alongside the latency in the fleet-overview payload.
|
||||
- A **latency** column on the right of online proxy nodes, in milliseconds. Pilot-agent nodes (which receive their commands over an outbound tunnel rather than being polled) read `n/a` since round-trip latency is not meaningful for them.
|
||||
- A **last-seen** timestamp on offline rows; pilot-agent nodes use their tunnel heartbeat, regular proxy nodes use the local Sencho instance's last successful contact.
|
||||
|
||||
The card refreshes every 30 seconds. When you have not registered any node it shows a `No nodes registered.` empty state.
|
||||
|
||||
### Stack Restarts (7d)
|
||||
|
||||
When the dashboard renders this variant, the same column shows a 7-day rolling restart map. Each stack with at least one restart in the window appears as a row carrying:
|
||||
|
||||
- The **stack name** in mono.
|
||||
- A **category badge** for the dominant restart reason: `crash` (rose), `auto-heal` (green), or `manual` (brand). Ties resolve in favor of `crash` over `auto-heal` over `manual`.
|
||||
- The **total restart count** on the right, with a proportional fill bar tinted to match the badge.
|
||||
|
||||
Stacks with zero restarts are not listed individually; they are summarized in a single footer row that reads `<n> stacks stable`. When no stack has restarted in the window the card shows a green check and `No restarts in the last 7 days.` This panel refreshes every five minutes.
|
||||
|
||||
## Recent Alerts
|
||||
|
||||
The bottom-most card on the dashboard collects every triggered notification on the active node, sorted by recency.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/dashboard/recent-alerts.png" alt="Recent Alerts card with pagination chevrons reading 1 / 12, listing eight rows tagged 'sencho-pilot-test' with severity icons (warning triangle for the threshold breach, info circles for deploys, destructive octagons for the two crash-detected entries), each carrying a relative timestamp, and a 'Clear All Notifications' button anchored at the bottom right." />
|
||||
</Frame>
|
||||
|
||||
Each row shows:
|
||||
|
||||
- A **severity icon**: an info circle (brand) for `info`, a warning triangle (amber) for `warning`, an octagonal alert (rose) for `error`.
|
||||
- An optional **node badge** when the alert came from a remote node, showing the source node's name. Local-node alerts skip the badge.
|
||||
- The **alert message**, dimmed once the alert has been read.
|
||||
- A relative **timestamp** on the right, for example `2h ago`.
|
||||
|
||||
A **Clear All Notifications** button is anchored at the bottom right of the card. Clicking it deletes every alert in the feed by issuing a `DELETE /api/notifications` against each node that contributed an entry; while the request is in flight the button label switches to `Clearing...` and the button disables. The button stays visible while there are alerts; when none remain the card collapses to a green check and the message `No recent alerts.`
|
||||
|
||||
The list paginates at eight rows; chevrons and an `N / M` indicator appear in the header when there is more than one page.
|
||||
|
||||
Alert rules and severity routing are configured per stack in **Settings · Notifications**. See [Alerts and notifications](/features/alerts-notifications) for channel setup, per-stack rules, and the routing engine.
|
||||
|
||||
## How health is derived
|
||||
|
||||
The masthead's state word and reasons line both come from a single rule against five signals: host CPU, host RAM, host disk, exited container count, and unread `error`-level notifications.
|
||||
|
||||
| State | Meaning |
|
||||
|-------|---------|
|
||||
| **Healthy** | All five signals nominal. CPU, RAM, and disk are all under 80%, no container is in the `exited` state, and there are no unread error alerts. |
|
||||
| **Degraded** | At least one of CPU, RAM, or disk is at or above 80%, OR at least one container is in the `exited` state, OR at least one unread error alert exists. |
|
||||
| **Critical** | At least one of CPU, RAM, or disk is at or above 90%, OR at least one container is in the `exited` state AND at least one unread error alert is present at the same time. |
|
||||
|
||||
Every signal that contributes to a non-Healthy state appears in the reasons line, separated by middle dots (`·`), so you can read the cause directly without hovering anything.
|
||||
|
||||
## Refresh cadence
|
||||
|
||||
The dashboard is built from several independent polling loops so that fast-moving values stay live without flooding the API for slowly-changing ones.
|
||||
|
||||
| Surface | Polling cadence |
|
||||
|---------|-----------------|
|
||||
| Container counts and host system stats (masthead, gauges) | 5 seconds |
|
||||
| Stack statuses (status dots and uptime in the table) | 10 seconds |
|
||||
| Historical metrics (CPU sparklines and per-stack 10m series) | 60 seconds |
|
||||
| Configuration Status card | 60 seconds |
|
||||
| Fleet Heartbeat | 30 seconds |
|
||||
| Stack Restarts (7d) | 5 minutes |
|
||||
| Recent Alerts | Pushed live over the notifications WebSocket, with a 60-second safety-net reconcile poll |
|
||||
|
||||
In addition, every Docker container event (start, stop, die, restart, health-status change) on the active node fires a `sencho:state-invalidate` browser event that immediately refetches container counts, system stats, stack statuses, and the Configuration Status card so the masthead and the table react in well under a second instead of waiting for the next polling tick.
|
||||
|
||||
When you switch the active node from the node switcher, the dashboard resets every panel to its loading state and starts a new round of fetches against the new node so you never see stale data from the previous one.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Stack health is empty even though I have stacks on disk">
|
||||
The table reads from the active node's `COMPOSE_DIR`. If you mounted the compose volume to a different path inside the container, set the environment variable in your `docker-compose.yml` so it matches what is on the host. Per-node overrides live on the **Compose directory** field of the node entry; the resolution chain is `node.compose_dir` → `process.env.COMPOSE_DIR` → `/app/compose`. After fixing the mount or the field, the table populates on the next 10-second poll.
|
||||
</Accordion>
|
||||
<Accordion title="Fleet Heartbeat shows a node as 'unreachable' or with a red dot">
|
||||
The Heartbeat card calls `/api/fleet/overview` on the local Sencho instance every 30 seconds; that endpoint records the last contact for each remote node. A red dot means the local instance has not been able to reach that node's `api_url`, or that the long-lived API token attached to the node entry has expired or been revoked. Re-test the connection from **Settings · Nodes**, regenerate the token on the remote instance, paste the new value into the node form, and the heartbeat clears on the next tick. See [Multi-node management](/features/multi-node) for the full setup flow and [Pilot agent](/features/pilot-agent) when the offline node is in agent mode.
|
||||
</Accordion>
|
||||
<Accordion title="A pilot-agent node shows 'n/a' for latency">
|
||||
Pilot-agent nodes connect outbound to the primary over a reverse tunnel, so there is no synchronous request-response loop to measure. The Heartbeat card uses the tunnel's last heartbeat to set the dot color, and intentionally renders `n/a` in the latency column. If the dot is red, check the agent container's logs for the first `[Pilot]` line and confirm the agent can reach the primary URL it is dialing. See [Pilot agent](/features/pilot-agent).
|
||||
</Accordion>
|
||||
<Accordion title="Configuration Status still shows the old value after I changed a setting">
|
||||
The card refreshes every 60 seconds on its own schedule. Most settings changes also dispatch a live invalidation event that triggers an immediate refetch, but some flows (cloud backup provider switch, SSO provider change) only update on the next polling tick. If the row remains stale after a minute, hard-reload the dashboard tab to force a fresh fetch.
|
||||
</Accordion>
|
||||
<Accordion title="Recent Alerts shows '1 / 12' but I want to scrub through history">
|
||||
The card paginates at eight rows. Use the chevrons in the header to flip pages. The full alert history (with filters and search) lives in the bell menu in the top bar; the dashboard card is a tape view of the latest entries on the active node. **Clear All Notifications** deletes the whole feed in one call, fanned out across every node that contributed an entry.
|
||||
</Accordion>
|
||||
<Accordion title="The masthead reads 'Critical' but my CPU is low">
|
||||
Critical can fire on RAM, disk, or the combination of any exited container with an unread error alert, not only on CPU. Read the reasons line right under the meta line; it lists every signal pushing the state up. Common cases: RAM at or above 90% from an oversized service, disk at or above 90% from runaway log retention, or a crashed container that no one has acknowledged in the alerts panel.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
@@ -1,95 +1,130 @@
|
||||
---
|
||||
title: "Deploy Enforcement"
|
||||
description: "Block deploys that violate a scan policy before docker compose up runs, with an admin bypass path and full audit trail."
|
||||
description: "Block deploys that violate a scan policy before docker compose up runs, with an admin bypass path and a full audit trail."
|
||||
---
|
||||
|
||||
Deploy enforcement is the pre-flight half of Sencho's vulnerability workflow. When a [scan policy](/features/vulnerability-scanning#scan-policies) with **Block on deploy** enabled matches a stack, Sencho scans every image referenced by the stack's compose file before starting any container. If any image meets or exceeds the policy's severity threshold, the deploy is rejected and the compose stack is never brought up. Detection always continues post-deploy (and on schedule), so images that develop new vulnerabilities after the initial deploy still surface through alerts.
|
||||
Deploy enforcement is the pre-flight half of Sencho's vulnerability workflow. When a [scan policy](/features/vulnerability-scanning#scan-policies) with **Block on deploy** enabled matches a stack, Sencho scans every image referenced by the stack's compose file before starting any container. If any image meets or exceeds the policy's severity threshold, the deploy is rejected and the stack never starts. Detection always continues post-deploy and on a schedule, so images that develop new vulnerabilities after the initial deploy still surface through alerts.
|
||||
|
||||
<Note>
|
||||
Deploy enforcement requires a **Skipper** or **Admiral** license. Policies on Community are evaluation-only and cannot block deploys.
|
||||
</Note>
|
||||
|
||||
## Configuring a block policy
|
||||
|
||||
Policies are managed under **Settings → Advanced → Security → Scan Policies**. The **Add policy** button opens the editor; existing policies appear as a list of cards with `max: <SEVERITY>` and `block` badges, the configured stack-pattern scope, and pencil and trash buttons.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/deploy-enforcement/policy-list.png" alt="Scan Policies card showing a configured policy with the name 'Production block on critical', a max: CRITICAL badge, a block badge, and the Scope demo-blocked-* set as the stack pattern" />
|
||||
</Frame>
|
||||
|
||||
The editor exposes the five fields that govern enforcement:
|
||||
|
||||
<Frame>
|
||||
<img src="/images/deploy-enforcement/policy-edit-modal.png" alt="New policy modal with kicker SECURITY · NEW POLICY, fields for Name, Stack pattern, Max severity (Critical), Block on deploy toggle ON, and Enabled toggle ON" />
|
||||
</Frame>
|
||||
|
||||
| Field | Purpose |
|
||||
|-------|---------|
|
||||
| **Name** | A descriptive label that appears on the block dialog and in audit log entries. |
|
||||
| **Stack pattern (optional)** | Glob-style match against stack names. `prod-*` matches `prod-api` but not `production-api`. Leave blank to apply to every stack on the node. |
|
||||
| **Max severity** | The threshold. If any image's highest finding meets or exceeds this severity, the policy fires. |
|
||||
| **Block on deploy** | When on, the pre-flight gate hard-rejects deploys that violate the threshold. When off, the policy still evaluates post-deploy and scheduled scans and dispatches warning alerts on violations. |
|
||||
| **Enabled** | Disabled policies are skipped during evaluation. |
|
||||
|
||||
The editor sets the pattern, severity, and toggles. Per-node scoping is set via the [Security API](/api-reference/security#scan-policies) (`node_id`) or replicated from a control node via [Fleet Federation](/features/fleet-federation). The most specific enabled policy that matches the stack on the target node is the one that runs.
|
||||
|
||||
## How enforcement runs
|
||||
|
||||
Sencho applies the pre-flight gate on every code path that starts a compose stack:
|
||||
Sencho applies the pre-flight gate on every code path that can start a compose stack:
|
||||
|
||||
- **Deploy** and **Redeploy** from the stack page
|
||||
- **Update** (re-pull + redeploy)
|
||||
- **Deploy from Git source** and **Apply** on a managed git source
|
||||
- **Template deploy** from the App Store
|
||||
- **Recreate** from the stack actions menu
|
||||
- **Deploy** from the stack page.
|
||||
- **Update** (re-pull plus restart).
|
||||
- **Deploy from a Git Source** when the initial deploy is requested at create time.
|
||||
- **Template deploy** from the App Store.
|
||||
- **Bulk deploy** from the [Stack Labels](/features/stack-labels) page.
|
||||
- **Auto-update scheduler** (the gate is hard-enforced here; see [Auto-update scheduler interaction](#auto-update-scheduler-interaction)).
|
||||
|
||||
On every one of these actions, Sencho:
|
||||
|
||||
1. Looks up the most specific enabled policy that matches the stack on the target node.
|
||||
2. If the policy has **Block on deploy** off, lets the deploy proceed and evaluates the post-deploy scan against the policy for alerting.
|
||||
3. If **Block on deploy** is on, enumerates the stack's images with `docker compose config --images`, runs a pre-flight Trivy scan against each one, and compares the highest severity in each scan against the policy threshold.
|
||||
4. If every image is below the threshold, the deploy proceeds exactly as before. A post-deploy drift scan still runs in the background.
|
||||
5. If any image violates the threshold, the deploy is rejected with HTTP `409 Conflict` and the stack never starts. The UI opens a dialog listing the offending images.
|
||||
4. If every image is below the threshold, the deploy proceeds. A post-deploy drift scan still runs in the background.
|
||||
5. If the compose file fails to parse, enforcement fails closed: the deploy is rejected with a single synthetic violation labeled `(compose parse error)` so a malformed file cannot slip past the gate.
|
||||
6. If any image violates the threshold, the deploy is rejected with HTTP `409 Conflict` and the stack never starts. The UI opens a dialog listing the offending images.
|
||||
|
||||
Pre-flight scans use the same 24-hour digest cache as on-demand scans, so the second deploy of the same image does not pay the full scan time.
|
||||
|
||||
## What the block dialog shows
|
||||
|
||||
<Frame>
|
||||
<img src="/images/vulnerability-scanning/deploy-blocked-dialog.png" alt="AlertDialog shown when a deploy is blocked, listing the policy name and the offending images with severity chips" />
|
||||
<img src="/images/deploy-enforcement/block-dialog.png" alt="Deploy blocked dialog with kicker DEMO-BLOCKED-APP · SCAN POLICY · BLOCKED, title 'Deploy blocked by security policy', a row listing the offending image redis:7.0-alpine with 8 critical · 59 high counts and a CRITICAL severity chip, a Close button, and a destructive Deploy anyway button" />
|
||||
</Frame>
|
||||
|
||||
The dialog shows:
|
||||
|
||||
- The policy that fired, with its max severity.
|
||||
- Every image that violated the threshold, with its severity chip and the counts of critical and high findings.
|
||||
- A kicker with the stack name, the words **SCAN POLICY**, and **BLOCKED** in the destructive accent color.
|
||||
- The policy name and a sentence naming the threshold the image crossed.
|
||||
- One row per offending image, with the image reference in monospace, the counts of critical and high findings, and a severity chip in the matching color.
|
||||
- A **Close** button that dismisses the dialog without deploying.
|
||||
- A **Deploy anyway** button when the current user is an admin (see bypass below); the button is disabled and relabeled **Admin required to bypass** for non-admin roles.
|
||||
- A destructive **Deploy anyway** button when the current user is an admin (see bypass below). For non-admin sessions the primary slot is replaced by a disabled outline button labeled **Admin required to bypass**.
|
||||
|
||||
Clicking a violation row's severity chip takes you to the scan details sheet for that image, where you can review every finding and decide whether to upgrade the base image or [suppress](/features/cve-suppressions) a specific CVE.
|
||||
The dialog is informational; the rows are not interactive. To dig into individual findings, open the stack's image in the [Resources Hub](/features/resources) and review the scan from there or jump straight from the Vulnerability Scanner [scan results drawer](/features/vulnerability-scanning#the-scan-results-drawer).
|
||||
|
||||
## Bypassing a block
|
||||
|
||||
Blocks can be overridden by admins on a per-deploy basis. The button is only visible when:
|
||||
Blocks can be overridden by admins on a per-deploy basis. The **Deploy anyway** button is only active when:
|
||||
|
||||
- The current user has the `admin` role.
|
||||
- The deploy came from the UI (the bypass flag is ignored when the caller role is not admin server-side, so non-admin sessions cannot forge the header).
|
||||
- The deploy came from the UI. The bypass flag is ignored when the caller role is not admin server-side, so non-admin sessions cannot forge the header.
|
||||
|
||||
Every bypass is recorded in the [Audit Log](/features/audit-log) with:
|
||||
|
||||
- `method` and `path` of the originating request (so you can tell whether the bypass came from a deploy, update, template, or git-apply call).
|
||||
- `method` and `path` of the originating request, so you can tell whether the bypass came from a deploy, update, template, or git-from call.
|
||||
- `username` of the admin who bypassed.
|
||||
- A `policy.bypass` summary with the stack name, policy name, violation count, and the list of image references that were bypassed.
|
||||
- A `policy.bypass` summary in the format `policy.bypass stack="<name>" policy="<policyName>" violations=<n> images=[<comma-joined-refs>]`.
|
||||
|
||||
API callers can pass `?ignorePolicy=true` on any deploy endpoint to request a bypass. The flag is only honored when the token's session resolves to an admin user; API tokens without admin scope cannot bypass a policy. See the [Security API reference](/api-reference/security#bypassing-a-block) for the exact request shape.
|
||||
API callers can pass `?ignorePolicy=true` on any deploy endpoint to request a bypass. The flag is only honored when the token's session resolves to a user with the `admin` role; tokens whose user is not an admin cannot bypass a policy. See the [Security API reference](/api-reference/security#bypassing-a-block) for the exact request shape.
|
||||
|
||||
### Auto-update scheduler interaction
|
||||
|
||||
The auto-update scheduler runs deploys without an interactive user, so the bypass flag does not apply. When a scheduled auto-update is rejected by a policy, Sencho dispatches a `scan_finding` warning alert with the stack name, the policy name, and the offending images, then skips that stack and continues the rest of the schedule. The audit-log entry records the actor as `auto-update:<username>` for runs invoked by a logged-in user, or `auto-update:scheduler` for unattended runs. Re-enabling that stack in auto-update means either bringing the image down to compliant severity or relaxing the policy.
|
||||
|
||||
## Drift detection keeps running
|
||||
|
||||
Deploy enforcement prevents a new deploy from introducing known vulnerabilities at the front door. Long-running stacks whose images were clean at deploy time can still develop new CVEs as upstream feeds update. Two surfaces catch this:
|
||||
|
||||
- **Post-deploy scans** run automatically on every successful deploy (whether the pre-flight gate was tripped or not). When a post-deploy scan violates an enabled policy, Sencho dispatches a warning alert and the scan details sheet shows a policy-violation banner.
|
||||
- **Post-deploy scans** run automatically on every successful deploy, whether the pre-flight gate was tripped or not. When a post-deploy scan violates an enabled policy, Sencho dispatches a warning alert and the scan details sheet shows a policy-violation banner.
|
||||
- **[Scheduled fleet scans](/features/vulnerability-scanning#scheduled-fleet-scans)** re-scan every image on a node at a cron schedule you pick. Any scan that violates a matching policy produces the same warning alert and banner, so you never need to babysit running stacks.
|
||||
|
||||
Neither drift mechanism blocks, stops, or quarantines a running stack automatically. The gate is intentionally limited to deploy time.
|
||||
|
||||
## FAQ
|
||||
## Troubleshooting
|
||||
|
||||
**CI pipelines feel slower after I enabled a block policy.**
|
||||
<AccordionGroup>
|
||||
<Accordion title="CI pipelines feel slower after I enabled a block policy">
|
||||
Only the first deploy of a new image pays the full Trivy runtime. Sencho caches scan results by image digest for 24 hours, so repeat deploys of the same image hit the cache in under a second. For greenfield CI where every deploy ships a new image tag, budget 30 to 120 seconds per image on the first deploy depending on image size and whether Trivy's vulnerability database is already seeded on the node.
|
||||
</Accordion>
|
||||
<Accordion title="The gate let a deploy through even though I have a block policy">
|
||||
Check the following in order:
|
||||
|
||||
Only the first deploy of a new image pays the full Trivy runtime. Sencho caches scan results by image digest for 24 hours, so repeat deploys of the same image hit the cache in under a second. For greenfield CI where every deploy ships a new image tag, budget 30 to 120 seconds per image on the first deploy depending on image size and whether Trivy's vulnerability database is already seeded on the node.
|
||||
|
||||
**The gate let a deploy through even though I have a block policy.**
|
||||
|
||||
Check the following in order:
|
||||
|
||||
1. Open **Settings → Security** on the target node and confirm Trivy is installed. Sencho fails open when Trivy is missing, dispatching a warning alert instead of blocking. [Install Trivy](/operations/trivy-setup) to enforce the policy.
|
||||
2. Confirm the policy is enabled and the stack pattern matches the stack name. `prod-*` matches `prod-api` but not `production-api`. An empty pattern matches every stack on the node.
|
||||
3. Check the highest severity in the latest scan for each image. If no image reached the threshold, the gate correctly allowed the deploy.
|
||||
|
||||
**A deploy is blocked and I cannot bypass as a non-admin.**
|
||||
|
||||
Only users with the `admin` role can bypass a block. Ask an admin to review the violations and either bypass the single deploy, upgrade the base image, or suppress the offending CVE with an expiry.
|
||||
|
||||
**The block dialog mentions an image I do not recognize.**
|
||||
|
||||
Pre-flight enumerates images via `docker compose config --images`, which expands any `extends`, `env_file`, or variable substitution in the compose file. An image pulled by a dependency you did not author may appear in the list. Review the stack's compose file and confirm the reference.
|
||||
|
||||
**I want a block policy that only applies to a subset of my fleet.**
|
||||
|
||||
Combine two mechanisms: scope the policy to a specific node (policies scoped to a node win over global ones) and tighten the stack pattern glob. For example, a policy with `stack_pattern=prod-*` scoped to your production node fires only on `prod-*` stacks deployed to that node.
|
||||
1. Open **Settings → Advanced → Security** on the target node and confirm Trivy is installed. Sencho fails open when Trivy is missing, dispatching a warning alert instead of blocking. [Install Trivy](/operations/trivy-setup) to enforce the policy.
|
||||
2. Confirm the policy is enabled and the stack pattern matches the stack name. `prod-*` matches `prod-api` but not `production-api`. An empty pattern matches every stack on the node.
|
||||
3. Check the highest severity in the latest scan for each image. If no image reached the threshold, the gate correctly allowed the deploy.
|
||||
</Accordion>
|
||||
<Accordion title="A deploy is blocked and I cannot bypass as a non-admin">
|
||||
Only users with the `admin` role can bypass a block. Ask an admin to review the violations and either bypass the single deploy, upgrade the base image, or [suppress](/features/cve-suppressions) the offending CVE with an expiry.
|
||||
</Accordion>
|
||||
<Accordion title="The block dialog mentions an image I do not recognize">
|
||||
Pre-flight enumerates images via `docker compose config --images`, which expands any `extends`, `env_file`, or variable substitution in the compose file. An image pulled by a dependency you did not author may appear in the list. Review the stack's compose file and confirm the reference.
|
||||
</Accordion>
|
||||
<Accordion title="The block dialog says (compose parse error)">
|
||||
Enforcement fails closed when the compose file cannot be parsed. The synthetic `(compose parse error)` violation prevents a malformed file from slipping past the gate. Open the stack's [file explorer](/features/stack-file-explorer) and fix the YAML; the deploy succeeds once the file parses cleanly and every image clears the threshold.
|
||||
</Accordion>
|
||||
<Accordion title="The auto-update scheduler skipped a stack I expected to roll forward">
|
||||
A scheduled auto-update that pulls an image violating a matching policy is skipped, not blocked with a 409. The skip is announced through a `scan_finding` warning alert and recorded in the [Audit Log](/features/audit-log) with the actor `auto-update:<username>` (the user who launched the run) or `auto-update:scheduler` for unattended runs. To unblock the schedule, either upgrade the image to a compliant version or relax the policy threshold for that stack.
|
||||
</Accordion>
|
||||
<Accordion title="I want a block policy that only applies to a subset of my fleet">
|
||||
Combine two mechanisms: scope the policy to a specific node (policies scoped to a node win over global ones) and tighten the stack-pattern glob. For example, a policy with `stack_pattern=prod-*` scoped to your production node fires only on `prod-*` stacks deployed to that node.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
@@ -1,84 +1,213 @@
|
||||
---
|
||||
title: Fleet Actions
|
||||
description: Run fleet-wide bulk operations from one place — stop stacks across nodes by label, or apply labels to many stacks at once.
|
||||
title: "Fleet Actions"
|
||||
description: "Bulk operations across the fleet from one tab: stop stacks by label, replace labels on a batch of stacks, and reclaim Docker disk space on every node."
|
||||
---
|
||||
|
||||
The **Fleet Actions** tab on the Fleet view groups bulk operations that touch more than a single stack on a single node. Each action lives in its own card, runs from the control instance, and reports per-node and per-stack results inline so you never have to click through a modal to learn what happened.
|
||||
|
||||
Three cards ship today: **Stop fleet by label**, **Bulk label assign**, and **Prune Docker resources fleet-wide**.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/fleet-actions/fleet-actions-overview.png" alt="Fleet view with the Fleet Actions tab selected. Three side-by-side cards: Stop fleet by label (rose rail), Bulk label assign (purple rail), and Prune Docker resources fleet-wide (amber rail). Each card carries its own form controls and a calm warning callout beneath the primary button." />
|
||||
</Frame>
|
||||
|
||||
<Note>
|
||||
Fleet Actions require a Sencho **Skipper** or **Admiral** license.
|
||||
Fleet Actions is a Skipper feature. Every card requires an admin user role.
|
||||
</Note>
|
||||
|
||||
The **Fleet Actions** tab inside Fleet centralizes bulk operations that span more than a single stack and would otherwise need many manual clicks. It lives next to **Deployments** and **Routing** in the Fleet view, after the separator that follows Status.
|
||||
## What Fleet Actions covers (and what it doesn't)
|
||||
|
||||
Fleet Actions covers the operations that aren't already exposed elsewhere in Sencho:
|
||||
|
||||
- **Stop fleet by label** dispatches a stop to every stack labeled with a given name across every node in the fleet.
|
||||
- **Bulk label assign** applies the same label set to many stacks on one node in a single round trip.
|
||||
- **Prune Docker resources fleet-wide** reclaims disk space on every node by removing unused images, volumes, and networks in one submit.
|
||||
|
||||
Other bulk actions you may be looking for live in their natural homes:
|
||||
Fleet Actions is the home for operations that span the fleet but don't fit anywhere else. If your goal is in the table below, the dedicated surface is the right place.
|
||||
|
||||
| Goal | Where to do it |
|
||||
|------|----------------|
|
||||
| Restart, stop, or update a subset of stacks | **Bulk mode** in the stack sidebar (press `B` to enable, then check the stacks you want) |
|
||||
| Schedule a recurring restart, update, or snapshot | **Schedules** in the primary navigation |
|
||||
| Trigger a Sencho self-update on remote nodes | **Check Updates** button on the Fleet tab |
|
||||
| Restart or update a hand-picked subset of stacks on one node | **Bulk mode** in the [stack sidebar](/features/sidebar#bulk-mode) (press `B` to enable, then check the stacks you want) |
|
||||
| Schedule a recurring restart, stop, update, or snapshot | [Scheduled Operations](/features/scheduled-operations) |
|
||||
| Trigger a Sencho self-update across remote nodes | **Check Updates** button on the Fleet masthead |
|
||||
| Steer where new blueprint deployments land | [Fleet Federation](/features/fleet-federation) |
|
||||
| Replicate scan policies and CVE suppressions to remotes | [Fleet Sync](/features/fleet-sync) |
|
||||
|
||||
## Three cards, three execution paths
|
||||
|
||||
The cards share a tab and a tier gate, but they don't share an execution path. Knowing which path runs explains the result panels and the failure modes.
|
||||
|
||||
| Card | Endpoint | Where it runs | Scope |
|
||||
|---|---|---|---|
|
||||
| Stop fleet by label | `POST /api/fleet/labels/fleet-stop` | Control instance orchestrates; fans out to each node | Every configured node |
|
||||
| Bulk label assign | `POST /api/fleet-actions/labels/bulk-assign` | Target node (request is proxied) | The single node you pick |
|
||||
| Prune Docker resources fleet-wide | `POST /api/fleet/labels/fleet-prune` | Control instance orchestrates; fans out to each node | Every reachable node |
|
||||
|
||||
The two fan-out cards (Stop and Prune) iterate every node in **Settings → Nodes**, including offline remotes; unreachable nodes show up in the results with a transport error rather than blocking the rest of the batch. The single-node card (Bulk label assign) proxies the request through the standard `x-node-id` header to the node you select, so the work happens locally on that node.
|
||||
|
||||
## Stop fleet by label
|
||||
|
||||
This card sends a stop to every stack on every node that has a label matching the name you type. Labels are matched **by name** across the fleet, so a label called `production` on one node and a separate label also called `production` on another node both match.
|
||||
Stop every stack that carries a given label name on every node where that label exists. Labels are matched **by name** across the fleet, so a label called `production` on one node and an independently-authored `production` label on another node both match. See [Stack Labels](/features/stack-labels) for how to author the selector taxonomy.
|
||||
|
||||
1. Open **Fleet > Fleet Actions**.
|
||||
2. In the **Stop fleet by label** card, type a label name. The input suggests names that already exist on any node.
|
||||
### Step by step
|
||||
|
||||
1. Open **Fleet → Fleet Actions**.
|
||||
2. Type a label name in the **Label name** field. The input autocompletes against label names that already exist on any **online** node; an offline node with an unseen label still receives the request, it just won't show up in the suggestions.
|
||||
3. Click **Stop matching stacks**.
|
||||
4. Confirm the destructive action. Sencho dispatches stops per node.
|
||||
5. Results appear inline below the card, grouped by node, with per-stack success or failure rows.
|
||||
4. A confirmation appears with the kicker **Fleet stop** and the title `Stop all stacks labeled "<name>"?`. Click **Stop fleet** to commit.
|
||||
|
||||
A node that has no label by that name appears in the results with **no matching label** instead of an error.
|
||||
<Frame>
|
||||
<img src="/images/fleet-actions/fleet-actions-stop-confirm.png" alt="Fleet stop confirmation dialog. Kicker 'Fleet stop' in red mono, italic title 'Stop all stacks labeled "docs-preview"?', Cancel and Stop fleet buttons in the footer." />
|
||||
</Frame>
|
||||
|
||||
### Reading the per-node breakdown
|
||||
|
||||
When the request finishes, the results render below the form, grouped by node. Each node row carries a colored icon and either a stack count or a `(no matching label)` annotation; the indented children below each row are the per-stack results.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/fleet-actions/fleet-actions-stop-results.png" alt="Per-node breakdown after running fleet stop against a label that no node has. The header reads PER-NODE BREAKDOWN with two badges, '0 ok' and '7 failed'. Seven rows follow, one per node (Local, Opsix, Pitt-Moba, SLX-Mars, sencho-pilot-test, sencho-test-01, sencho-test-02), each annotated '(no matching label) · Label not present'." />
|
||||
</Frame>
|
||||
|
||||
A few quirks worth knowing:
|
||||
|
||||
- A node that has no label by that name appears as `<node> (no matching label)` and is counted in the **failed** badge. This is not a transport failure, it just means the label was not present on that node.
|
||||
- A node where the label exists but no stacks are assigned to it appears with a matched count of zero stacks. No per-stack rows render.
|
||||
- When the control instance reaches a remote node, the per-stack result you see comes from the remote node's own response. If the remote returns a non-2xx for the whole label, every stack on that node renders with the same error message.
|
||||
|
||||
### Behaviour and partial-failure semantics
|
||||
|
||||
- The endpoint always returns 200 with a `results` array. Partial failures live inside that array; the HTTP status is not the place to look.
|
||||
- Each remote node call carries a 60-second timeout. A slow remote with many stacks can produce a clean per-stack list or a timeout row, depending on whether the remote streamed before the timeout fired.
|
||||
- Local nodes share a per-node bulk-action lock with the per-label action endpoint, so a fleet stop and a per-label stop initiated against the same node serialize cleanly instead of double-stopping the same containers.
|
||||
|
||||
## Bulk label assign
|
||||
|
||||
This card replaces the label set on many stacks at once on a single node.
|
||||
Replace the label set on a batch of stacks on a single node, in one round trip. Use this when you've decided on a taxonomy change (e.g. splitting `prod` into `prod-edge` and `prod-core`) and need to relabel many stacks without clicking through each stack's editor.
|
||||
|
||||
1. Pick the node you want to update from the dropdown.
|
||||
2. Check the stacks you want to update. **Select all** is available once stacks load.
|
||||
3. Click each label pill to toggle it. The chosen pills become the new label set for every selected stack.
|
||||
4. Click **Apply to N stacks** and confirm.
|
||||
<Frame>
|
||||
<img src="/images/fleet-actions/fleet-actions-bulk-assign.png" alt="Bulk label assign card with the Local node selected, three stacks checked (bazarr, plex; counter reads 'Stacks (3/14)'), two labels toggled on (Media and Network; counter reads 'Labels (2/3)'), and the primary button reading 'Apply to 3 stacks'." />
|
||||
</Frame>
|
||||
|
||||
Selecting no labels and applying clears existing label assignments on the chosen stacks. The selected label set always **replaces** the existing one rather than appending to it.
|
||||
### Single-node scope (and why)
|
||||
|
||||
Unlike the other two cards, Bulk label assign targets exactly one node. The request is proxied via `x-node-id` to the node you pick from the **Select a node** dropdown, so the label rewrites happen on that node's local database. There is no fleet-wide variant in v1; if you need to retag the same stacks on multiple nodes, run the card once per node.
|
||||
|
||||
### Step by step
|
||||
|
||||
1. Pick a node from the **Select a node** dropdown. The card loads that node's stacks and labels in parallel.
|
||||
2. Check the stacks you want to update. The counter in the **STACKS** header reads `Stacks (selected/total)`; the **Select all** affordance flips to **Clear** once everything is selected.
|
||||
3. Toggle the label pills you want as the new label set. The counter in the **LABELS** header tracks the selection.
|
||||
4. Click **Apply to N stacks**. A confirmation appears titled `Apply N labels to M stacks?` (or the singular forms). Confirm to commit.
|
||||
|
||||
### Replace, not append
|
||||
|
||||
The selected label set **replaces** each chosen stack's existing label set on this node. Selecting zero labels clears assignments on those stacks; the confirmation modal calls this out so the destructive case is hard to miss.
|
||||
|
||||
### Batch ceiling
|
||||
|
||||
The endpoint accepts up to **1,000 stack assignments per call** and returns `400` over the limit. In practice this is well above any sensible UI selection. Per-entry validation is lenient: an invalid stack name returns a per-entry failure row in the **Per-stack results** card, and the rest of the batch still applies.
|
||||
|
||||
## Prune Docker resources fleet-wide
|
||||
|
||||
This card reclaims disk space on every reachable node by running Docker prune across the targets you select. Each target runs serially on a node; nodes are processed in parallel.
|
||||
Reclaim disk space on every reachable node by deleting unused images, volumes, and networks. The control instance fans out to each node and reports reclaimed bytes per node and per target.
|
||||
|
||||
1. Check the targets you want to prune: **Images**, **Volumes**, and **Networks**. At least one is required.
|
||||
2. Pick a **Scope**:
|
||||
- **Managed only** (default): removes only resources owned by stacks Sencho manages. Safe for shared Docker hosts.
|
||||
- **All unused**: runs `docker prune --all` on every reachable node. Removes every unused image, volume, or network, including resources from workloads Sencho does not manage.
|
||||
3. Click **Prune across fleet** and confirm. The confirmation copy escalates when **All unused** is selected.
|
||||
4. Results appear inline below the card, grouped by node with one sub-row per target showing approximate bytes reclaimed.
|
||||
<Frame>
|
||||
<img src="/images/fleet-actions/fleet-actions-prune.png" alt="Prune Docker resources fleet-wide card. Description reads 'Reclaim space on 7 nodes by removing unused images, volumes, and networks. Reclaimed bytes are approximate.' Below: TARGETS checkboxes (Images checked by default; Volumes; Networks). SCOPE segmented control (Managed only / All unused) with helper text 'Restricts to resources owned by stacks Sencho manages.' Primary button 'Prune across fleet'. A calm callout reads 'Prune is destructive and cannot be undone. Each target is run serially per node; reclaimed bytes appear per node and per target below.'" />
|
||||
</Frame>
|
||||
|
||||
Reclaimed-byte totals are best-effort. Docker does not report bytes for network removals, so the networks row always reports `0 B`.
|
||||
### Pick what to prune
|
||||
|
||||
## Permissions
|
||||
The **TARGETS** checkboxes are independent and at least one must be ticked: **Images**, **Volumes**, **Networks**. The card defaults to Images alone, which is the cheapest and most common case.
|
||||
|
||||
Both cards require the **admin** role and a **Skipper** or **Admiral** license. Community-tier users see a calm explainer card in this tab instead of the action surface.
|
||||
### Managed only versus All unused
|
||||
|
||||
Scope is a segmented control with two options:
|
||||
|
||||
- **Managed only** (default). Sencho looks up the list of stacks it knows about on the node, then prunes only resources owned by those stacks. Active containers and resources that other tools placed on the host are untouched. Helper text reads *Restricts to resources owned by stacks Sencho manages.*
|
||||
- **All unused**. Sencho runs `docker prune --all` for each selected target. Any image, volume, or network not currently in use is deleted, including resources from workloads Sencho does not manage. The confirmation title flips to **Prune ALL unused resources across the fleet?** and the confirm button reads **Prune everything unused**.
|
||||
|
||||
### Behaviour and reclaimed bytes
|
||||
|
||||
- Each remote node receives one `POST /api/system/prune/system` per selected target, with a 120-second timeout. If a transport error fires for one target, the remaining targets on that node are short-circuited with the same error so a dead node does not absorb the full multi-target timeout budget.
|
||||
- Local nodes serialize against a per-node lock (`bulk-prune:<nodeId>`). A second prune launched against the same local node while the first is still in flight returns *A prune is already running on this node* for each target.
|
||||
- Reclaimed bytes are reported by the Docker daemon and are approximate. Per-node rows in the results panel sum the per-target reclaim; the per-target children show how much each prune actually freed.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
| Requirement | Why it matters |
|
||||
|---|---|
|
||||
| **Skipper or Admiral license on the control instance** | Every card gates on the control instance's tier. Remote nodes inherit the paid tier through the proxy's tier header, so a paid control plane covers the whole fleet. |
|
||||
| **Admin role for the active user** | Operator and viewer roles cannot reach any of the three endpoints. |
|
||||
| **Configured remote nodes in Settings → Nodes** | The two fan-out cards iterate the configured node list. A node missing its `api_url` or `api_token` shows up in the results with *Remote node not configured* per stack or per target. |
|
||||
| **Labels you intend to target** | Stop fleet by label and the autocomplete depend on labels existing on at least one node. See [Stack Labels](/features/stack-labels) for the authoring flow. |
|
||||
|
||||
## Behaviour and lifecycle
|
||||
|
||||
- **Always returns 200.** Every card is structured so that the HTTP status reflects the request shape, not the operational outcome. Partial failure is encoded in per-row fields, not in the status code.
|
||||
- **No retry, no scheduling, no undo.** Fleet Actions runs synchronously and is operator-driven; there is no background scheduler and no roll-back. For recurrence, use [Scheduled Operations](/features/scheduled-operations).
|
||||
- **Offline remotes still receive the request.** A node that is down at the moment of the action returns a transport-error row but does not block the fan-out across the rest of the fleet.
|
||||
- **Concurrent runs serialize per node.** Both fan-out cards take per-node locks before hitting Docker, so kicking off a second prune (or a second fleet stop and a per-label stop) while the first is in flight yields a calm *A bulk action is already running on this node* / *A prune is already running on this node* row rather than silent overlap.
|
||||
|
||||
## Limitations and non-goals
|
||||
|
||||
Fleet Actions is intentionally narrow in v1. The following are deliberately out of scope:
|
||||
|
||||
- **No fleet-wide bulk label assign.** Bulk label assign targets one node at a time. Re-tagging the same stacks on multiple nodes is two clicks of the node selector and two confirmations.
|
||||
- **No fleet-wide bulk start, restart, or update.** Stop is the only fleet-wide stack action today. Per-node multi-stack start, restart, and update live in the sidebar's Bulk mode.
|
||||
- **No label-set selectors.** Stop fleet by label matches one label name. Combinations like "stacks labelled A AND B" are not supported.
|
||||
- **No partial-stop ceiling.** The Stop card stops every stack the label matches; there is no "stop the first N" knob.
|
||||
- **No undo.** A stopped stack stays stopped until you start it again; a pruned image is gone until it is pulled or rebuilt.
|
||||
- **Approximate reclaim numbers.** The bytes the Prune card reports come from the Docker daemon and are best-effort, not authoritative.
|
||||
- **60-second remote timeout on fleet-stop, 120-second remote timeout per prune target.** A remote with many stacks or a very slow filesystem may produce timeout rows before the work fully completes. The action itself usually still finishes on the remote; the control instance just stopped waiting.
|
||||
|
||||
## Practical workflows
|
||||
|
||||
### Pause everything before a power event
|
||||
|
||||
Tag the stacks you want to bring down with a dedicated label (for example `evening-shutdown`). Run **Stop fleet by label** with that label name. The per-node breakdown confirms each stack's stop result; restart from the sidebar when power is back.
|
||||
|
||||
### Migrate a label taxonomy on one node
|
||||
|
||||
Split a coarse label like `prod` into `prod-edge` and `prod-core` by creating the new labels in **Settings → Labels** on the affected node, then opening **Bulk label assign**, picking that node, checking the stacks that should move, and toggling the new label set. The replace-not-append semantic guarantees the old label is removed in the same write.
|
||||
|
||||
### Free disk before a heavy deploy
|
||||
|
||||
Run **Prune Docker resources fleet-wide** with **Images** selected and **Managed only** scope. The reclaim shows up per node and gives you a quick read on which hosts had accumulated the most stale layers. Switch to **All unused** if you want the prune to reach workloads that Sencho does not manage.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
**Nothing happens when I click "Stop matching stacks".**
|
||||
Verify a label by that name exists on at least one node. The autocomplete suggestions are aggregated from every reachable node; if your label only exists on a node that is currently offline, it will not appear in the suggestions but the request still tries every node.
|
||||
<AccordionGroup>
|
||||
<Accordion title="A node shows '(no matching label)' but I created the label there">
|
||||
The fleet-stop match is by **label name**, not label ID. Confirm the label name on the affected node under **Settings → Labels**; a typo, a case mismatch, or a trailing space will leave the node out. Labels are scoped per node, so renaming the label on one node does not propagate to the others.
|
||||
</Accordion>
|
||||
<Accordion title="The autocomplete didn't suggest a label I know exists">
|
||||
The autocomplete fans out to **online** nodes only and aggregates label names from their label list response. A label that exists only on an offline node won't appear in the suggestions. The fleet-stop request still iterates every configured node, so typing the name by hand and submitting will reach the offline node when it returns.
|
||||
</Accordion>
|
||||
<Accordion title="Bulk label assign reports 'Invalid stack name' for one row">
|
||||
Stack names must match the standard validator: alphanumeric plus dash and underscore, no spaces, no path separators. The endpoint validates each assignment independently, so a single bad name does not block the rest of the batch. Fix the offending entry and re-run; the rows that already succeeded won't be re-applied.
|
||||
</Accordion>
|
||||
<Accordion title="Apply to N stacks button is disabled">
|
||||
The button requires at least one stack selected; the label selection can be empty (which clears assignments). Check that the node selector resolved (the **Loading…** spinner has cleared) and that the stack list is populated.
|
||||
</Accordion>
|
||||
<Accordion title="Fleet stop timed out on one remote node">
|
||||
Each remote node call carries a 60-second timeout. A remote with many stacks or a slow Docker daemon can outlast that budget. Check the affected remote's logs; the stop usually completed on the remote even though the control instance stopped waiting. Re-running the same fleet-stop is safe: stacks that are already stopped return the per-stack error *No containers found for this stack* and do not toggle anything else.
|
||||
</Accordion>
|
||||
<Accordion title="Prune across fleet reports 'A prune is already running on this node'">
|
||||
The per-node prune lock is held while the first prune is in flight. Wait for the first run to finish or click **Refresh** on the Fleet masthead to confirm it has cleared, then re-run. The lock is released automatically when the first run exits.
|
||||
</Accordion>
|
||||
<Accordion title="The Fleet Actions tab opens but the cards aren't there">
|
||||
Fleet Actions is a Skipper feature. On a Community license, the tab still opens and surfaces a calm explainer card pointing at the upgrade. Confirm the active license under **Settings → License**.
|
||||
</Accordion>
|
||||
<Accordion title="A node returns a transport error row for every stack or every target">
|
||||
The node is in **Settings → Nodes** but its `api_url` or `api_token` is missing, expired, or unreachable. Open **Settings → Nodes** on the control instance and click **Test connection** for the remote; fix the credential or the reachability, then re-run the action.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
**A node shows "Label not present" but I know I created the label there.**
|
||||
Labels are stored per node. The fleet stop matches by **name** only. If the label was renamed on one node, the new name is what gets matched. Check **Settings > Labels** on the affected node.
|
||||
## Where Fleet Actions fits
|
||||
|
||||
**Bulk label assign reports "Invalid stack name" for one row.**
|
||||
Stack names must be alphanumeric with dashes and underscores. The endpoint validates each entry independently, so a single bad name does not block the rest of the batch.
|
||||
Fleet Actions is one slice of the Fleet view. Each adjacent surface answers a different question.
|
||||
|
||||
**The Fleet Actions tab is missing from Fleet.**
|
||||
Confirm the active license is **Skipper** or **Admiral** under **Settings > License**. The tab itself is always visible, but the action cards only render at paid tiers.
|
||||
|
||||
**A node row reports "unreachable" after a fleet prune.**
|
||||
Sencho was unable to dispatch the prune to that remote node within the per-target timeout. Open the node detail and confirm it is online and that its API token is still valid. Sencho short-circuits later targets on the same node once one fails so a dead remote does not slow the whole operation.
|
||||
|
||||
**A target row shows `0 B` reclaimed.**
|
||||
That target had nothing to prune at the time of the run. For networks this is also the expected reading on every successful run, because Docker does not report bytes for network removals.
|
||||
| Related feature | What it covers | Why it is not Fleet Actions |
|
||||
|---|---|---|
|
||||
| [Stack Labels](/features/stack-labels) | Author the per-node label taxonomy that Fleet Actions targets. | Fleet Actions consumes labels; Stack Labels create and rename them. |
|
||||
| [Stack Sidebar · Bulk mode](/features/sidebar#bulk-mode) | Per-node, per-stack bulk start, stop, restart, and update. | Bulk mode handles a hand-picked subset of stacks on one node; Fleet Actions runs by selector across many nodes. |
|
||||
| [Scheduled Operations](/features/scheduled-operations) | Recurring or one-shot scheduled stack operations. | Scheduled Operations owns recurrence; Fleet Actions is always operator-initiated. |
|
||||
| [Fleet Federation](/features/fleet-federation) | Operator-driven placement controls (cordon, pin) for [Blueprints](/features/blueprint-model). | Federation steers declarative placement; Fleet Actions runs imperative operations on what already exists. |
|
||||
| [Fleet Sync](/features/fleet-sync) | Push-only replication of security rules from a control instance to its replicas. | Fleet Sync replicates state; Fleet Actions runs operations. |
|
||||
| [Multi-Node Management](/features/multi-node) | How nodes get added to the fleet (proxy or pilot mode) and how the license tier propagates. | Multi-Node Management is the prerequisite; Fleet Actions runs against whatever Multi-Node Management already configured. |
|
||||
| [Fleet View](/features/fleet-view) | The masthead, tab strip, and node grid that host the Fleet Actions tab. | Fleet Actions is one tab inside Fleet View. |
|
||||
| [Licensing](/features/licensing) | The full tier matrix and what each tier unlocks. | The single source of truth for the Skipper requirement called out at the top of this page. |
|
||||
|
||||
@@ -11,13 +11,13 @@ Create point-in-time snapshots of every `compose.yaml` and `.env` file across yo
|
||||
|
||||
## Creating a snapshot
|
||||
|
||||
1. Navigate to **Fleet View** and select the **Snapshots** tab
|
||||
2. Click **Create Snapshot** in the top-right corner
|
||||
3. An inline form appears with an optional description field (e.g. "Before v2 migration")
|
||||
1. Navigate to **Fleet** and select the **Snapshots** tab
|
||||
2. Click **Create Snapshot**
|
||||
3. An inline form appears above the snapshot list with an optional description field (e.g. "Before v2 migration")
|
||||
4. Click **Create** to capture files from every reachable node
|
||||
|
||||
<Frame>
|
||||
<img src="/images/fleet-backups/create-snapshot.png" alt="Inline form for creating a fleet snapshot with an optional description field" />
|
||||
<img src="/images/fleet-backups/create-snapshot.png" alt="Inline create snapshot form with optional description field, showing Create and Cancel buttons above the snapshot table" />
|
||||
</Frame>
|
||||
|
||||
During creation, Sencho connects to each node in parallel:
|
||||
@@ -35,23 +35,25 @@ Fleet snapshots can also be created automatically on a recurring schedule. Navig
|
||||
The snapshot list shows each snapshot in a table with the following columns:
|
||||
|
||||
- **Date** - when the snapshot was taken
|
||||
- **Description** - your optional label, or a prefix like "Scheduled snapshot" for automated ones
|
||||
- **Scope** - how many nodes and stacks were captured (e.g. "2 nodes, 22 stacks")
|
||||
- **Description** - your optional label, or a prefix like "Scheduled snapshot" for automated ones. If Cloud Backup is configured, an upload icon in this column marks snapshots that have been mirrored off-site.
|
||||
- **Scope** - how many nodes and stacks were captured (e.g. "3 nodes, 21 stacks")
|
||||
- **Warnings** - a warning icon with a count if any nodes were skipped, or "None"
|
||||
- **Actions** - **View** to open the detail view, and a delete button for admins
|
||||
- **Actions** - **View** to open the detail view, a cloud-upload icon for snapshots not yet mirrored (Admiral only), and a delete button for admins
|
||||
|
||||
<Frame>
|
||||
<img src="/images/fleet-backups/browse-snapshots.png" alt="Snapshot list showing date, description, scope, warnings, and action buttons" />
|
||||
<img src="/images/fleet-backups/browse-snapshots.png" alt="Snapshot list showing paginated rows with date, description with cloud upload indicator, scope, warnings, and action buttons" />
|
||||
</Frame>
|
||||
|
||||
## Snapshot detail view
|
||||
|
||||
Click **View** on any snapshot to open the detail view. The header card displays the snapshot's title (or "Untitled Snapshot"), who created it, when, and badge counts for nodes and stacks captured.
|
||||
Click **View** on any snapshot to open the detail view. Use **Back to Snapshots** in the top-left to return to the list.
|
||||
|
||||
The header shows the snapshot's title (or "Untitled Snapshot"), who created it, when, and badge counts for nodes and stacks captured.
|
||||
|
||||
Below the header, each node appears as a collapsible card. Expand a node to see its stacks, then expand a stack to see individual files. Each file has a **Preview** button that renders the file contents inline.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/fleet-backups/snapshot-detail.png" alt="Snapshot detail view showing a collapsible tree of nodes, stacks, and files with preview buttons" />
|
||||
<img src="/images/fleet-backups/snapshot-detail.png" alt="Snapshot detail view with a node expanded, showing a stack expanded with a compose file, Preview button, and Restore button" />
|
||||
</Frame>
|
||||
|
||||
If any nodes were unreachable during snapshot creation, a warning banner appears at the top of the detail view listing each skipped node and the reason it was skipped.
|
||||
@@ -62,12 +64,12 @@ Admins can restore individual stacks from any snapshot:
|
||||
|
||||
1. Open a snapshot's detail view
|
||||
2. Expand the node and stack you want to restore
|
||||
3. Click **Restore** below the file list
|
||||
3. Click **Restore** below the stack's file list
|
||||
4. A confirmation dialog appears. Optionally check **Redeploy stack after restore** to immediately apply the restored configuration.
|
||||
5. Click **Restore** to confirm
|
||||
|
||||
<Frame>
|
||||
<img src="/images/fleet-backups/restore-dialog.png" alt="Restore confirmation dialog with a redeploy checkbox" />
|
||||
<img src="/images/fleet-backups/restore-dialog.png" alt="Restore confirmation dialog showing the target stack and node name, a redeploy checkbox, and Cancel and Restore buttons" />
|
||||
</Frame>
|
||||
|
||||
Sencho writes the snapshot's files back to the target node:
|
||||
@@ -85,18 +87,24 @@ Admins can delete snapshots from the list view by clicking the trash icon on the
|
||||
## Cloud Backup
|
||||
|
||||
<Note>
|
||||
Cloud Backup requires an Admiral license. Configure it in **Settings → Cloud Backup**.
|
||||
Cloud Backup requires an Admiral license. Configure it in **Settings → System → Cloud Backup**.
|
||||
</Note>
|
||||
|
||||
Cloud Backup mirrors every fleet snapshot to off-site storage so your snapshots survive local disk failure. Two storage modes are supported.
|
||||
Cloud Backup mirrors every fleet snapshot to off-site storage so your snapshots survive local disk failure. The Cloud Backup settings page (reached via **Settings → System → Cloud Backup**) shows a header with your current scope, provider, storage used, and total snapshot count in the cloud. Two storage modes are supported.
|
||||
|
||||
### Sencho Cloud Backup (included)
|
||||
|
||||
A managed 500 MB allowance backed by Cloudflare R2, included with every Admiral license. Open **Settings → Cloud Backup**, choose **Sencho Cloud Backup**, and click **Activate**. Sencho exchanges your license key for scoped storage credentials and starts replicating new snapshots automatically. The settings panel shows your storage usage and lets you reprovision credentials if needed.
|
||||
A managed 500 MB allowance backed by Cloudflare R2, included with every Admiral license. Open **Settings → System → Cloud Backup**, choose **Sencho Cloud Backup (Included)**, and click **Activate**. Sencho exchanges your license key for scoped storage credentials and starts replicating new snapshots automatically.
|
||||
|
||||
Once active, the settings page shows a storage gauge (used / 500 MB and object count), a status message confirming auto-upload is on, and a **Reprovision** button to refresh credentials if needed. You can verify connectivity at any time with the **Test** button.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/cloud-backup/cloud-backup-active.png" alt="Cloud Backup settings page with Sencho Cloud Backup active, showing the storage gauge, auto-upload status, and the Cloud Snapshots list" />
|
||||
</Frame>
|
||||
|
||||
### Custom S3 (BYOB)
|
||||
|
||||
Bring any S3-compatible bucket: AWS S3, MinIO, Backblaze B2, Wasabi, or your own Cloudflare R2 token. Choose **Custom S3** in the storage-mode dropdown and fill in:
|
||||
Bring any S3-compatible bucket: AWS S3, MinIO, Backblaze B2, Wasabi, or your own Cloudflare R2 token. Choose **Custom S3 (BYOB)** in the storage-mode dropdown and fill in:
|
||||
|
||||
- **Endpoint URL** (e.g. `https://s3.us-east-1.amazonaws.com`, `https://my-minio.example.com:9000`)
|
||||
- **Region** (e.g. `us-east-1`, or `auto` for R2)
|
||||
@@ -107,34 +115,24 @@ Bring any S3-compatible bucket: AWS S3, MinIO, Backblaze B2, Wasabi, or your own
|
||||
|
||||
Click **Test** to verify connectivity, then **Save**. Secret keys are encrypted at rest. Sencho only sends them to your configured endpoint.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/cloud-backup/cloud-backup-custom.png" alt="Custom S3 configuration form with fields for Endpoint URL, Region, Bucket, Path Prefix, Access Key ID, Secret Access Key, and an Auto-upload toggle" />
|
||||
</Frame>
|
||||
|
||||
### Manual upload vs auto-upload
|
||||
|
||||
When auto-upload is on, every fleet snapshot is replicated as soon as it is created. Manual snapshots from the **Fleet → Snapshots** view upload asynchronously so the UI returns immediately; scheduled snapshots block on the upload so the task's success status reflects cloud durability.
|
||||
|
||||
To upload a single snapshot on demand, open the **Snapshots** tab in Fleet View. Each row that hasn't been mirrored yet shows a cloud-upload action next to the **View** button. Once a snapshot is in the cloud, a small cloud icon appears next to its description.
|
||||
To upload a single snapshot on demand, open the **Snapshots** tab in Fleet View. Each row that hasn't been mirrored yet shows a cloud-upload action in the Actions column. Once a snapshot is in the cloud, an upload icon appears next to its description.
|
||||
|
||||
### Browsing and downloading cloud snapshots
|
||||
|
||||
The **Cloud Snapshots** panel in **Settings → Cloud Backup** lists every archive currently in your bucket, with size and last-modified timestamp. Click the download icon to save a `.tar.gz` archive locally for off-host disaster recovery. Each archive contains a `metadata.json` describing the snapshot and a `nodes/` tree with the captured compose and environment files, organised by node and stack.
|
||||
The **Cloud Snapshots** panel in **Settings → System → Cloud Backup** lists every archive currently in your bucket, with size and last-modified timestamp. Click the download icon to save a `.tar.gz` archive locally for off-host disaster recovery. Each archive contains a `metadata.json` describing the snapshot and a `nodes/` tree with the captured compose and environment files, organised by node and stack.
|
||||
|
||||
### Restoring from a cloud snapshot
|
||||
|
||||
For in-place rollback, use the **Restore** action on the snapshot detail view as described above; the local copy is the source of truth for live restore. Cloud snapshots cover the disaster-recovery case where the local disk is gone: download the archive, extract it, and bring up a fresh Sencho instance pointed at the recovered files.
|
||||
|
||||
### Cloud Backup troubleshooting
|
||||
|
||||
#### Bad credentials
|
||||
|
||||
If **Test** reports an authentication error, double-check the Access Key ID, Secret Access Key, and bucket. Some providers require you to enable S3-compatible API access on the bucket separately. For MinIO, verify the user has read/write permission on the target bucket.
|
||||
|
||||
#### Over quota (Sencho Cloud Backup)
|
||||
|
||||
Sencho Cloud Backup has a 500 MB allowance per license. When you hit the cap, new uploads fail with a quota error. Delete older cloud snapshots from the **Cloud Snapshots** panel to free space. The local copies are unaffected.
|
||||
|
||||
#### Network timeout or 5xx error
|
||||
|
||||
Transient errors surface as a notification. Retry by clicking the cloud-upload action on the snapshot row, or wait for the next scheduled snapshot which will retry on its own. Persistent failures usually indicate an endpoint outage; verify the storage provider is reachable from your Sencho host.
|
||||
|
||||
## Access control
|
||||
|
||||
| Action | Admin | Node Admin | Deployer | Auditor | Viewer |
|
||||
@@ -144,6 +142,12 @@ Transient errors surface as a notification. Retry by clicking the cloud-upload a
|
||||
| Create snapshot | Yes | No | No | No | No |
|
||||
| Restore from snapshot | Yes | No | No | No | No |
|
||||
| Delete snapshot | Yes | No | No | No | No |
|
||||
| Upload snapshot to cloud | Yes | No | No | No | No |
|
||||
| Delete cloud snapshot | Yes | No | No | No | No |
|
||||
|
||||
<Note>
|
||||
Cloud backup actions (upload, delete cloud snapshots) also require an Admiral license.
|
||||
</Note>
|
||||
|
||||
## Storage
|
||||
|
||||
@@ -151,25 +155,26 @@ Snapshots are stored in Sencho's SQLite database. Compose files are typically sm
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### Snapshot shows skipped nodes
|
||||
|
||||
If a remote node is offline, unreachable, or its API token has expired, the node is skipped during snapshot creation. The snapshot list shows a warning icon with a count of skipped nodes. Open the snapshot's detail view to see which nodes were skipped and the reason for each.
|
||||
|
||||
Common causes:
|
||||
- The remote Sencho instance is stopped or restarting
|
||||
- The node's API URL or token was changed after it was added
|
||||
- A firewall or network issue is blocking the connection between nodes
|
||||
|
||||
To resolve, verify the remote node is running and reachable, then update the node's API URL and token in the Fleet settings if needed. Create a new snapshot after fixing the connectivity issue.
|
||||
|
||||
### Restore fails with "Target node no longer exists"
|
||||
|
||||
This error occurs when the node recorded in the snapshot has been removed from your fleet since the snapshot was taken. You cannot restore to a node that is no longer registered. Re-add the node first, then retry the restore.
|
||||
|
||||
### Restore fails with "No files found for this stack"
|
||||
|
||||
The stack you're trying to restore may not have been captured in the snapshot (for example, if the compose file was missing or unreadable at the time the snapshot was taken). Open the snapshot's detail view to verify which stacks and files are available.
|
||||
|
||||
### Diagnostic logging
|
||||
|
||||
If you need to investigate snapshot operations in detail, enable **Developer Mode** in **Settings > Developer**. This activates diagnostic logging for snapshot creation (per-node capture timing and file counts), restore operations, and scheduled snapshot execution. Diagnostic logs appear in the server's standard output with a `:debug` suffix.
|
||||
<AccordionGroup>
|
||||
<Accordion title="A snapshot shows skipped nodes">
|
||||
If a remote node is offline, unreachable, or its API token has expired, the node is skipped during snapshot creation. The list shows a warning icon with a count of skipped nodes, and the snapshot's detail view names each one along with the reason. Common causes are the remote Sencho instance being stopped or restarting, the node's API URL or token having been changed after it was added, or a firewall or network issue blocking the connection. Verify the remote is running and reachable, update the node's API URL and token in **Settings → System → Nodes** if needed, then create a new snapshot.
|
||||
</Accordion>
|
||||
<Accordion title="Restore fails with 'Target node no longer exists'">
|
||||
The node recorded in the snapshot has been removed from the fleet since the snapshot was taken. Snapshots reference nodes by registry ID, so a node that was deleted and re-added is treated as a different target. Re-add the node first, then retry the restore against the freshly registered row.
|
||||
</Accordion>
|
||||
<Accordion title="Restore fails with 'No files found for this stack'">
|
||||
The stack was not captured in the snapshot, usually because its compose file was missing or unreadable on disk at the time the snapshot was taken. Open the snapshot's detail view to verify which stacks and files are available, and pick a different snapshot if the one you have is incomplete.
|
||||
</Accordion>
|
||||
<Accordion title="Cloud Backup Test reports an authentication error">
|
||||
Double-check the Access Key ID, Secret Access Key, and bucket name; one wrong character is the most common cause. Some providers require S3-compatible API access to be enabled on the bucket separately from the credentials. For MinIO, confirm the user has read/write permission on the target bucket. After correcting the values, click **Test** again before saving.
|
||||
</Accordion>
|
||||
<Accordion title="Cloud uploads fail after hitting the 500 MB quota">
|
||||
Sencho Cloud Backup carries a 500 MB allowance per license. Once you hit the cap, new uploads fail with a quota error and the storage gauge in the settings page reads at or near 500 MB. Free space by deleting older archives from the **Cloud Snapshots** panel in **Settings → System → Cloud Backup**. Local snapshots are unaffected by cloud deletions, so the in-place restore path stays intact. To raise the ceiling, switch the storage mode to **Custom S3 (BYOB)** and point at a bucket you control.
|
||||
</Accordion>
|
||||
<Accordion title="Cloud upload returns a network timeout or 5xx error">
|
||||
Transient errors surface as a notification and leave the local snapshot in place. Retry by clicking the cloud-upload action on the snapshot row, or wait for the next scheduled snapshot which retries on its own. Persistent failures usually indicate an endpoint outage; verify the storage provider is reachable from your Sencho host (custom S3 endpoints often sit behind a different DNS or firewall path than the rest of your traffic).
|
||||
</Accordion>
|
||||
<Accordion title="I need more detail about a snapshot or restore operation">
|
||||
Enable **Developer Mode** under **Settings → Developer** to activate diagnostic logging for snapshot creation (per-node capture timing and file counts), restore operations, and scheduled snapshot execution. Diagnostic logs appear in the server's standard output with a `:debug` suffix.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
@@ -3,12 +3,16 @@ title: "Fleet Federation"
|
||||
description: "Operator-driven placement controls: cordon nodes and pin blueprints to specific nodes."
|
||||
---
|
||||
|
||||
The **Federation** tab is a placement-control surface for fleets running [Blueprints](/features/blueprint-model). It lets you steer where new deployments land without rewriting selectors or labels: mark a node unschedulable for new work, or force a specific blueprint to remain on a specific node regardless of selector matches.
|
||||
The **Federation** tab is the placement-control surface for fleets running [Blueprints](/features/blueprint-model). It lets you steer where new deployments land without rewriting selectors or labels: mark a node unschedulable for new work, or force a specific blueprint to remain on a specific node regardless of selector matches. Federation builds on the Blueprints reconciler, so the controls described here only affect blueprint-managed deployments.
|
||||
|
||||
Federation lives under **Fleet → Federation**.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/fleet-federation/federation-tab.png" alt="Fleet → Federation tab on a fleet with one cordoned node. The Cordoned nodes section reads '1 of 7' with the hint 'TOGGLE ON EACH NODE CARD' and lists sencho-test-01 with a REMOTE chip, a timestamp ('since 5/18/2026, 5:50:28 PM'), and the operator-supplied reason 'Patching kernel, back online in 30m'. The Pin policy section below reads 'Force a blueprint onto a specific node, overriding its selector.' and shows the empty-state message 'No blueprints yet. Create one in the Deployments tab to manage placement here.'" />
|
||||
</Frame>
|
||||
|
||||
<Note>
|
||||
Federation is an Admiral feature. The tab is hidden at the Community and Skipper tiers. Cordon and pin actions require an admin user role.
|
||||
Federation is an Admiral feature. Cordon and pin actions require an admin user role.
|
||||
</Note>
|
||||
|
||||
## Placement control, not placement automation
|
||||
@@ -17,76 +21,184 @@ Sencho's blueprint reconciler resolves selectors automatically: when a node grow
|
||||
|
||||
The model is deliberate: Sencho proposes placements; you confirm or override them. The reconciler never moves an existing deployment in response to cordon or pin. Cordon affects only *new* placements, and pin only changes which node the reconciler considers desired. Eviction from non-pinned nodes still flows through the existing state-review and confirmation prompts.
|
||||
|
||||
## Cordon a node
|
||||
In practice, Federation hands you four operator decisions:
|
||||
|
||||
Cordon marks a node as unschedulable. From the moment a node is cordoned:
|
||||
- **Don't schedule new work here for a while.** Cordon a node.
|
||||
- **Anchor this blueprint to one host.** Pin a blueprint.
|
||||
- **Hand the decision back to the reconciler.** Uncordon or unpin.
|
||||
- **Ride out a migration.** Combine pin (on the destination) with cordon (on the source) so the source drains naturally and the destination becomes the new home.
|
||||
|
||||
- **New blueprint deployments skip it.** A blueprint whose selector matches the cordoned node will not deploy a fresh stack there.
|
||||
- **Existing deployments on the node are unchanged.** Active stacks keep running. Drift checks keep running. Revision bumps still redeploy in place. The reconciler does not initiate a withdraw or evict because of a cordon.
|
||||
- **The cordon is visible to everyone.** Lower-tier viewers see a "Cordoned" pill on the node card so they understand why the node is not picking up new work, even though they cannot toggle the state.
|
||||
## Key capabilities
|
||||
|
||||
To cordon a node, open the node's card on **Fleet → Overview**, click the kebab menu (`⋯`) in the top-right corner, and choose **Cordon node**. You can attach an optional one-line reason (up to 256 characters); it surfaces in the Federation tab summary and in the audit log. Use **Uncordon node** from the same menu to lift the restriction.
|
||||
### Cordon a node
|
||||
|
||||
| Without cordon | With cordon |
|
||||
|---|---|
|
||||
| Selector match → new stack deploys | Selector match → reconciler skips this node |
|
||||
| Stack already deployed → drift-check + redeploy on revision | Same: existing deployment is unaffected |
|
||||
| Stack leaves selector → withdraw / evict_blocked | Same: cordon does not change the selector |
|
||||
Cordon marks a node as unschedulable. From the moment a node is cordoned, the reconciler skips it for new placements; existing stacks keep running, drift checks keep running, and revision bumps still redeploy in place. The cordon is visible to every tier (the read-only `Cordoned` pill on the node card) so non-Admiral operators understand why a node is not picking up new work. Uncordon to lift the restriction.
|
||||
|
||||
The Federation tab shows a read-only summary of currently cordoned nodes (name, type, when cordoned, optional reason). The action lives on the node card; the summary is for awareness.
|
||||
The cordon reason is free-form text (up to 256 characters). It surfaces in the Federation tab summary, in the cordon pill's tooltip on the node card, and on the audit log row for the action.
|
||||
|
||||
## Pin a blueprint to a node
|
||||
### Pin a blueprint to a node
|
||||
|
||||
Pinning a blueprint forces the reconciler to treat that blueprint as desired only on a single specific node, regardless of what its selector says. Use a pin when:
|
||||
Pinning a blueprint forces the reconciler to treat that blueprint as desired only on a single specific node, regardless of what its selector says. The pin overrides the selector entirely. Pin is the right tool when a workload depends on local state, must run on a specific host (a gateway, an integration target), or is in the middle of a host-to-host migration.
|
||||
|
||||
- You want one blueprint to stay on one host (a workload that depends on local state, a service that must run on the gateway node, a host-specific integration).
|
||||
- You are migrating a blueprint between nodes and want to hold it on the destination while you tear down the source.
|
||||
- You need to override an unintended selector match without rewriting the selector.
|
||||
When the target node is removed from the fleet, the pin clears automatically and the blueprint reverts to its selector behaviour on the next reconciliation tick.
|
||||
|
||||
Set or clear pins from **Fleet → Federation**, in the **Pin policy** table:
|
||||
### Audit visibility
|
||||
|
||||
Cordon, uncordon, and pin actions all flow through Sencho's standard audit log. Each action records the actor, the affected resource, and the optional reason (for cordon). Filter the **Audit** view by `node.cordon`, `node.uncordon`, or `blueprint.pin` to see the full history.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
| Requirement | Why it matters |
|
||||
|-------------|----------------|
|
||||
| **Admiral license active on the control instance** | Federation enforcement, the tab itself, and the cordon and pin controls all gate on the control instance's tier. Remote nodes inherit Admiral through the [proxy's tier assertion](/features/multi-node#license-enforcement-across-nodes), so a paid control plane covers the whole fleet. |
|
||||
| **Admin role for the active user** | Operator and viewer roles can read cordon state but cannot toggle it. Pin policy edits are also admin-only. |
|
||||
| **At least one Blueprint defined** | The Pin policy table is empty until you create a blueprint under **Fleet → Deployments**. Cordon does not require any blueprints; it only suppresses *new* placements from blueprints that exist later. |
|
||||
| **Active connection to each remote node** | Cordon state is written to the control instance's local database, but it only takes effect once the reconciler runs against the fleet's current node set. A remote node that is `Offline` still carries its cordon flag and resumes honouring it as soon as it comes back. |
|
||||
|
||||
## Step by step
|
||||
|
||||
### Cordon a node
|
||||
|
||||
Open **Fleet → Overview**. On the card for the node you want to cordon, click the kebab menu (`⋯`) in the top-right corner and choose **Cordon node**.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/fleet-federation/node-card-kebab-cordon.png" alt="Fleet Overview grid with the kebab menu open on a sencho-test-01 node card. The menu surfaces three items: Edit node, Delete node, and Cordon node (Ban icon). Other node cards (Local, Opsix, Pitt-Moba, sencho-pilot-test, sencho-test-02, SLX-Mars) are visible behind the open menu." />
|
||||
</Frame>
|
||||
|
||||
A confirmation dialog opens with an optional reason field. The reason is free-form text up to 256 characters; type a short note that will help the rest of your team understand why the node is out of rotation.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/fleet-federation/cordon-modal-with-reason.png" alt="Cordon confirmation dialog titled 'Cordon sencho-test-01'. The description reads 'Mark this node as unschedulable. New blueprint deployments will skip it. Existing deployments remain in place.' A 'Reason (optional)' text field is populated with 'Patching kernel, back online in 30m'. Cancel and 'Cordon node' buttons sit in the dialog footer." />
|
||||
</Frame>
|
||||
|
||||
Click **Cordon node** to commit. The card immediately gains a `Cordoned` pill, and the Federation tab's Cordoned nodes summary picks the node up on the next refresh.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/fleet-federation/node-card-cordoned-badge.png" alt="Fleet Overview grid with the sencho-test-01 node card showing the amber 'Cordoned' pill alongside the existing Online, remote, version, and Update available badges. Surrounding node cards (Local, Opsix, Pitt-Moba, sencho-pilot-test, sencho-test-02, SLX-Mars) have no cordon indicator." />
|
||||
</Frame>
|
||||
|
||||
Use **Uncordon node** from the same menu to lift the restriction. Uncordoning opens a short confirmation that reads *"Re-enable this node for new blueprint placements. Existing deployments are unchanged."* Confirming clears the flag and the pill disappears on the next refresh.
|
||||
|
||||
### Pin a blueprint to a node
|
||||
|
||||
Open **Fleet → Federation** and find the blueprint row in the **Pin policy** table.
|
||||
|
||||
| Column | What it shows |
|
||||
|---|---|
|
||||
| **Blueprint** | Name and short description. |
|
||||
| **Selector** | The selector you would otherwise match against (kept for context). |
|
||||
| **Pinned to** | A dropdown listing every node in the fleet plus an "(unpinned)" option. Changing it saves immediately and triggers a reconciliation. |
|
||||
| **Selector** | The selector you would otherwise match against, kept for context. |
|
||||
| **Pinned to** | A dropdown listing every node in the fleet plus an *(unpinned)* option. Changing it saves immediately and triggers a reconciliation. |
|
||||
| **Effective** | The desired set the reconciler will actually use: the pinned node when set, the selector summary otherwise. |
|
||||
|
||||
When a pin is set, the blueprint's [drift mode](/features/blueprint-model#drift-policy) and stateful classification still apply on the pinned node. Pin only changes *where* the blueprint is desired, not *how* it is reconciled there.
|
||||
Select a node from the **Pinned to** dropdown to pin. Select *(unpinned)* to clear the pin and hand the decision back to the reconciler. The change saves on selection (no separate Save button), and the **Effective** column updates in place.
|
||||
|
||||
The blueprint detail sheet renders a small read-only banner (`Pinned to <node>. Selector is overridden.`) and the deployment table marks the pinned row with a "Pinned" indicator so the override is obvious anywhere a blueprint surfaces.
|
||||
When a pin is active, the blueprint's [drift mode](/features/blueprint-model#drift-policy) and stateful classification still apply on the pinned node. Pin only changes *where* the blueprint is desired, not *how* it is reconciled there. The blueprint detail sheet renders a small read-only banner (`Pinned to <node>. Selector is overridden.`) and the deployment table marks the pinned row with a `Pinned` indicator so the override is obvious anywhere the blueprint surfaces.
|
||||
|
||||
### Pin overrides cordon
|
||||
### Verify
|
||||
|
||||
Cordon governs automatic placement; pin is an explicit operator decision. When the two collide (a blueprint pinned to a cordoned node), pin wins. The blueprint stays on (or deploys onto) the pinned node even though the node is otherwise unschedulable. This keeps cordon's cost predictable: cordoning a node cannot silently break a workload you previously chose to anchor there. If you want to remove the workload too, unpin the blueprint or withdraw it explicitly.
|
||||
After cordoning or pinning, check the verification surfaces:
|
||||
|
||||
- **Cordon:** the node card carries a `Cordoned` pill in the Fleet → Overview grid, and the node appears in the Federation tab's *Cordoned nodes* section with the timestamp and reason. The audit log records a `node.cordon` event.
|
||||
- **Pin:** the Pin policy row's **Effective** column reads the pinned node, the blueprint detail sheet shows the override banner, and the deployment table marks the pinned row. The audit log records a `blueprint.pin` event.
|
||||
|
||||
## Behaviour and lifecycle
|
||||
|
||||
| Action | Takes effect | Persists across restart |
|
||||
|---|---|---|
|
||||
| Cordon | Next reconciliation tick, for new placement and state-review decisions only | Yes |
|
||||
| Uncordon | Next reconciliation tick | Yes (the flag is cleared, not just hidden) |
|
||||
| Pin | Immediate; the next reconciliation tick treats the pinned node as the only desired target | Yes |
|
||||
| Unpin | Immediate; the reconciler reverts to selector behaviour | Yes (cleared) |
|
||||
| Pinned node deleted | Pin clears automatically as part of node deletion housekeeping | Yes (cleared) |
|
||||
|
||||
The reconciler only re-evaluates cordon when it is deciding where a blueprint should run *next* or whether the live state matches the desired state. Drift checks against existing containers, revision-driven redeploys, and manual deploys initiated from the stack editor all bypass the cordon flag by design. The flag is a hint to the placement layer, not a quarantine on the node.
|
||||
|
||||
The **Effective** column in the Pin policy table is computed live each time the page loads. It is not stored; only the pin itself persists.
|
||||
|
||||
## Security and audit
|
||||
|
||||
Federation actions require both an Admiral license and an admin user role. The `Cordoned` pill on the node card stays visible at every tier as a read-only signal, so a non-Admiral operator can still understand why a node is skipping new work.
|
||||
|
||||
Every cordon, uncordon, and pin action is captured in the audit log. Each row carries the actor, the action (`node.cordon`, `node.uncordon`, `blueprint.pin`), the affected resource, and the cordon reason where applicable. Filter the **Audit** view by these action names for the full history.
|
||||
|
||||
Federation is hub-only. The cordon flag and the pin live on the control instance, and the blueprint reconciler that reads them also runs on the control instance. Cordoning a remote node does not require any change on the remote itself.
|
||||
|
||||
## Limitations and non-goals
|
||||
|
||||
Federation v1 ships the two controls above and nothing more. The following are deliberately out of scope:
|
||||
|
||||
- **Drain node.** Evacuating all blueprints from a node depends on volume migration, which is operator-driven for stateful workloads in the current release.
|
||||
- **Capacity planning.** Predictive resource utilisation belongs to a later iteration once there is real-world fleet usage data to calibrate against.
|
||||
- **Auto-eviction on cordon.** Cordon never withdraws or evicts an existing deployment. Use the deployment table for explicit withdraw or eviction confirmations.
|
||||
- **In-place-redeploy block.** Cordon does not stop revision-driven redeploys on the cordoned node. Bumping a blueprint's revision still redeploys the existing stack in place on every node where it already runs, cordoned or not.
|
||||
- **Manual deploy interception.** Cordon affects the blueprint reconciler only. A manual deploy initiated from the stack editor still lands on whichever node you target.
|
||||
- **Drift-mode override.** Pin does not change the blueprint's drift policy or stateful classification on the pinned node. Both flow through unchanged.
|
||||
- **Reason length.** The cordon reason is capped at 256 characters and stored verbatim.
|
||||
|
||||
## Practical workflows
|
||||
|
||||
### Take a node out of rotation for OS patching
|
||||
|
||||
Cordon the node with a reason that names the work (e.g. *"Patching kernel, back online in 30m"*). The Federation tab summary and the node-card tooltip surface that reason for the rest of your team. Existing stacks keep running while you reboot and patch; new blueprint placements skip the node until you uncordon.
|
||||
|
||||
### Migrate a stack between hosts
|
||||
|
||||
Pin the blueprint to the destination node first. The next reconciliation tick deploys onto the destination. Once health checks pass on the destination, withdraw the original deployment from the source via the deployment table. The pin holds the workload on the destination while you tear down the source, so there is no window where the reconciler can reverse the move.
|
||||
|
||||
### Anchor a gateway-only blueprint
|
||||
|
||||
Pin a gateway blueprint (reverse proxy, ingress, tunnel) to the node that owns the public IP. Even if the selector matches a second node, the pin overrides; only the gateway host receives the deployment. If the gateway node is also cordoned for maintenance, the pin wins: the blueprint stays on the pinned node even while it is otherwise unschedulable.
|
||||
|
||||
## Pin overrides cordon
|
||||
|
||||
Cordon governs automatic placement; pin is an explicit operator decision. When the two collide (a blueprint pinned to a cordoned node), pin wins. The blueprint stays on (or deploys onto) the pinned node even though the node is otherwise unschedulable. This keeps cordon's cost predictable: cordoning a node cannot silently break a workload you already chose to anchor there. If you want to remove the workload too, unpin the blueprint or withdraw it explicitly.
|
||||
|
||||
### Pinning a stateful blueprint
|
||||
|
||||
Pinning a stateful blueprint that is currently deployed on multiple nodes shrinks the desired set to one node. On the next reconciliation tick, the non-pinned deployments enter `evict_blocked` and wait for an explicit eviction confirmation. This is the same flow that protects stateful workloads from automatic withdraw when a selector changes. Confirm each eviction from the deployment table or unpin the blueprint to restore the original desired set.
|
||||
|
||||
## Out of scope
|
||||
|
||||
Federation v1 ships the two controls above and nothing more. Items deliberately deferred:
|
||||
|
||||
- **Drain node.** Evacuating all blueprints from a node depends on volume migration, which is operator-driven for stateful workloads in the current release.
|
||||
- **Capacity planning.** Predictive resource utilisation belongs to a later iteration once there is real-world fleet usage data to calibrate against.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="A blueprint refuses to deploy on a node that matches its selector">
|
||||
Check whether the node is cordoned. Cordoned nodes are excluded from new placements; the Federation tab summary lists every cordoned node and the reason. Uncordon the node to re-enable automatic placement, or pin the blueprint explicitly to deploy onto a cordoned node.
|
||||
</Accordion>
|
||||
<Accordion title="A cordoned node still received a new deployment">
|
||||
Check whether the blueprint is pinned to that node. Pin overrides cordon by design, so a blueprint pinned to a cordoned node still deploys on the next reconciliation tick. Unpin the blueprint to restore selector-driven placement, then cordon will be honoured.
|
||||
</Accordion>
|
||||
<Accordion title="A pinned blueprint deployed somewhere unexpected">
|
||||
The pin is the source of truth. Open Federation and confirm the pinned node matches your intent. The Effective column shows what the reconciler will use. Selector matches are ignored while a pin is set.
|
||||
The pin is the source of truth. Open Federation and confirm the **Pinned to** value matches your intent. The **Effective** column shows what the reconciler will use. Selector matches are ignored while a pin is set.
|
||||
</Accordion>
|
||||
<Accordion title="Pinning a blueprint left rows in evict_blocked on the other nodes">
|
||||
That is the stateful guard working as intended. The reconciler does not auto-evict stateful workloads; each leftover deployment must be confirmed from the deployment table, just like a selector change would require. Unpinning the blueprint restores the original desired set if you want to keep all of them.
|
||||
</Accordion>
|
||||
<Accordion title="I cordoned a node but the badge has not appeared on the card">
|
||||
The Fleet Overview grid polls every 30 seconds, so the `Cordoned` pill can lag the action by up to that long. The Federation tab summary and the audit log update immediately; if the audit log shows the cordon action but the badge is still missing, click **Refresh** on the Fleet page to force an immediate re-fetch.
|
||||
</Accordion>
|
||||
<Accordion title="The Federation tab is not visible">
|
||||
Federation requires an Admiral license. On Skipper, Federation is hidden by design and the rest of the blueprint surface (catalog, deployments, drift) remains available. The cordoned pill on a node card is visible to all tiers; only the toggle is gated.
|
||||
Federation requires an Admiral license. The rest of the blueprint surface (catalog, deployments, drift) is available on lower tiers, and the read-only `Cordoned` pill on a node card is visible everywhere. If Federation is missing on an Admiral license, sign in as an admin user; cordon and pin actions are admin-only.
|
||||
</Accordion>
|
||||
<Accordion title="The Pin policy table is empty even though I created blueprints">
|
||||
Federation reads from the same Blueprints registry that powers **Fleet → Deployments**. If the Deployments tab shows blueprints but Federation does not, refresh the Fleet page (the Pin policy table caches its source list when the tab mounts). If Deployments is also empty, no blueprint has been saved yet; create one there first.
|
||||
</Accordion>
|
||||
<Accordion title="A pinned node was deleted">
|
||||
The pin clears automatically when its target node is removed from the fleet. The blueprint reverts to its selector behaviour on the next reconciliation tick.
|
||||
The pin clears automatically when its target node is removed from the fleet. The blueprint reverts to its selector behaviour on the next reconciliation tick. There is no manual cleanup step.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Where Federation fits
|
||||
|
||||
Federation is one tab in a larger Fleet view, and it focuses on a narrow slice of fleet operations: placement decisions for declarative blueprints. The other surfaces that share the Fleet view cover orthogonal concerns:
|
||||
|
||||
| Related feature | What it covers | Why it is not Federation |
|
||||
|---|---|---|
|
||||
| [Fleet View](/features/fleet-view) | The masthead, tab strip, and node-grid that contains the Federation tab. Read this for the broader Fleet UI. | Federation is one of the tabs hosted here. |
|
||||
| [Multi-Node Management](/features/multi-node) | How nodes are added, how the control plane reaches them (Proxy or Pilot Agent), and how the license tier propagates. | Federation operates on the control plane only; node connectivity is handled before Federation runs. |
|
||||
| [Pilot Agent](/features/pilot-agent) | The outbound WebSocket tunnel mode used by nodes behind NAT or restrictive ingress. | Federation does not change the transport; cordon and pin both work regardless of node mode. |
|
||||
| [Sencho Mesh](/features/sencho-mesh) | Cross-node container networking so a service on one node can reach a service on another by alias. | Mesh routes packets between containers; Federation decides which nodes receive blueprint deployments. |
|
||||
| [Fleet Actions](/features/fleet-actions) | Bulk operations across labelled nodes (restart, stop). | Fleet Actions runs imperative operations on existing deployments; Federation steers declarative placement decisions. |
|
||||
| [Fleet Sync](/features/fleet-sync) | Push-only replication of security policies (scan policies, CVE suppressions) from a control instance to replica instances. | Fleet Sync replicates *security state*, not placement; the two features do not interact. |
|
||||
| [Blueprints](/features/blueprint-model) | The declarative deployment model whose reconciler Federation steers. | Required reading: without Blueprints, Federation has nothing to do. |
|
||||
| [Licensing](/features/licensing) | The full tier matrix and what each tier unlocks. | The single source of truth for the Admiral requirement called out at the top of this page. |
|
||||
|
||||
Federation is not Fleet Sync, not Mesh, and not the remote-node proxy. The cordon flag affects only declarative blueprint deployments, not manually deployed stacks. If you are looking for a way to take an entire node fully out of service (existing deployments included), see the deployment table's withdraw flow on each affected blueprint; cordon by itself is intentionally non-destructive.
|
||||
|
||||
@@ -1,101 +1,204 @@
|
||||
---
|
||||
title: "Fleet Secrets"
|
||||
description: "Centralized, encrypted, versioned env-var bundles you can push to labeled nodes' stacks."
|
||||
description: "Author one set of environment-variable bundles on the control instance, push them to labeled nodes' stacks, and keep an audit trail of every change."
|
||||
---
|
||||
|
||||
Fleet Secrets gives you one place to author the environment variables a stack needs, and one action to push them to every node that runs that stack. Bundles are encrypted at rest, every save bumps a version, and every push is recorded in the audit log with a per-node diff.
|
||||
The **Secrets** tab on the Fleet view is where you bundle environment variables once and push them, version by version, to every node in the fleet that runs a copy of the same stack. Each bundle is a named set of `KEY=value` pairs encrypted at rest, every save is a new immutable version with a change note, and every push records a per-node diff in the audit log.
|
||||
|
||||
Fleet Secrets is available on **Skipper** and **Admiral** licenses. The tab is hidden on Community.
|
||||
The unit of work is the **bundle**. One bundle has one current `kv` payload; pushing it overlays that payload onto a chosen env file inside a chosen stack on every node a label selector matches. The wizard is three tabs (**Target**, **Preview**, **Results**) and reports each per-node outcome inline.
|
||||
|
||||
## What's in scope
|
||||
<Frame>
|
||||
<img src="/images/fleet-secrets/overview.png" alt="Fleet view with the Secrets tab selected. Heading 'Secret bundles' with a 'New bundle' button on the right; below it a table with columns Name, Description, Version, Keys, Updated, and three icon actions per row (edit, send, delete). One row visible: app-config · Shared environment for the app stack · v2 · 5 keys, last updated a few moments ago." />
|
||||
</Frame>
|
||||
|
||||
| Today | Not yet |
|
||||
|-------|---------|
|
||||
| Bundles of `KEY=value` env pairs | Certificates, key files, credential JSON blobs |
|
||||
| Push to nodes by label selector | Pin a bundle to specific node IDs |
|
||||
| Versioning on every save with a change note | Scheduled or auto-rotation |
|
||||
| Diff preview before each push | Multi-region or per-environment overlays |
|
||||
<Note>
|
||||
Fleet Secrets is a Skipper feature. Every action requires an admin user role.
|
||||
</Note>
|
||||
|
||||
If a feature is not in the table above, it's deferred to a later release. Fleet Secrets in 1.0 is intentionally focused on env-var drift across copies of the same stack.
|
||||
## What Fleet Secrets covers (and what it doesn't)
|
||||
|
||||
Fleet Secrets owns env-var bundles that span more than one node. If your goal is in the table below, the dedicated surface is the right place.
|
||||
|
||||
| Goal | Where to do it |
|
||||
|------|----------------|
|
||||
| Edit `.env` for one stack on one node | The [stack editor](/features/stack-file-explorer) on that node |
|
||||
| Ship a TLS cert, a key file, or a JSON credential blob to a stack | The stack editor (Fleet Secrets only carries `KEY=value` pairs in v1) |
|
||||
| Rotate a value on a recurring schedule | Not yet. Pushes are operator-initiated in v1; track the rotation in [Scheduled Operations](/features/scheduled-operations) if you want a reminder |
|
||||
| Restart the targets after a push | [Fleet Actions](/features/fleet-actions) |
|
||||
| Replicate scan policies and CVE suppressions to remote control instances | [Fleet Sync](/features/fleet-sync) |
|
||||
| Inspect who pushed which bundle when | [Audit Log](/features/audit-log) |
|
||||
|
||||
## The bundle is the unit
|
||||
|
||||
A **bundle** has a name, a description, a current version number, and a `kv` payload of `KEY=value` pairs. Each save creates a new version row that is immutable; the bundle's `current_version` field points at the newest one. The payload is encrypted with the instance's AES-256-GCM data key, the same key that protects MFA secrets and registry credentials, so the kv values never sit on disk in plaintext.
|
||||
|
||||
A **push** is a separate action. It reads the bundle's current version, walks every node the selector resolves to, computes a per-node diff against the chosen env file inside the chosen stack, and writes the merged result. Pushes are sequential per bundle, so only one push for a given bundle runs at a time. Each per-node write is independent, so one failure does not roll back the others.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
| Requirement | Why it matters |
|
||||
|---|---|
|
||||
| Skipper or Admiral license on the control instance | The tab and the underlying actions are paid; this is the same gate that opens Blueprints and Fleet Federation |
|
||||
| Admin user role | Bundle CRUD and push run as the signed-in operator and write authored-by rows into the audit log |
|
||||
| At least one stack on at least one node | Pushes target an existing stack directory; the wizard does not create stacks |
|
||||
| The target stack's compose declares the env file via `env_file:` | The env-file dropdown in the push wizard reads `env_file:` entries from a representative node's compose; a stack with only an inline `environment:` block will not show up |
|
||||
| The control instance can reach the remote node's API URL | Each remote write is an HTTP call from the control instance to the remote's `/api/stacks/.../env`; an unreachable remote is reported as a per-node failure, not a transport error for the whole push |
|
||||
|
||||
## Create a bundle
|
||||
|
||||
1. Open **Fleet → Secrets**.
|
||||
2. Click **New bundle**.
|
||||
3. Give it a name (letters, digits, dot, dash, underscore; 2–64 chars) and an optional description.
|
||||
4. Add `KEY=value` rows. Keys must match the env-var convention (`A-Z`, digits, underscore; cannot start with a digit). The eye icon toggles value visibility; the copy icon copies a single value to the clipboard.
|
||||
5. Add a **Change note** (optional but recommended) explaining why this version exists.
|
||||
6. Click **Save**. The bundle is now version `v1`.
|
||||
3. Give it a name. Names are 2-64 characters, alphanumerics plus space, dot, dash, and underscore, and must start and end with an alphanumeric.
|
||||
4. Optionally add a description; the description is a free-text field and is shown in the bundle list.
|
||||
5. Add `KEY=value` rows. Keys follow the env-var convention: start with a letter or underscore, then any mix of letters, digits, and underscores; case-sensitive. The eye icon toggles per-row visibility, and the copy icon copies a single value to the clipboard.
|
||||
6. Add a **Change note** if you want one (it appears next to the version row on the **Versions** tab).
|
||||
7. Click **Save**. The bundle is now version `v1`.
|
||||
|
||||
Saving the bundle encrypts the payload with the instance's data key (AES-256-GCM via the same key that protects MFA secrets and registry credentials). The plaintext only ever exists in memory while you're editing.
|
||||
<Frame>
|
||||
<img src="/images/fleet-secrets/create.png" alt="New secret bundle sheet. Breadcrumb 'Fleet › Secrets › New bundle', heading 'New secret bundle', Save button in the action bar. Identity section with Name 'app-config' and Description 'Shared environment for the app stack'. Key/value pairs section showing three rows: PUID=1000, PGID=1500, TZ=America/Los_Angeles, each row carrying eye, copy, and delete icons; an 'Add key' button beneath. Change note (optional) field reads 'Initial app config bundle'." />
|
||||
</Frame>
|
||||
|
||||
Saving encrypts the payload in memory and writes the ciphertext to the `secret_versions` table. The plaintext values only live in the editor while you're filling them in; once saved they are read back by decrypt only when the bundle is opened or a push references that version.
|
||||
|
||||
## Edit and version
|
||||
|
||||
Editing a bundle creates a new version. Each version is immutable once saved. Open a bundle and switch to the **Versions** tab to see who saved each version and any change notes attached.
|
||||
Open an existing bundle from the list and edit the keys, the description, or the change note. Saving creates a new version row and bumps the bundle's `current_version`; the previous version stays in history and is never rewritten. There is no row-level lock, so concurrent edits use last-write-wins semantics: both saves succeed, both appear in **Versions**, and the higher version number is current.
|
||||
|
||||
Saves are last-write-wins. There's no row-level lock; if two operators edit the same bundle concurrently, the later save creates the higher version number and the earlier save remains visible in history.
|
||||
The **Versions** tab on the bundle sheet lists every version with its key count, change note, who saved it, and when.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/fleet-secrets/versions.png" alt="Bundle sheet for app-config in edit mode with the Versions tab selected. Two version rows visible: v2 with 5 keys, change note 'Switch app to debug logging and wire Sentry', authored by admin at the timestamp on the right; v1 with 3 keys, change note 'Initial app config bundle', authored by admin at an earlier timestamp." />
|
||||
</Frame>
|
||||
|
||||
The Versions tab is read-only history. There is no **Restore v1** action in v1; if you need to revert, open the current version, edit the rows back to the values you want, add a change note explaining the revert, and save. The Versions tab will then show the revert as a new version on top of the others.
|
||||
|
||||
## Push to nodes
|
||||
|
||||
A push reads the bundle's current version, computes the diff against each target's existing env file, and writes the merged result.
|
||||
The push wizard walks Target → Preview → Results in order. Once a stage is populated you can click back into earlier tabs to adjust; you cannot skip Preview before executing.
|
||||
|
||||
1. From the bundle row, click the **Send** action to open the push wizard.
|
||||
2. **Target tab:**
|
||||
- Pick one or more node labels.
|
||||
- Choose **any** (a node matches if it has any of the selected labels) or **all** (a node must have all selected labels).
|
||||
- Enter the **Stack name** that the bundle should be applied to. The stack must exist on each target node by that name.
|
||||
- Pick the **Env file** to write. The dropdown lists every file declared by the stack's compose on a representative target node — the canonical `.env` plus any files referenced via `env_file:` directives. Per-node compose files can differ; nodes that do not declare the chosen file are reported as failed for that push.
|
||||
3. Click **Preview**.
|
||||
4. **Preview tab:** the table shows one row per matched node. Each row reports a quick header (`+added · ~changed · ·unchanged`) and an expander with the per-key diff. Drift — keys present on the target that are absent from the bundle — is shown as informational `removed` rows. The push will not delete those keys.
|
||||
5. Click **Push to N nodes**.
|
||||
6. **Results tab:** per-node status pills (`ok`, `failed`, `skipped`). Hover a failure to see the error.
|
||||
### Target tab
|
||||
|
||||
Pushes are sequential per bundle to make audit ordering predictable. Only one push for a given bundle can run at a time; a second concurrent push attempt returns `409 Conflict` until the first finishes.
|
||||
Pick the **label selector** that matches the nodes you want to push to. Toggle **any** (a node matches if it carries any of the selected labels) or **all** (a node must carry every selected label). Enter the **Stack name** that the bundle should be applied to; the stack must exist on each matched node by that exact name. Pick the **Env file** from the dropdown. The dropdown is populated from the `env_file:` directives on a representative target's compose, so per-node differences are possible.
|
||||
|
||||
## Merge semantics: overlay
|
||||
<Frame>
|
||||
<img src="/images/fleet-secrets/target.png" alt="Push wizard Target tab. Breadcrumb 'Fleet › Secrets › app-config › Push', heading 'Push app-config' with 'v2 · 5 keys' subtitle. Preview button at the top. Target nodes section: 'Match any / all of these labels' with the any toggle selected; below it the labels combobox is open showing one option 'app-tier' with a checkmark. Target stack section: Stack name 'secrets-demo', Env file dropdown showing '.env', helper text below explaining that env-file options come from a representative target." />
|
||||
</Frame>
|
||||
|
||||
Fleet Secrets uses **overlay** merge:
|
||||
Click **Preview** when you're done. Preview is a read-only walk: it asks each matched node what's currently in the chosen env file and computes the diff against the bundle. It does not write anything.
|
||||
|
||||
- A key in the bundle that is missing from the target is **added**.
|
||||
- A key in both with a different value is **changed** to the bundle's value.
|
||||
- A key in both with the same value is **unchanged**.
|
||||
- A key on the target that is **not in the bundle** is preserved on the target. The diff shows it under `removed (drift)` so you can see the divergence, but the push leaves it alone.
|
||||
### Preview tab
|
||||
|
||||
This is deliberately conservative for a v1 push action. If a key truly needs to be removed from a target's env file, do it in the stack editor on that node, then push the bundle to bring the rest of the keys in line.
|
||||
Each row is a matched node with a short header reading `+added · ~changed · ·unchanged`, plus a `drift N` count when the target has keys the bundle doesn't. Expand a row to see the per-key diff: status labels are **ADDED**, **CHANGED**, **REMOVED**, and **UNCHANGED**. A `removed` row is informational; the push will not delete it.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/fleet-secrets/preview.png" alt="Push wizard Preview tab with two node rows expanded. Local: header reads +2 ~2 ·1 drift 1, with per-key rows ADDED LOG_LEVEL, REMOVED OLD_LEGACY_FLAG, CHANGED PGID, UNCHANGED PUID, ADDED SENTRY_DSN, CHANGED TZ. Opsix: header reads +5 ~0 ·0, with five ADDED rows for LOG_LEVEL, PGID, PUID, SENTRY_DSN, TZ. 'Push to 2 nodes' button at the top." />
|
||||
</Frame>
|
||||
|
||||
Preview is also the safety net. The diff you see is the diff the next push will apply, so if a row looks wrong, fix the bundle first and re-run Preview.
|
||||
|
||||
### Results tab
|
||||
|
||||
Once you click **Push to N nodes**, the wizard fans out to each matched node sequentially. Each per-node row reports `ok` with the same `+added · ~changed · ·unchanged` header it had in Preview, or `failed` with the exact backend error string. Hover a failed row for the full message.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/fleet-secrets/results.png" alt="Push wizard Results tab. Two rows: Local with a green check icon and counts +2 ~2 ·1; Opsix with a red error icon and the message 'failed to write env (HTTP 500: {error: Failed to save env file...})'. 'Done' button at the top." />
|
||||
</Frame>
|
||||
|
||||
Per-bundle serialization means a second push attempt while the first is running returns `A push for this secret is already running`; wait for the first to finish, then retry. Per-node serialization within a single push keeps the audit-log ordering predictable.
|
||||
|
||||
## Import from a stack
|
||||
|
||||
If a node already has a working env file, you don't need to retype it into a new bundle. Open the bundle in edit mode and click **Import from stack** next to the **Add key** button. Pick a **Node**, type the **Stack name** and **Env file** basename, and click **Import**.
|
||||
|
||||
The import reads the named env file from the chosen node and overlays its kv into the editor: keys that match existing rows have their values updated in place, keys that don't appear yet append at the bottom. The editor state stays dirty until you click **Save**, so you can review, tweak, or discard the import before it becomes a version.
|
||||
|
||||
Import is the natural seeder for a fresh bundle (read the working env off one node, then save it as v1) and a useful read-back for triage (point the bundle at a deployed stack to see what's currently set).
|
||||
|
||||
## Behaviour and lifecycle
|
||||
|
||||
| Aspect | Behaviour |
|
||||
|---|---|
|
||||
| Push execution | Sequential per bundle; sequential per node within a push |
|
||||
| Concurrency lock | One in-flight push per bundle; second attempt returns `A push for this secret is already running` |
|
||||
| Partial failure | Each per-node write is independent; one failure does not roll back others |
|
||||
| Per-node timeout | 10 seconds per remote HTTP call |
|
||||
| Encryption at rest | AES-256-GCM with the instance data key; only the kv payload is encrypted, bundle metadata (name, description, version numbers, change notes, who saved when) is plaintext SQLite |
|
||||
| Versioning | Immutable; new save creates a new row in `secret_versions`, bundle's `current_version` is bumped |
|
||||
| Concurrent edits | Last-write-wins (no row-level lock); both versions land in history, the later one is current |
|
||||
| Drift handling | Keys present on the target but missing from the bundle are preserved and reported as informational `drift` in Preview; the push never deletes them |
|
||||
| Live progress | None. The wizard is request/response; Results renders only after every per-node write completes |
|
||||
| Audit | Every CRUD, import, preview, and execute action writes a row into the audit log; every executed push also writes one `secret_pushes` row per target node |
|
||||
|
||||
## Audit trail
|
||||
|
||||
Every mutating action shows up in **Audit Log**:
|
||||
Every mutating action surfaces in the [Audit Log](/features/audit-log) with the exact summary string below.
|
||||
|
||||
| Action | Audit summary |
|
||||
|--------|---------------|
|
||||
|---|---|
|
||||
| Create bundle | `Created secret` |
|
||||
| Update bundle (new version) | `Updated secret: <id>` |
|
||||
| Update bundle | `Updated secret: <id>` |
|
||||
| Delete bundle | `Deleted secret: <id>` |
|
||||
| Import env from a stack | `Imported env into secret: <id>` |
|
||||
| Preview a push | `Previewed secret push: <id>` |
|
||||
| Execute a push | `Pushed secret: <id>` |
|
||||
|
||||
The execute-push action also writes one row per target node in `secret_pushes`, capturing the bundle id, version, push id (a UUID grouping all rows from the same push), node id, stack name, env file, status, error message, and per-node `added` / `changed` / `unchanged` counts.
|
||||
In addition to the audit-log row, every executed push writes one row per target node into a `secret_pushes` table. Each row captures the bundle id, the version that was pushed, a UUID that groups all rows from the same push, the node id, the stack name, the env file basename, the status (`ok` or `failed`), the error message when the status is `failed`, and the per-node added / changed / unchanged counts. The audit-log summary points at the bundle and version; the `secret_pushes` rows are the per-target detail.
|
||||
|
||||
## Limitations and non-goals
|
||||
|
||||
v1 is intentionally narrow.
|
||||
|
||||
- **Payloads are `KEY=value` only.** Certificates, key files, and JSON credential blobs are out of scope. The same compose-level `env_file:` mechanism is what makes the push possible; non-env-file targets need a different surface.
|
||||
- **Targeting is by label selector.** Pinning a bundle to a specific list of node IDs is not exposed in the wizard.
|
||||
- **No scheduled rotation.** Every push is operator-initiated. Use [Scheduled Operations](/features/scheduled-operations) for a reminder if you need recurring rotation.
|
||||
- **One bundle, one kv set.** Multi-region overlays, per-environment variants, or merged-from-multiple-bundles pushes are out of scope; if you need three different kv shapes, create three bundles.
|
||||
- **Drift keys are never deleted.** A key the target has but the bundle doesn't is shown as informational `drift` and preserved. If you want it gone, edit that node's env file directly in the [stack editor](/features/stack-file-explorer), then re-push the bundle.
|
||||
- **No live progress.** The wizard does not stream per-node updates while a push runs; Results renders only after the last node has been written.
|
||||
- **No version restore action.** Versions are immutable history; to revert, edit the current version back to the old values and save a new version with a change note explaining the revert.
|
||||
|
||||
## Practical workflows
|
||||
|
||||
- **Rotate a database password across every replica.** Edit the bundle's `DB_PASSWORD` row, add a change note ("rotate 2026-Q2"), and push to nodes labeled `db-replica` against the `app` stack's `.env`. Once Results is clean, follow with **Fleet Actions → Stop fleet by label** + redeploy on the affected nodes so the running containers pick up the new value.
|
||||
- **Seed a fresh node from a known-good one.** Create an empty bundle, click **Import from stack** in edit mode, point it at the node that already has the working env, save as v1, and push to the new node.
|
||||
- **Bring a divergent fleet back in line.** Run **Preview** without pushing to see, node by node, exactly which keys have drifted. The diff is the source of truth for what the next push will write; share it with the team before committing.
|
||||
- **Triage a failed push.** Open the **Audit Log** and find the `Pushed secret: <id>` row. The matching `secret_pushes` entries name every target, its status, and the error. A common cause is `env file '.env' not found` when the target's compose lacks an `env_file:` line; declare the file in compose, redeploy that stack, and re-run the push.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### "env file '<name>' not found on node <X>"
|
||||
<AccordionGroup>
|
||||
<Accordion title="A node landed in the 'failed' column">
|
||||
Hover the row for the full error. The two most common causes are (1) the target stack does not declare the chosen env file via `env_file:` in its compose, so the remote write target does not resolve, and (2) the stack directory does not exist on that node. Open the stack on the affected node, confirm both, fix the cause, and re-run the push. Pushes are idempotent under the overlay merge, so re-running is safe.
|
||||
</Accordion>
|
||||
<Accordion title="The wizard says 'A push for this secret is already running'">
|
||||
The per-bundle lock is held while the first push is in flight. Wait for it to finish; the lock is released automatically when the push exits, success or failure. Per-bundle serialization is what keeps the audit-log ordering predictable, so the lock is intentional.
|
||||
</Accordion>
|
||||
<Accordion title="The Env file dropdown is empty or missing the file I expected">
|
||||
The dropdown is populated by reading `env_file:` directives from a representative target's compose. If no target declares any env file, the dropdown falls back to the default `.env` only. If your stack uses an inline `environment:` block in compose, the file does not exist and the dropdown can't list it; add an `env_file:` entry in compose and redeploy the stack, then reopen the wizard.
|
||||
</Accordion>
|
||||
<Accordion title="Values on the target look wrong after a push">
|
||||
Open the bundle's **Versions** tab and confirm which version is current. Then run **Preview** without pushing: the diff you see is exactly what the next push will write. If the bundle is correct but the running container still shows the old values, the container needs to be recreated to pick up the new env, since `env_file:` is read at create-time, not on a restart. Stop and redeploy the stack on the affected node, or use [Fleet Actions → Stop fleet by label](/features/fleet-actions) for a batch redeploy.
|
||||
</Accordion>
|
||||
<Accordion title="Two operators saved the bundle at the same time">
|
||||
Last-write-wins. Both saves succeed and both are visible in the **Versions** tab; only the later one is the current version. Open the bundle, reconcile the kv with whichever operator should "win" the merge, and save once more with a change note explaining the reconciliation.
|
||||
</Accordion>
|
||||
<Accordion title="Preview shows 'drift' rows; should I worry?">
|
||||
No. A `drift N` count means the target has N keys the bundle doesn't, which is informational only. The push will leave those keys alone. If a drift key is something you want gone, delete it from that node's env file in the [stack editor](/features/stack-file-explorer), then re-push the bundle to bring the rest of the keys back in line.
|
||||
</Accordion>
|
||||
<Accordion title="A node in the label selector resolved to zero matches">
|
||||
The labels combobox lists labels that exist somewhere in the fleet, but a label needs to be assigned to at least one online node for the selector to match it. Open **Settings → Nodes** on the control instance, confirm the label is assigned to the node you expect, and re-run **Preview**.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
The node's compose file for that stack does not declare the env file you asked Fleet Secrets to write to. Either pick a different env file in the wizard's **Env file** dropdown, or update that node's compose to reference the file via `env_file:`.
|
||||
## Where Fleet Secrets fits
|
||||
|
||||
### "stack not found"
|
||||
Fleet Secrets is one tab inside the Fleet view. Each adjacent surface answers a different question.
|
||||
|
||||
The named stack does not exist on that node. Fleet Secrets pushes to existing stacks only — it does not create stack directories.
|
||||
|
||||
### A push partially succeeded
|
||||
|
||||
Per-node failures are independent. The bundle and other targets are unaffected; only the failing node missed the push. Fix the underlying cause (file declared in compose, stack created, node reachable) and re-run the push. Idempotent overlay means re-pushing is safe.
|
||||
|
||||
### "A push for this secret is already running"
|
||||
|
||||
Only one push per bundle runs at a time. Wait for the previous push to finish, then retry.
|
||||
|
||||
### Values look wrong on the target
|
||||
|
||||
Open the bundle and switch to **Versions** — confirm the version you intended to push is the current one. Open the wizard, run **Preview** without pushing, and read the diff. The preview is a faithful representation of what the next push will write.
|
||||
| Related feature | What it covers | Why it is not Fleet Secrets |
|
||||
|---|---|---|
|
||||
| [Stack File Explorer](/features/stack-file-explorer) | The per-node editor for compose files, env files, and any other file in the stack directory | Stack File Explorer edits one file on one node; Fleet Secrets pushes one kv to many nodes |
|
||||
| [Fleet Actions](/features/fleet-actions) | Imperative bulk operations across labeled nodes: stop, prune, relabel | Fleet Actions runs operations; Fleet Secrets pushes configuration |
|
||||
| [Fleet Federation](/features/fleet-federation) | Operator-driven placement controls (cordon, pin) for [Blueprints](/features/blueprint-model) | Federation steers where new declarative deployments land; Fleet Secrets distributes existing-stack env updates |
|
||||
| [Fleet Sync](/features/fleet-sync) | Push-only replication of security policies from a control instance to its replicas | Fleet Sync replicates scan policies and CVE suppressions; Fleet Secrets replicates environment-variable bundles |
|
||||
| [Scheduled Operations](/features/scheduled-operations) | Recurring or one-shot scheduled stack operations | Scheduled Operations owns recurrence; Fleet Secrets pushes are operator-initiated in v1 |
|
||||
| [Audit Log](/features/audit-log) | Every mutating action on the control instance, including all Fleet Secrets actions | Fleet Secrets writes into the audit log; the Audit Log is where you read it back |
|
||||
| [Multi-Node Management](/features/multi-node) | Adding nodes to the fleet, the proxy model, and how tier propagates | Fleet Secrets rides on the same proxy; the bundles are authored on the control instance and pushed to remotes via the remote-node HTTP proxy |
|
||||
|
||||
@@ -1,44 +1,73 @@
|
||||
---
|
||||
title: "Fleet Sync"
|
||||
description: "How security rules replicate from a control Sencho instance to remote nodes."
|
||||
description: "Push-only replication of security rules from a control Sencho instance to its replicas."
|
||||
---
|
||||
|
||||
When you manage several Sencho instances as a fleet, the control instance (the one where you added remote nodes in **Settings → Nodes**) acts as the source of truth for security configuration. Rules you create on the control replicate automatically to every remote so the entire fleet enforces the same policies.
|
||||
When you manage several Sencho instances as a fleet, the control instance (the one whose **Settings → Nodes** holds the rest of the fleet as remotes) owns the authoritative copy of your security rules. Every time you create, edit, or delete one of those rules on the control, it streams to every reachable remote so the whole fleet enforces the same posture. The remotes are *replicas*: they hold a mirrored read-only copy until you demote them.
|
||||
|
||||
Today this covers **vulnerability scan policies** and **CVE suppressions**. Both replicate over the same channel and follow the same rules.
|
||||
Fleet Sync replicates three resources today, all over the same channel and with the same lifecycle:
|
||||
|
||||
## Control vs replica
|
||||
- **Scan policies** ([Vulnerability Scanning](/features/vulnerability-scanning)).
|
||||
- **CVE suppressions** ([CVE Suppressions](/features/cve-suppressions)).
|
||||
- **Misconfig acknowledgements**.
|
||||
|
||||
Every Sencho instance has a **role** that determines whether it accepts writes for replicated resources:
|
||||
|
||||
| Role | Behavior |
|
||||
|------|----------|
|
||||
| **Control** | Default for any standalone instance and for the node you point your browser at when managing a fleet. Accepts create, edit, and delete for security rules. Pushes changes to every remote. |
|
||||
| **Replica** | An instance that has received at least one sync push from a control. Shows rules read-only with a banner indicating they are managed upstream. Returns `403 Forbidden` for direct write attempts. |
|
||||
|
||||
Role detection is automatic: a Sencho instance becomes a replica the first time it accepts a sync push. To switch a replica back to a standalone control, an admin uses **Demote to control** in **Settings → Security** (described below).
|
||||
Fleet Sync lives under **Settings → Security** on both the control (where you author rules) and the replica (where you see them, read-only).
|
||||
|
||||
<Frame>
|
||||
<img src="/images/fleet-sync/replica-banner.png" alt="Settings page on a replica showing the 'Managed by control node' banner above the policies list" />
|
||||
<img src="/images/fleet-sync/control-security.png" alt="Settings → Security on the control. The Vulnerability Scanner card shows Trivy installed with Auto-update Trivy turned on. Below it, the empty 'No scan policies configured' state and the CVE Suppressions section listing two active suppressions against github.com/docker/docker, each annotated 'by admin · expires Never'." />
|
||||
</Frame>
|
||||
|
||||
## Control and replica roles
|
||||
|
||||
Every Sencho instance carries a `fleet_role` flag that is either `control` or `replica`. The flag is consulted on every write path for the replicated resources and by the **Settings → Security** UI when it decides whether to render the edit controls.
|
||||
|
||||
| Role | Behaviour |
|
||||
|---|---|
|
||||
| **Control** | The default for any fresh install and for the instance whose **Settings → Nodes** lists the rest of the fleet. Accepts create, edit, and delete on scan policies, CVE suppressions, and misconfig acknowledgements. Pushes the full current state of each resource to every reachable remote on every write. |
|
||||
| **Replica** | An instance that has received at least one Fleet Sync push. Renders replicated rules as read-only with a "Managed by control node" banner above the policy editor on **Settings → Security**. Returns `403 Forbidden` for any direct write attempt against the replicated tables. |
|
||||
|
||||
The transition from control to replica happens automatically the first time a replica accepts a push: the apply transaction sets `fleet_role = 'replica'` atomically with the row replacement, so the role flip and the new rows land together or not at all. Going the other way is explicit: an admin clicks **Demote to control** on the replica (see [Demote a replica](#demote-a-replica) below).
|
||||
|
||||
## What replicates
|
||||
|
||||
The three replicated resources share one wire protocol, one retry queue, and one anchor. Each push carries the full current state of one resource (not a delta); the receiver replaces every `replicated_from_control = 1` row in a single transaction.
|
||||
|
||||
What does *not* replicate:
|
||||
|
||||
- **Local rules created directly on a replica.** A replica can still hold its own local rules; sync only touches rows that originated on the control. The two coexist in the same table and are distinguished by an internal flag.
|
||||
- **Trivy itself.** The scanner binary is installed independently on each instance. The Security panel on a remote shows a "Scanner is per-node" callout in place of the full editor, since the scanner lifecycle is a node concern, not a fleet concern.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/fleet-sync/remote-security-via-proxy.png" alt="Settings → Security on a remote, opened via the control's node switcher. The Vulnerability Scanner card shows 'Not installed' with an Install Trivy button; below it, a 'Scanner is per-node' status callout reads 'Trivy is installed independently on each Sencho instance. Scan policies and CVE suppressions are managed on the control node.'" />
|
||||
</Frame>
|
||||
- **Everything outside the three resources above.** API tokens, audit logs, blueprints, secrets, alert rules, users, SSO config, and general settings stay per-instance.
|
||||
- **Pilot-agent nodes.** Sync over the [pilot tunnel](/features/pilot-agent) is not part of v1; the control logs a one-time warning per pilot node and skips it during fanout. The pilot node's local rules are unaffected.
|
||||
|
||||
## How replication works
|
||||
|
||||
1. You create, edit, or delete a security rule on the control.
|
||||
2. The control commits the change to its local database.
|
||||
3. The control iterates every remote node registered in **Settings → Nodes** that has an API URL and token configured.
|
||||
4. For each remote, the control posts the full current rule list to `/api/fleet/sync/<resource>`, tagged with the remote's own identity, a monotonic timestamp, and a fingerprint that identifies the control.
|
||||
5. The remote validates the payload, checks the control fingerprint and timestamp, and replaces its replicated rows in a single transaction.
|
||||
Fleet Sync is a push-only fan-out from the control on every relevant write. There is no pull, no polling, no schedule, no autonomous reconciler.
|
||||
|
||||
Local-only rules created directly on a remote coexist with the replicated set. Sync only replaces rows flagged as coming from the control.
|
||||
1. The control commits a change locally (the operator clicks **Save** in the UI; the row is written in a database transaction).
|
||||
2. The write handler invokes `pushResourceAsync(resource)`, which loads the current full state of that resource from the local database.
|
||||
3. The control iterates every remote node in **Settings → Nodes** that is a proxy-mode remote with an `api_url` and `api_token` configured, then posts the payload to `POST /api/fleet/sync/<resource>` on each one in parallel.
|
||||
4. Each remote validates the payload, checks the control fingerprint, checks the per-resource watermark, replaces its replicated rows, sets `fleet_role = 'replica'`, and records an audit-log entry. All five steps run inside one SQLite transaction.
|
||||
5. Failures are queued for retry (see [Push ordering and retry](#push-ordering-and-retry)).
|
||||
|
||||
## Control anchor
|
||||
Pushes to the same remote are serialised: a per-`(node, resource)` mutex guarantees that two payloads for the same node never overlap, so a slow earlier push cannot land on top of a faster later one. Different resources fan out independently and can travel in parallel.
|
||||
|
||||
A replica binds to the first control that pushes to it. The fingerprint is a hash derived from the control's persistent install ID, so a hostname change does not flag drift.
|
||||
When you add a new remote to **Settings → Nodes**, the control performs an immediate backfill push for every resource so the new replica catches up to the current state without waiting for the next write event.
|
||||
|
||||
After the first sync, pushes from any other control are rejected with `409 CONTROL_IDENTITY_MISMATCH`. This prevents an operator from accidentally pointing a second control at an existing replica and overwriting its mirrored state.
|
||||
The control identifies itself in every payload with a fingerprint derived from its persistent install ID (SHA-256 truncated to 16 hex characters). The fingerprint is stable across hostname or `api_url` changes; only a SQLite reset or an explicit re-anchor on the replica resets it.
|
||||
|
||||
To re-bind a replica to a different control (for example, after a control rebuild), an admin sends:
|
||||
## Anchor and re-anchor
|
||||
|
||||
A replica binds to the first non-empty fingerprint it sees. After that:
|
||||
|
||||
- A push carrying the same fingerprint is accepted.
|
||||
- A push carrying a different non-empty fingerprint is rejected with `HTTP 409` and the body code `CONTROL_IDENTITY_MISMATCH`. This prevents an operator from accidentally pointing a second control at an existing replica and overwriting its mirrored state.
|
||||
- A push with no fingerprint at all is accepted (this is the legacy code path; the receiver treats absence as "unknown control" and applies the rows without binding).
|
||||
|
||||
When you need to re-bind a replica to a different control (for example, the original control was rebuilt and now carries a fresh install ID), an admin on the replica calls the re-anchor endpoint with an explicit override flag:
|
||||
|
||||
```bash
|
||||
curl -X POST https://<replica-url>/api/fleet/role/reanchor \
|
||||
@@ -47,76 +76,104 @@ curl -X POST https://<replica-url>/api/fleet/role/reanchor \
|
||||
-d '{"override": true}'
|
||||
```
|
||||
|
||||
The next push from any control becomes the new anchor.
|
||||
Re-anchor clears the cached fingerprint, clears the per-resource watermarks, and drops every replicated row inside one transaction. The replica stays a replica (it remains a passive receiver) and the next push from any control becomes the new anchor.
|
||||
|
||||
## Push ordering
|
||||
## Push ordering and retry
|
||||
|
||||
Each push carries a strictly-increasing timestamp. The receiver rejects strictly-older pushes with `409 STALE_SYNC_PUSH` so a slow push that arrives after a faster one cannot overwrite the newer state. The control's normal write path generates a fresh timestamp on every change.
|
||||
Each push carries a `pushedAt` timestamp generated by the control. The control keeps a process-wide monotonic counter so two rapid writes never produce the same number, and so a clock skew that would otherwise rewind time still emits a strictly-increasing value.
|
||||
|
||||
Controls that predate this protocol (older Sencho versions in mixed-version fleets) send no timestamp, and the receiver treats those pushes as legacy and accepts them.
|
||||
On the receiver, the watermark is stored per resource. An incoming `pushedAt` that is strictly older than the current watermark for that resource is rejected with `HTTP 409 STALE_SYNC_PUSH`. Watermarks are independent: a stale push for scan policies does not block a newer push for CVE suppressions.
|
||||
|
||||
## Automatic retry
|
||||
If a push fails (network error, replica offline, `409 STALE_SYNC_PUSH`, `409 CONTROL_IDENTITY_MISMATCH`, or any other non-2xx response), the failure is recorded against `(node, resource)` in the `fleet_sync_status` table. The retry service then takes over:
|
||||
|
||||
If a remote is offline when a write happens, the control records the failure on the **Settings → Security** sync-status panel and retries every 5 minutes for the next 24 hours. Once the remote comes back online, the next retry catches it up to the latest state.
|
||||
- **Tick:** every 5 minutes.
|
||||
- **Window:** any failure within the last 24 hours is eligible for retry.
|
||||
- **Stale-target warning:** once a previously-working remote has been failing continuously for more than 1 hour, the control dispatches a `warning`-level notification. The warning fires at most once per hour-long failure window so it does not spam alert routing.
|
||||
- **Identity drift:** when a replica's view of its own URL (as seen by the control) changes between pushes, the replica dispatches a `warning` notification flagging the new identity. Identity-scoped rules are reapplied automatically on the next sync.
|
||||
|
||||
If a previously-working remote stays unreachable for more than an hour, the control dispatches a **warning notification** so the operator knows the node has fallen behind. The notification fires once per hour-long failure window.
|
||||
Stale pushes (`409 STALE_SYNC_PUSH`) are not counted as failures, since a newer push has already landed for the same resource on that node. The retry loop treats those targets as healthy.
|
||||
|
||||
## Policy scope across a fleet
|
||||
## Demote a replica
|
||||
|
||||
When you create a policy on the control, its scope is preserved during replication:
|
||||
An admin on a replica can take the instance back to a standalone control from **Settings → Security**. The button is "Demote to control" and sits inside the "Managed by control node" callout. The exact modal copy:
|
||||
|
||||
- **Fleet-wide** (no specific node selected): applies on every instance in the fleet.
|
||||
- **Specific node**: applies only on the instance whose identity matches. A policy targeting `Node A` evaluates only during scans that run on Node A, even after it has replicated to every other remote.
|
||||
> **Demote replica to control**
|
||||
>
|
||||
> Removes every replicated scan policy and CVE suppression mirrored from the control. Local edits to security policies on this instance become available again.
|
||||
|
||||
Only one policy is evaluated per deploy. The most specific match wins, in this order:
|
||||
Demote runs inside a transaction. It:
|
||||
|
||||
1. Node-scoped and stack-scoped policies.
|
||||
2. Node-scoped, stack-wildcard policies.
|
||||
3. Fleet-wide, stack-scoped policies.
|
||||
4. Fleet-wide, stack-wildcard policies.
|
||||
- Drops every row marked `replicated_from_control = 1` across all three resources.
|
||||
- Clears the cached control fingerprint and the three per-resource watermarks.
|
||||
- Flips `fleet_role` back to `control`.
|
||||
|
||||
When two rules tie on scope class, the lowest policy id wins so every replica resolves the same winner.
|
||||
Local rules authored on this instance directly (rows that did not come from the control) are untouched. From the control's side, the demoted instance is still in **Settings → Nodes**; the next write event triggers another push, which the now-standalone control rejects until the operator removes it from the original control's node list or accepts a new anchor.
|
||||
|
||||
On a replica, the **Settings → Security** panel shows local rules, fleet-wide replicated rules, and replicated rules that target this replica. Identity-scoped rules meant for a different replica do not appear.
|
||||
Demote requires `{"confirm": true}` in the request body to prevent a misclick from wiping mirrored state.
|
||||
|
||||
## Demote to control
|
||||
## Prerequisites
|
||||
|
||||
A replica admin can demote the instance back to a standalone control from **Settings → Security**. The button sits next to the "Managed by control node" banner. Demote:
|
||||
| Requirement | Why it matters |
|
||||
|---|---|
|
||||
| **A paid Sencho tier on the control instance** | Creating scan policies, CVE suppressions, and misconfig acknowledgements is a paid feature. Fleet Sync simply replicates rules that were created on the control, so a paid tier on the control is what enables the whole flow. Replicas accept pushes regardless of their own tier. |
|
||||
| **Admin user role on the control** | Authoring the rules that replicate, and operating the re-anchor and demote endpoints on a replica, are all admin-only actions. Operator and viewer roles can read rules but cannot create or remove them. |
|
||||
| **Proxy-mode remotes with `api_url` and `api_token` configured in Settings → Nodes** | Fleet Sync pushes over HTTPS to each remote's Sencho API using its long-lived bearer token. Remotes without an `api_url` or `api_token`, or remotes that connect over the pilot tunnel, are skipped. |
|
||||
| **Network reachability from the control to each remote** | Pushes are HTTP requests originating on the control. A remote that is firewalled off, behind NAT without a forwarded port, or otherwise unreachable will queue retries until it returns. |
|
||||
|
||||
- Drops every replicated scan policy and CVE suppression mirrored from the control.
|
||||
- Clears the cached fleet identity and control fingerprint.
|
||||
- Re-enables local edits on this instance.
|
||||
## Limitations
|
||||
|
||||
The control loses this remote as a replica. Re-adding the remote in the control's **Settings → Nodes** restarts replication from scratch on the next write.
|
||||
Fleet Sync v1 ships the three replicated resources and the control mechanics described above. The following are deliberately out of scope today:
|
||||
|
||||
## Pilot-agent nodes
|
||||
|
||||
Pilot-agent nodes are not part of fleet sync today. The control logs a one-time warning per pilot node and skips it during pushes. Pilot-tunnel-based replication is on the post-1.0 roadmap.
|
||||
- **No first-party sync-status panel in the UI.** Sync activity is visible in two places: the receiving replica's audit log (every applied push records `POST /api/fleet/sync/<resource>` with username `system` and the source fingerprint in the summary) and the control's server logs. A built-in status panel is on the post-1.0 roadmap.
|
||||
- **Stack-pattern scope only in the form UI.** The scan policy and CVE suppression forms expose a stack-pattern glob (e.g. `prod-*`) but no node selector. The underlying data model supports node-scoped rules, and they replicate faithfully when present, but authoring them today requires writing directly to the control's API.
|
||||
- **5,000 rows per resource per push.** If the control accumulates more than 5,000 active rules in any one resource, the sender truncates to the first 5,000 and emits a `warning` notification (cool-down: 6 hours). In practice no fleet should approach this cap; if you see the warning, treat it as a signal to consolidate rules.
|
||||
- **Pilot-agent nodes are not part of Fleet Sync today.** The control logs a one-time warning per pilot node and skips it during fanout. Replicating security rules over the [pilot tunnel](/features/pilot-agent) is on the post-1.0 roadmap.
|
||||
- **One control per replica at a time.** The anchor binds on the first non-empty fingerprint received and stays bound until re-anchor. Two controls pushing to the same replica is not a supported topology; the second control will receive `409 CONTROL_IDENTITY_MISMATCH` for every push.
|
||||
- **Pushes carry the full current state, not deltas.** Each push replaces every replicated row for that resource on the receiver. The cost is bounded by the row cap above; the benefit is that the replica always converges to the control's current state without needing to reconcile a long edit history.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### Security rules are missing on this node
|
||||
<AccordionGroup>
|
||||
<Accordion title="Rules I created on the control are missing on a remote">
|
||||
The remote is either offline, was offline at the moment of the last write and has not yet been retried, or its bearer token in **Settings → Nodes** is invalid. Open **Settings → Nodes** on the control and click **Test connection** for the remote in question. If the connection check passes, wait up to 5 minutes for the next retry tick; the next successful push will reconcile the remote to the current state. If the connection check fails, fix the underlying reachability or token issue and Fleet Sync will catch up on the next retry.
|
||||
</Accordion>
|
||||
<Accordion title="The Security panel on a remote is empty even though the control has rules">
|
||||
This usually means the remote has not yet received its first push (it is still a fresh control from its own perspective). Any write on the control triggers a fresh push, or you can add the remote again in **Settings → Nodes** to trigger the add-node backfill. Once the first push lands, the remote flips to replica and shows the "Managed by control node" callout above the policy list.
|
||||
</Accordion>
|
||||
<Accordion title="A push returns 409 CONTROL_IDENTITY_MISMATCH">
|
||||
The replica is already anchored to a different control's fingerprint. Either point your browser at the original control and continue authoring there, or sign in as admin on the replica and call the re-anchor endpoint with `{"override": true}` to clear the anchor. The next push from any control then becomes the new anchor.
|
||||
</Accordion>
|
||||
<Accordion title="A push returns 409 STALE_SYNC_PUSH">
|
||||
A newer push for the same resource has already landed on this replica, so the older retry is silently dropped. No action needed; the next write on the control will produce a fresher timestamp and succeed.
|
||||
</Accordion>
|
||||
<Accordion title="I deleted a rule on the control but it is still visible on a replica">
|
||||
The push for the delete probably failed. Check **Settings → Nodes** on the control; a remote that is offline or whose token is invalid will not have received the delete. Once the underlying issue is resolved, the next retry catches the remote up to the current state (which no longer includes the deleted rule).
|
||||
</Accordion>
|
||||
<Accordion title="I cannot see whether sync succeeded">
|
||||
There is no first-party sync-status panel in the UI today (see [Limitations](#limitations)). Two surfaces give you the information indirectly: the replica's **Audit** view records every applied push as a `POST /api/fleet/sync/<resource>` entry with username `system` and the source fingerprint in the summary, and the control's server logs record the push outcome per remote. The administrative API endpoint `GET /api/fleet/sync-status` returns the per-resource per-node status table for tooling that needs to consume it directly.
|
||||
</Accordion>
|
||||
<Accordion title="A warning says my suppression list was truncated">
|
||||
The control side of Fleet Sync caps each push at 5,000 rows per resource (see [Limitations](#limitations)). If your suppression list, scan policy list, or acknowledgement list exceeds the cap, the sender truncates to the first 5,000 and emits a single warning every 6 hours. Consolidate rules by replacing repetitive entries with pattern-scoped ones, or remove suppressions that are no longer needed.
|
||||
</Accordion>
|
||||
<Accordion title="Misconfig acknowledgements did not replicate">
|
||||
Misconfig acknowledgements ride the same channel as scan policies and CVE suppressions; if one resource replicates and another does not, it is almost always a per-resource watermark race. Re-save the acknowledgement on the control to produce a fresh `pushedAt`, then the next push catches the remote up. If the symptom persists, check the control's server logs for the `[FleetSync]` push outcome on that remote.
|
||||
</Accordion>
|
||||
<Accordion title="I want this replica to be a standalone control again">
|
||||
Open **Settings → Security** on the replica and click **Demote to control** in the "Managed by control node" callout. The action requires confirmation and drops every replicated row from this instance. Local rules authored directly on this instance are kept. The control loses this remote as a replica; remove it from the control's **Settings → Nodes** as well if you no longer want the control to push to it.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
You are looking at a replica. Navigate to the control Sencho instance and view its **Settings → Security** page; rules you create there replicate here automatically.
|
||||
## Where Fleet Sync fits
|
||||
|
||||
If rules are expected to be present but are not, check **Settings → Security** on the control: it surfaces sync status per node and shows the most recent error. The retry loop catches up automatically once the underlying issue (network, expired token) is resolved.
|
||||
Fleet Sync is one of several features that span multiple nodes in a Sencho fleet. Each handles a different concern; they do not overlap.
|
||||
|
||||
### A remote rejects the sync
|
||||
|
||||
A remote rejects a push if its bearer token is invalid or the request does not reach it. Verify the node's API URL and token in **Settings → Nodes** on the control, then use **Test Connection** to confirm the remote is reachable.
|
||||
|
||||
### A push returns 409 CONTROL_IDENTITY_MISMATCH
|
||||
|
||||
The replica is anchored to a different control. Either point your browser at the original control, or call the reanchor endpoint on the replica with `{"override": true}` to re-bind.
|
||||
|
||||
### A push returns 409 STALE_SYNC_PUSH
|
||||
|
||||
A newer push has already landed on the replica for the same resource. The control's retry loop will skip the stale outcome on the next tick. No action needed.
|
||||
|
||||
### A rule I deleted on the control is still visible on a replica
|
||||
|
||||
Sync happens on every write. If you deleted a rule but the replica still shows it, the push probably failed. Check the sync-status panel on the control or wait for the next retry tick (within 5 minutes).
|
||||
|
||||
### I want this replica to be a standalone control
|
||||
|
||||
Use **Demote to control** in **Settings → Security** on the replica. The button is admin-only and asks for explicit confirmation because demote wipes mirrored rules.
|
||||
| Related feature | What it covers | Why it is not Fleet Sync |
|
||||
|---|---|---|
|
||||
| [Fleet View](/features/fleet-view) | The masthead, tab strip, and node-grid that show fleet status at a glance. | Fleet View is observability; Fleet Sync is state replication. |
|
||||
| [Multi-Node Management](/features/multi-node) | Adding remotes to **Settings → Nodes**, proxy versus pilot mode, license inheritance across the fleet. | Multi-Node Management is how nodes become reachable; Fleet Sync uses the same plumbing to ship security rules. |
|
||||
| [Pilot Agent](/features/pilot-agent) | The outbound WebSocket tunnel mode used by remotes behind NAT. | Fleet Sync does not flow over the pilot tunnel today. |
|
||||
| [Vulnerability Scanning](/features/vulnerability-scanning) | Trivy installation, scan execution, and posture grading on a single node. | Scanning runs per node; Fleet Sync replicates the policies and acknowledgements that shape what scans surface across the fleet. |
|
||||
| [CVE Suppressions](/features/cve-suppressions) | Marking specific CVEs as known-benign so they stop triggering alerts. | Suppressions are one of the three resources Fleet Sync replicates; this page documents the replication mechanics. |
|
||||
| [Fleet Federation](/features/fleet-federation) | Operator-driven placement controls for blueprints (cordon and pin). | Federation steers blueprint placement; Fleet Sync replicates security state. The two do not interact. |
|
||||
| [Fleet Actions](/features/fleet-actions) | Bulk imperative operations (fleet stop, bulk label) across labelled nodes. | Fleet Actions is imperative and operator-driven; Fleet Sync is declarative and automatic. |
|
||||
| [Licensing](/features/licensing) | Tier matrix and what each tier unlocks. | Single source of truth for the paid-tier requirement on the control. |
|
||||
|
||||
@@ -1,44 +1,138 @@
|
||||
---
|
||||
title: Fleet View
|
||||
description: Monitor all your nodes from a single dashboard with real-time health metrics, search, filtering, and container drill-down.
|
||||
description: Monitor every node in your Sencho deployment from one screen, drill into stacks and containers, and orchestrate Sencho version updates across the fleet.
|
||||
---
|
||||
|
||||
The **Fleet** tab gives you a bird's-eye view of every node in your Sencho deployment, local and remote, on one screen. It is available to all tiers, with advanced features unlocked by Skipper and Admiral.
|
||||
The **Fleet** tab is the single page you open when you want to see the whole estate at once: every node's health, every stack, every container, and which boxes need a Sencho update. It is the home for fleet-wide tabs that go beyond monitoring (Snapshots, Status, Deployments, Routing, Federation, Fleet Actions, Secrets), each linking out to its own dedicated page.
|
||||
|
||||
<Note>
|
||||
Fleet is hub-only and is hidden from the nav strip when a remote node is the active selection. See [Multi-Node Management](/features/multi-node#what-top-level-views-show-when-a-remote-node-is-active).
|
||||
</Note>
|
||||
|
||||
<Frame>
|
||||
<img src="/images/fleet-view/fleet-overview.png" alt="Fleet Overview with the fleet masthead, grid/topology toggle, and pinned local node" />
|
||||
<img src="/images/fleet-view/fleet-overview.png" alt="Fleet page showing the masthead with the editorial 'The fleet' word, a meta line reading '8 nodes · 8 online · last sync 12s', a '1 critical' reasons line, and CPU / MEM / CONTAINERS stat tiles next to a 1-alert counter. Below, the eight-tab strip starts at Overview and ends at Secrets. Under that, the toolbar (search, sort, filters, Grid/Topology toggle) sits above a grid of node cards: a pinned Local card with a cyan ★ Local rail and an Online badge plus version chip, then Opsix, Pitt-Moba, sencho-pilot-test (with an Update available badge and Update to v0.76.3 button), and four more remote cards." />
|
||||
</Frame>
|
||||
|
||||
## Page layout
|
||||
|
||||
The page has three rails stacked top to bottom: a status masthead, a tab strip plus action buttons, and the active tab's content.
|
||||
|
||||
### Fleet masthead
|
||||
|
||||
The top of the page is a status masthead that summarises the state of the entire fleet at a glance:
|
||||
A single rail summarises the state of every registered node so you can read the whole estate without scrolling.
|
||||
|
||||
- **Status word** in italic display type: **The fleet** with a pulsing coloured rail (green when all nodes are online and healthy, amber when any node is offline, rose when any node is critical)
|
||||
- **Meta line** showing node count, online count, and last sync time (e.g. "2 nodes · 1 online · last sync 7s")
|
||||
- **Reasons** line appears when the fleet is degraded, summarising what's wrong ("1 offline · 1 critical")
|
||||
- **Stats cluster** with three tiles:
|
||||
- **CPU**: average across online nodes, with the peak node and percentage called out below
|
||||
- **MEM**: total RAM used across the fleet, with total capacity and percentage
|
||||
- **CONTAINERS**: active running count; hover the tile to see a breakdown of running vs total
|
||||
- **Alerts** indicator on the right shows the current critical count, highlighted in rose when above zero
|
||||
- A **state word** that always reads `The fleet`, set in the editorial italic display face. The word changes color with the fleet's overall state: foreground when **Healthy** (every node online, none critical), warning when **Degraded** (any node offline), destructive when **Critical** (any online node above the CPU or disk threshold).
|
||||
- A **pulsing dot** in the same colour, with a soft halo. The dot is solid when Healthy and pulses when Degraded or Critical, so motion only appears when something needs attention.
|
||||
- A **rail tint** down the left edge of the masthead that mirrors the state colour, plus a subtle gradient wash across the bar.
|
||||
- A **meta line** in uppercase mono tracking that reads `<n> nodes · <m> online · last sync Xs`, where the sync delta uses compact units (`s` / `m` / `h`). While data is loading the line reads `syncing…` instead.
|
||||
- A **reasons line** below the meta line that names exactly what is wrong, e.g. `1 offline · 1 critical`. The reasons line is hidden when the fleet is Healthy.
|
||||
- Three **stat tiles** on the right edge, in mono tabular numerals:
|
||||
|
||||
### Grid / Topology toggle
|
||||
| Tile | What it shows |
|
||||
|------|---------------|
|
||||
| **CPU** | Average CPU across online nodes, with a sub-line that names the peak node (`peak <name> <percent>%`). The value tints amber once average CPU is at or above 80%. |
|
||||
| **MEM** | Total RAM used across online nodes, with `of <total> · <percent>%` underneath. |
|
||||
| **CONTAINERS** | Active running container count across online nodes, with `of <total> total` underneath. Hover the tile to surface a cursor-following tooltip with the running / total split. |
|
||||
|
||||
Below the masthead, switch between two layouts:
|
||||
- An **alerts indicator** pinned to the far right: a bell icon next to the current critical-node count, with an `alert` / `alerts` mono label. The icon and number tint destructive while the count is above zero.
|
||||
|
||||
- **Grid** (default): one card per node. The local node is pinned at the top with a cyan accent rail and a ★ Local badge so it is never confused with a remote.
|
||||
- **Topology**: an interactive node-map view. Each node renders as a rack card with a status pill (Online, Critical, or Offline), a Local or Remote badge, CPU, memory, and disk bars, and a stack and running-container summary. Online remotes also show their gateway round-trip latency; cordoned nodes carry an amber **Cordoned** banner; pilot-agent nodes whose heartbeat has gone stale flag a small clock glyph. The local node is framed with a cyan ring. Connector lines colour by link health (cyan when online, amber when critical, dashed when offline). Pan, zoom, and use the minimap to navigate large fleets. Click a node to jump to its details.
|
||||
### Tabs
|
||||
|
||||
The Fleet view is a tab strip. Four tab triggers are visible to every tier; the Deployments, Routing, Federation, and Secrets triggers only render when the active license unlocks them. A vertical separator after **Status** divides the per-node monitoring tabs from the fleet-wide orchestration tabs.
|
||||
|
||||
| Tab | Tier | What it does |
|
||||
|-----|------|--------------|
|
||||
| **Overview** | Community | The grid or topology view of every node and its health. Covered in the next section. |
|
||||
| **Snapshots** | Community (manual) / Skipper (scheduled) | Snapshot every compose file across the fleet. See [Fleet-Wide Backups](/features/fleet-backups). |
|
||||
| **Status** | Community | One card per node summarising which automations and security features are configured. Covered below. |
|
||||
| **Deployments** | Skipper | Blueprint deployments and reconciler state. See [Blueprints](/features/blueprint-model). |
|
||||
| **Routing** | Admiral | Cross-node service routing via Sencho Mesh. See [Sencho Mesh](/features/sencho-mesh). |
|
||||
| **Federation** | Admiral | Cordon nodes and pin blueprints to specific hosts. See [Fleet Federation](/features/fleet-federation). |
|
||||
| **Fleet Actions** | Community trigger / Skipper content | The tab is visible to every tier; opening it on Community surfaces a Skipper upgrade prompt. See [Fleet Actions](/features/fleet-actions). |
|
||||
| **Secrets** | Skipper | Encrypted env-var bundles you push to labeled nodes. See [Fleet Secrets](/features/fleet-secrets). |
|
||||
|
||||
### Action buttons
|
||||
|
||||
Three buttons sit next to the tab strip in the top-right corner. They are visible from every tab, not just Overview.
|
||||
|
||||
| Button | What it does |
|
||||
|--------|--------------|
|
||||
| **Check Updates** | Opens the [Node Updates sheet](#node-updates) where you can manage Sencho version updates across the fleet. |
|
||||
| **Refresh** | Forces an immediate re-fetch of fleet data. The icon spins while the refresh is in flight; the button is disabled until it completes. |
|
||||
| **Add node** | Opens the same Add Node dialog as **Settings · System · Nodes** so you can register a new node without leaving the Fleet page. Admin-only. After saving a remote proxy node, a connection test runs automatically and the result toasts in (success or saved-but-unreachable). |
|
||||
|
||||
## The Overview tab
|
||||
|
||||
Overview is the default tab and is where most operators spend their time. It offers two layouts of the same data: a card grid and an interactive topology graph.
|
||||
|
||||
### Toolbar
|
||||
|
||||
In **Grid** mode the toolbar exposes search, sort, filters, and the view toggle. In **Topology** mode only the view toggle remains (the grid-only controls collapse).
|
||||
|
||||
| Control | Behaviour |
|
||||
|---------|-----------|
|
||||
| **Search** | Real-time filter that matches against node names and the names of stacks deployed on each node. Typing `plex` keeps only nodes that have a `plex` stack; typing `dev` keeps only nodes whose name contains `dev`. |
|
||||
| **Sort** | Combobox with five orderings: Name, CPU Usage, Memory Usage, Containers, Status. The selection persists in the browser. |
|
||||
| **Direction toggle** | Arrow button next to the sort combobox. The icon rotates 180° and the title flips between *Switch to ascending* / *Switch to descending* to confirm the current direction. |
|
||||
| **Filters** | Popover with four sections: **Status** (All / Online / Offline), **Type** (All / Local / Remote), **Severity** (Critical Only toggle), and **Tags** (multi-select of fleet label palette, only visible if labels exist). The button shows a count badge for active filters; a **Clear all filters** action appears at the bottom of the popover when at least one filter is set. |
|
||||
| **Grid / Topology** | Segmented control on the right, with a Grid icon and a Topology icon. The control resets to Grid when the page is reloaded. |
|
||||
|
||||
### Grid view
|
||||
|
||||
Every node renders as a card. The local node is pinned at the top of the grid with a cyan accent rail, a brand-tinted gradient wash, and a `★ Local` mono label in the top-right so it can never be confused with a remote.
|
||||
|
||||
| Element | Description |
|
||||
|---------|-------------|
|
||||
| **Server icon tile** | Green when online, muted when offline. Sits next to the node name. |
|
||||
| **Online / Offline badge** | Online (success palette, Wi-Fi icon) or Offline (secondary palette, Wi-Fi-off icon). |
|
||||
| **Type badge** | Outline pill reading `local` or `remote`. |
|
||||
| **Version badge** | The node's Sencho version in mono tabular numerals (e.g. `v0.76.3`). Hidden if the node cannot report a version. |
|
||||
| **Update available** badge | Warning pill shown when a newer Sencho release is published for this node. |
|
||||
| **Critical** badge | Destructive pill with a triangle icon, surfaced when the online node is above 90% CPU or 90% disk. |
|
||||
| **Cordoned** badge | Warning pill with a Ban icon, surfaced when an Admiral has cordoned the node. The badge tooltip carries the cordon reason or the default *Unschedulable: new blueprint deployments skip this node*. See [Fleet Federation](/features/fleet-federation) for the full cordon and pin flow. |
|
||||
| **Updating / Updated / Failed** badge | Update progress indicator, shown only while or just after an update flows through. Failed states surface inline retry and dismiss buttons and a cursor-following error tooltip. |
|
||||
| **Container stats grid** | Three cells: **Running** (active containers), **Stopped** (exited containers), **Stacks** (count, or `-` if the node has not reported). Hidden on offline nodes. |
|
||||
| **CPU / RAM / Disk bars** | Each row shows the metric icon, the percent (CPU) or `used / total` (RAM, Disk), and a horizontal bar that tints amber at 60%, destructive at 80% (CPU/RAM), or amber at 75% / destructive at 90% (Disk). Hidden on offline nodes. |
|
||||
| **Update to v…** button | Outline button that runs along the bottom of the card when an update is available. The label includes the latest version. |
|
||||
| **Stack details** trigger | Footer button that toggles the stack drill-down. The label carries the stack count for the node. |
|
||||
|
||||
Offline nodes render dimmed, with no stats grid, no usage bars, and no update affordance.
|
||||
|
||||
### Node actions menu (admin)
|
||||
|
||||
Every card carries a three-dot **Node actions** kebab in the top-right corner. The menu surfaces the same lifecycle actions you would find in **Settings · System · Nodes**:
|
||||
|
||||
| Action | Notes |
|
||||
|--------|-------|
|
||||
| **Edit node** | Opens the Edit dialog prefilled with the node's connection details. For proxy-mode remotes, saving with a changed API URL or token re-runs the connection test automatically. |
|
||||
| **Delete node** | Opens a destructive confirmation. The local (default) node has no Delete option. Deleting a remote only removes it from this console; the remote instance and its containers are untouched. |
|
||||
| **Cordon node** / **Uncordon node** | Marks the node unschedulable so new blueprint deployments skip it (Admiral). Existing deployments keep running. |
|
||||
|
||||
The menu is admin-only. Non-admin users see no kebab on the card.
|
||||
|
||||
### Topology view
|
||||
|
||||
Switch the segmented control to **Topology** to swap the grid for a hub-and-spoke graph. The local node sits in the centre with a cyan ring and the remotes radiate outward; the layout is computed automatically from the registered fleet.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/fleet-view/fleet-topology.png" alt="Fleet Topology view showing the local node connected to a remote, each rendered as a rack card with CPU, memory, and disk bars" />
|
||||
<img src="/images/fleet-view/fleet-topology.png" alt="Fleet Topology view with the Local node card boxed in cyan in the centre showing 'CRITICAL · LOCAL' status pills, 16 cores CPU bar, RAM and DISK bars in destructive red, and '15 stacks · 17 running' below. Six remote node cards radiate to the right (SLX-Mars, sencho-test-03, sencho-test-02, sencho-test-01, sencho-pilot-test, Pitt-Moba, Opsix) each showing their CPU/MEM/DISK bars and stack summary. Cyan connector lines run from the local node to each remote. ReactFlow zoom controls sit at the bottom-left and the navigator minimap is pinned to the bottom-right." />
|
||||
</Frame>
|
||||
|
||||
Each node card in the topology graph carries:
|
||||
|
||||
- A **status pill** at the top reading `Online`, `Critical`, or `Offline`, with a coloured dot.
|
||||
- A **Local** or **Remote** label in the top-right.
|
||||
- A **lightning icon** when the node is critical; the card frame mutes when the node is offline.
|
||||
- The node name with a server icon.
|
||||
- Three compact **CPU / MEM / DISK** bars with the percent value at the right of each row.
|
||||
- A footer line summarising stacks and running containers, e.g. `3 stacks · 4 running`.
|
||||
|
||||
**Connector lines** colour by the link's health: cyan when both ends are online, warning when the remote is critical, dashed and muted when the remote is offline.
|
||||
|
||||
The graph is interactive: drag the canvas to pan, scroll to zoom, drag a node to reposition it, and use the **navigator minimap** in the bottom-right corner to jump around large fleets. The ReactFlow controls in the bottom-left expose explicit zoom in / zoom out / fit-to-view buttons. Click any node card to navigate into its node detail.
|
||||
|
||||
The graph re-lays out only when nodes are added, removed, or change type. Live metric updates on existing nodes do not move the layout, so an operator who has dragged nodes into a custom arrangement keeps it.
|
||||
|
||||
#### Topology layout modes (Skipper+)
|
||||
|
||||
A toolbar at the top of the topology canvas offers three layouts. Pick the one that matches how you reason about your fleet.
|
||||
@@ -49,255 +143,142 @@ A toolbar at the top of the topology canvas offers three layouts. Pick the one t
|
||||
| **Grouped** | Remotes cluster by their primary node label (alphabetically first when a node has multiple). The local node sits in its own cluster. Use this to read environment, region, or role groupings at a glance. Nodes without any label fall into an Unlabeled cluster. |
|
||||
| **Free** | Drag any node anywhere on the canvas. Positions persist in this browser, so the arrangement is there when you come back. |
|
||||
|
||||
Assign labels in **Settings → Nodes** to drive Grouped mode. Each node accepts multiple labels; the alphabetically first one is its cluster.
|
||||
|
||||
### Tabs
|
||||
|
||||
The Fleet page has three tabs:
|
||||
|
||||
- **Overview**: the monitoring view described on this page
|
||||
- **Snapshots**: fleet-wide backup snapshots (covered in [Fleet Backups](/features/fleet-backups))
|
||||
- **Status**: fleet-wide feature status rollup (covered below)
|
||||
|
||||
### Action buttons
|
||||
|
||||
Three buttons sit in the top-right corner:
|
||||
|
||||
| Button | What it does |
|
||||
|--------|--------------|
|
||||
| **Check Updates** | Opens the Node Updates modal to view and apply Sencho version updates across your fleet (Skipper+) |
|
||||
| **Refresh** | Re-fetches data from all nodes. Shows a spinner while loading |
|
||||
| **Add node** | Opens the same Add Node dialog as Settings → Nodes so you can register a new node without leaving the Fleet page. Admin-only. After saving a remote proxy node, a connection test runs automatically and the result toasts in (success or saved-but-unreachable). |
|
||||
|
||||
## Community features
|
||||
|
||||
Every Sencho installation gets the full fleet monitoring grid at no cost.
|
||||
|
||||
### Node grid
|
||||
|
||||
Each node appears as a card showing:
|
||||
|
||||
| Data | Description |
|
||||
|------|-------------|
|
||||
| **Status badge** | Online (green) or Offline (grayed out) |
|
||||
| **Type badge** | `local` or `remote` |
|
||||
| **Version badge** | The node's Sencho version (e.g. `v0.38.0`). Remote nodes running older versions that cannot report their version show no version badge |
|
||||
| **Update available badge** | Orange pill shown when a newer Sencho version is available for this node |
|
||||
| **Critical badge** | Red pill shown when CPU or disk usage exceeds 90% |
|
||||
| **Running containers** | Count of containers in `running` state |
|
||||
| **Stopped containers** | Count of containers in `exited` state |
|
||||
| **Stacks** | Total number of Compose stacks on the node |
|
||||
| **CPU usage** | Current percentage with colour-coded bar (green to amber to red) |
|
||||
| **RAM usage** | Used / total with percentage bar |
|
||||
| **Disk usage** | Used / total with percentage bar |
|
||||
|
||||
Nodes with an available Sencho update also show an **Update to vX.Y.Z** button at the bottom of the card.
|
||||
|
||||
Offline nodes are visually dimmed and show a "Node unreachable" placeholder instead of stats.
|
||||
|
||||
### Node actions menu (admin)
|
||||
|
||||
The kebab menu in the top-right of each card surfaces the same lifecycle actions you'd find in Settings → Nodes:
|
||||
|
||||
| Action | Notes |
|
||||
|--------|-------|
|
||||
| **Edit node** | Opens the Edit dialog prefilled with the node's connection details. For proxy-mode remotes, saving with a changed API URL or token re-runs the connection test automatically. |
|
||||
| **Delete node** | Opens a destructive confirmation. The local (default) node has no Delete option. Deleting a remote only removes it from this console; the remote instance and its containers are untouched. |
|
||||
| **Cordon node** / **Uncordon node** | Marks the node unschedulable so new blueprint deployments skip it (Admiral). Existing deployments keep running. |
|
||||
|
||||
The menu is admin-only. Non-admin users see no kebab on the card.
|
||||
|
||||
### Manual refresh
|
||||
|
||||
Click the **Refresh** button in the top-right to re-fetch data from all nodes. The button shows a spinner while loading.
|
||||
|
||||
---
|
||||
|
||||
## Fleet operations
|
||||
|
||||
<Note>
|
||||
The toolbar (search, sort, filters), stack drill-down, auto-refresh, and per-node update flow are available on every tier. The bulk **Update All** action inside the Node Updates modal is a Skipper or Admiral feature.
|
||||
</Note>
|
||||
|
||||
### Auto-refresh
|
||||
|
||||
Fleet data automatically refreshes every 30 seconds. A subtle indicator at the bottom of the page confirms this is active. When a node update is in progress, the refresh rate increases to every 5 seconds so you can watch status changes in near real-time.
|
||||
|
||||
### Search
|
||||
|
||||
The search bar filters the node grid in real time. It matches against:
|
||||
- Node names (e.g. typing `dev` shows only nodes with "dev" in the name)
|
||||
- Stack names (e.g. typing `plex` shows only nodes that have a "plex" stack)
|
||||
|
||||
### Sorting
|
||||
|
||||
Use the sort dropdown to order nodes by:
|
||||
- **Name** (alphabetical)
|
||||
- **CPU Usage** (highest first)
|
||||
- **Memory Usage** (highest first)
|
||||
- **Containers** (most first)
|
||||
- **Status** (online first)
|
||||
|
||||
Click the arrow button next to the dropdown to toggle ascending/descending.
|
||||
|
||||
Sort preferences are saved to your browser and persist across sessions.
|
||||
|
||||
### Filtering
|
||||
|
||||
Filter pills let you narrow the grid:
|
||||
|
||||
| Filter | Options |
|
||||
|--------|---------|
|
||||
| **Status** | All, Online, Offline |
|
||||
| **Type** | All Types, Local, Remote |
|
||||
| **Critical Only** | Show only nodes with CPU or disk above 90% |
|
||||
| **Tags** | Filter by stack labels assigned to nodes |
|
||||
|
||||
A "Clear filters" button appears when filters hide all nodes.
|
||||
Assign labels in **Settings · System · Nodes** to drive Grouped mode. Each node accepts multiple labels; the alphabetically first one is its cluster.
|
||||
|
||||
### Stack drill-down
|
||||
|
||||
Click **Stack details** on any online node card to expand the stack list. The button shows the total stack count (e.g. "18 stacks"). Each stack in the list shows a container count badge.
|
||||
|
||||
Click a stack name to expand it further and see individual containers with:
|
||||
- Container name
|
||||
- State badge (running, exited, restarting)
|
||||
- Image name (e.g. `linuxserver/plex:latest`)
|
||||
- Uptime (e.g. "Up 4 days")
|
||||
|
||||
Container data is fetched fresh each time you expand a stack, so you always see the current state.
|
||||
|
||||
Hover over any container row to reveal an **Open in editor** button that navigates you directly to that stack's editor on the corresponding node.
|
||||
Click **Stack details** on any online node card to expand the stack list. The right side of the row carries a `<n> stacks` counter once the list has loaded.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/fleet-view/fleet-drill-down.png" alt="Fleet View with Stack details expanded showing the stack list and a container drill-down" />
|
||||
<img src="/images/fleet-view/fleet-drill-down.png" alt="Fleet grid with the Opsix node card expanded. The node header shows 4 running, 0 stopped, 3 stacks, and CPU / RAM / Disk usage bars. Below, the Stack details list shows three stacks (saelix-app-ca, saelix-app-wa, saelix-db). saelix-app-ca and saelix-db are expanded, each showing one container row with a green running dot, the container name, a 'running' state badge, the image tag, and an 'Up 2 weeks' status string." />
|
||||
</Frame>
|
||||
|
||||
### Critical node detection
|
||||
Each stack row carries the stack name, an expand chevron, a `running / total` container count, and any **fleet label dots** colour-coded from the node's label palette (when labels are configured for the stack).
|
||||
|
||||
Nodes with CPU or disk usage above 90% automatically receive a red **Critical** badge. Combined with the **Critical Only** filter, this lets you quickly triage overloaded servers.
|
||||
Click a stack row to expand the per-container drill-down. Each container shows:
|
||||
|
||||
### Node Updates
|
||||
- A **state dot** that is green for running, warning for restarting, destructive for exited.
|
||||
- The **container name** and a **state badge** (`running`, `restarting`, or `exited`).
|
||||
- The **image tag** if known (e.g. `linuxserver/plex:latest`).
|
||||
- A **status string** carried by Docker (e.g. `Up 4 days`, `Restarting (1) 3s ago`).
|
||||
- An **external-link button** that appears on hover and navigates straight to that stack's editor on the corresponding node.
|
||||
|
||||
Click **Check Updates** in the header to open the Node Updates modal. This lets you manage Sencho version updates across your entire fleet from one place.
|
||||
Container data is fetched fresh each time you expand a stack, so you always see the current state. Stack data for a node is fetched once on first open and cached for the session; re-expanding a stack does not refetch the stack list.
|
||||
|
||||
### Auto-refresh cadence
|
||||
|
||||
The Overview data refreshes on its own:
|
||||
|
||||
| Loop | Cadence |
|
||||
|------|---------|
|
||||
| Per-node health, stats, stack counts | every 30 seconds |
|
||||
| Sencho update-status check | every 2 minutes |
|
||||
| Fast poll while any node is actively updating | every 5 seconds |
|
||||
|
||||
A subtle `Auto-refreshing every 30 seconds` line at the bottom of the page confirms the loop is alive. The **Refresh** button in the top-right forces an immediate poll without waiting.
|
||||
|
||||
## The Status tab
|
||||
|
||||
The **Status** tab gives you a fleet-wide rollup of which automations and security features are configured on every node, without having to open each node's settings individually.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/fleet-view/fleet-node-updates.png" alt="Node Updates modal showing update status for all fleet nodes" />
|
||||
<img src="/images/fleet-view/fleet-status-tab.png" alt="Fleet Status tab with eight node cards in a three-column grid. The Local card is in the top-left and shows the full eight-row summary (Agents None, Alert rules None, Auto-heal None, Webhooks None, MFA Off, Scanning None, Backup Enabled, Crash detect On). Opsix, Pitt-Moba, SLX-Mars, sencho-test-01/02/03 each show a six-row summary that omits MFA and Backup. The sencho-pilot-test card is offline and shows the muted message 'Node is unreachable. Configuration unavailable.'" />
|
||||
</Frame>
|
||||
|
||||
The modal shows:
|
||||
Online nodes render a two-column summary grid with up to eight rows:
|
||||
|
||||
- **Summary cards** at the top: counts of nodes that are Up to date, have updates Available, are currently Updating, or have Failed
|
||||
- **Latest version** label showing the newest available Sencho release (checked via GitHub Releases, cached for 30 minutes)
|
||||
- **Filter** search box to find specific nodes
|
||||
- **Node table** with columns: Node name, Type, Current version, Latest version, and Status (either an "Up to date" badge or an "Update" button)
|
||||
- **Recheck** button to refresh the latest version from GitHub and re-scan for available updates
|
||||
- **Update All** button to trigger updates on all remote nodes that have a pending update (Skipper or Admiral)
|
||||
| Row | What it shows | Visibility |
|
||||
|-----|---------------|------------|
|
||||
| **Agents** | Active notification agents, formatted `<n> active` or `None` | Always |
|
||||
| **Alert rules** | Per-stack alert rule count, formatted `<n> rule(s)` | Always |
|
||||
| **Auto-heal** | Enabled / total auto-heal policies, formatted `<enabled>/<total>` | Skipper or Admiral |
|
||||
| **Webhooks** | Active outbound webhooks, formatted `<n> active` | Always (when not gated) |
|
||||
| **MFA** | `On`, `Off`, or `Not set` | Local node only |
|
||||
| **Scanning** | Active vulnerability-scan policies, formatted `<n> policy/policies` | Always (when not gated) |
|
||||
| **Backup** | `Enabled` or `Disabled` | Local node only |
|
||||
| **Crash detect** | `On` or `Off` | Always |
|
||||
|
||||
When you click **Update** on a remote node, Sencho sends the update command to the remote instance. The remote pulls the latest Docker image, then spawns a short-lived helper container that performs the compose recreate. The node restarts with the new version, and the status badge transitions from "Updating" to "Updated" once the gateway detects the version change. The "Updated" badge remains visible for 60 seconds before the node returns to "Up to date".
|
||||
Offline nodes show a muted card with the heading and the message `Node is unreachable. Configuration unavailable.` The Online or Offline indicator in the card header confirms reachability at the time of the last fetch.
|
||||
|
||||
If you update the local node, a confirmation dialog appears first, then a reconnection overlay shows while your primary instance restarts.
|
||||
The tab fetches each node's configuration in parallel; a single dead node does not block the rest from rendering.
|
||||
|
||||
## Node Updates
|
||||
|
||||
Click **Check Updates** in the page header to open the **Node Updates** sheet. From here you can read every node's current Sencho version, see which nodes have an update available, and trigger updates one node at a time or across the whole fleet.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/fleet-view/fleet-node-updates.png" alt="Node updates sheet titled 'Node updates · 8 nodes · 1 update available'. A Recheck button and an 'Update all (1)' button sit in the header. Below them, four summary cards read '7 Up to date', '1 Available', '0 Updating', '0 Failed'. A 'Filter nodes…' search box sits above the node table, which has columns Node / Type / Current / Latest / Status. Seven rows show a green 'Up to date' badge; the sencho-pilot-test row reports Current 'unknown' and surfaces an Update button on the right. The footer reads 'LATEST VERSION v0.76.3'." />
|
||||
</Frame>
|
||||
|
||||
The sheet has four sections from top to bottom.
|
||||
|
||||
### Header actions
|
||||
|
||||
- **Recheck** re-queries every node's `/api/meta` endpoint and re-resolves the latest available Sencho version from GitHub Releases (with a Docker Hub fallback if the GitHub API is unreachable). The button disables and shows a spinner while the check is in flight.
|
||||
- **Update all (n)** ships only on Skipper or Admiral. The label carries the count of remote nodes with a pending update. Clicking it dispatches the update to every remote with a known current and latest version.
|
||||
|
||||
### Summary cards
|
||||
|
||||
Four tiles tell you how the fleet is split across update states:
|
||||
|
||||
| Card | Meaning |
|
||||
|------|---------|
|
||||
| **Up to date** | Nodes whose current version equals the latest published Sencho release. |
|
||||
| **Available** | Nodes with a published update they are not yet running. |
|
||||
| **Updating** | Nodes that are currently pulling and recreating with the new image. |
|
||||
| **Failed** | Nodes whose most-recent update attempt did not succeed (failure or timeout). |
|
||||
|
||||
### Node table
|
||||
|
||||
The table lists every registered node, filtered by the search box at the top. Columns:
|
||||
|
||||
| Column | Content |
|
||||
|--------|---------|
|
||||
| **Node** | Node name with a Monitor icon for local nodes and a Globe icon for remotes |
|
||||
| **Type** | `local` or `remote` outline pill |
|
||||
| **Current** | The node's reported Sencho version, in mono. Reads `unknown` if the node has not reported (offline, unreachable, or never connected). |
|
||||
| **Latest** | The newest published Sencho release. Highlighted when newer than Current. |
|
||||
| **Status** | Either an `Up to date` success badge, an `Update` button (per-row), or an in-progress / failed badge with retry and dismiss controls. |
|
||||
|
||||
The latest-version label is resolved from the GitHub Releases API (with a Docker Hub fallback) and cached for 30 minutes. **Recheck** flushes the cache and re-resolves immediately.
|
||||
|
||||
### What happens when you click Update
|
||||
|
||||
When you click **Update** on a remote row, Sencho dispatches the update command to that remote's API. The remote pulls the latest Docker image directly, then spawns a short-lived helper container that bind-mounts the compose working directory from the host and runs `docker compose up -d --force-recreate <service>`. The remote restarts with the new image; the table cell flips from **Updating** to a green **Updated** badge once the gateway detects the version change. The **Updated** state stays visible for a few seconds before the row settles back to **Up to date**.
|
||||
|
||||
When you click **Update** on the local row, a confirmation dialog appears first ("Update local node"). Confirming kicks off the same pull-and-recreate flow on the local Sencho instance. Because the gateway is restarting itself, a full-screen reconnecting overlay takes over the browser tab. The overlay polls `/api/health` every 3 seconds and dismisses itself once the new gateway answers; if the overlay does not clear within 5 minutes it surfaces a *Try Reloading* button instead of waiting indefinitely.
|
||||
|
||||
<Note>
|
||||
Triggering updates (both individual and bulk) requires admin privileges. Users with viewer or operator roles can see update status but cannot initiate updates.
|
||||
Triggering updates (per-row and **Update all**) requires the resolved user to have the admin role. Operator and viewer roles can read update status but cannot initiate updates.
|
||||
</Note>
|
||||
|
||||
**How self-update works:** Each Sencho instance reads its own Docker Compose labels to determine the image name, compose file path, and service name. It pulls the latest image directly, then spawns a helper container that mounts the compose directory from the host and runs `docker compose up --force-recreate`. This approach works regardless of the container's own volume mounts.
|
||||
|
||||
---
|
||||
|
||||
## Fleet Status tab
|
||||
|
||||
The **Status** tab gives you a fleet-wide summary of which automations and security features are active on each node, without having to open each node's settings individually.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/fleet/configuration-tab.png" alt="Fleet Configuration tab showing per-node configuration summaries with one online node and one offline node" />
|
||||
</Frame>
|
||||
|
||||
Each node appears as a card. Online nodes display a compact two-column summary grid:
|
||||
|
||||
| Field | What it shows |
|
||||
|-------|---------------|
|
||||
| **Agents** | Delivery agents configured on this node |
|
||||
| **Alert rules** | Number of per-stack alert rules |
|
||||
| **Auto-heal** | Active auto-heal policies |
|
||||
| **Webhooks** | Outbound webhook triggers |
|
||||
| **MFA** | Whether MFA is set up for the admin account |
|
||||
| **Scanning** | Active vulnerability scan policies |
|
||||
| **Backup** | Cloud backup provider, or Disabled |
|
||||
| **Crash detect** | Whether global crash detection is on |
|
||||
|
||||
Offline nodes show a muted card with "Node is unreachable. Configuration unavailable." The online/offline badge confirms the node's current reachability at the time of the last fetch.
|
||||
|
||||
The tab fetches configuration data from all nodes in parallel. A dead node does not block the others from rendering.
|
||||
|
||||
---
|
||||
|
||||
## How fleet data is fetched
|
||||
|
||||
Fleet View queries all registered nodes in parallel. Each node responds independently; one slow or offline node does not block the others. Local node data comes from the Docker socket and system stats directly. Remote node data is fetched over the Distributed API proxy using each node's Bearer token.
|
||||
Fleet View queries every registered node in parallel. Each node responds independently; one slow or offline node does not block the rest. Local node data comes from the host Docker socket and system stats directly. Remote node data is fetched over the [Distributed API proxy](/features/multi-node) using each node's bearer token.
|
||||
|
||||
<Note>
|
||||
Fleet View always runs on your primary (local) Sencho instance. It is never proxied through a remote node.
|
||||
Fleet View always runs on your control (local) Sencho instance. It is never proxied through a remote node.
|
||||
</Note>
|
||||
|
||||
---
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### Version shows as "unknown" for a remote node
|
||||
|
||||
This means the remote node's `/api/meta` endpoint did not return a valid version. Common causes:
|
||||
|
||||
- **Remote node is offline or unreachable.** Check that the node's API URL and token are correct in the node configuration.
|
||||
- **Remote node is running an older Sencho version** that predates the version reporting feature. Update the remote node manually (pull the latest Docker image and restart the container) to restore version reporting.
|
||||
|
||||
Once the remote node is reachable and running a current Sencho version, its version will appear automatically on the next fleet refresh.
|
||||
|
||||
### "Update All" says no updates available
|
||||
|
||||
The **Update All** button only triggers updates on remote nodes that:
|
||||
|
||||
1. Report a valid version lower than the latest available Sencho release
|
||||
2. Support the self-update capability (requires running in Docker)
|
||||
|
||||
If a remote node's version is unresolvable ("unknown"), the **Update** button on its individual card will still be available, but **Update All** requires both versions to be known for a safe comparison.
|
||||
|
||||
### Update fails with "no such file or directory"
|
||||
|
||||
If a remote node update fails with an error mentioning a compose file path (e.g. "open /path/to/compose.yaml: no such file or directory"), the remote node is running an older Sencho version (prior to v0.42.0) that attempted to access the compose file directly inside the container, where the host path does not exist.
|
||||
|
||||
**Resolution:** Update the remote node manually once by SSH-ing into the host and running:
|
||||
|
||||
```bash
|
||||
docker compose pull && docker compose up -d
|
||||
```
|
||||
|
||||
After this one-time manual update, the node will have the fixed self-update mechanism and all future updates can be triggered from Fleet Overview.
|
||||
|
||||
### Update shows "Remote node is unreachable"
|
||||
|
||||
This means the gateway could not connect to the remote node's `/api/meta` endpoint. Check that:
|
||||
|
||||
- The remote node is powered on and running
|
||||
- The API URL configured for this node is correct and reachable from the gateway
|
||||
- The network allows traffic between the gateway and remote node on the configured port
|
||||
- The remote node's Sencho container is healthy (`docker ps` should show it as running)
|
||||
|
||||
### Update button returns "Admin access required"
|
||||
|
||||
Fleet update operations (individual updates, Update All) require an admin account. If you see a 403 error when clicking **Update**, ask your Sencho administrator to either grant you admin privileges or perform the update on your behalf.
|
||||
|
||||
### Free-mode topology positions don't persist
|
||||
|
||||
Free mode stores node positions in your browser's local storage. The arrangement is per-browser, not synced across devices or users. If positions reset on reload, common causes are:
|
||||
|
||||
- **Browser local storage is disabled** (private/incognito windows, strict tracking-prevention modes). Sencho falls back to in-memory state for the session, but nothing carries over.
|
||||
- **Storage was cleared** by browser cleanup tools, profile reset, or by manually clearing site data.
|
||||
- **A different browser or profile** is being used. Each browser keeps its own arrangement.
|
||||
|
||||
To verify storage is writable, open DevTools → Application → Local Storage on the Sencho origin and confirm the `sencho-topology-preferences` key exists after switching modes or dragging a node.
|
||||
|
||||
### Grouped topology shows everything in one Unlabeled cluster
|
||||
|
||||
Grouped mode clusters remote nodes by their primary node label. When no remotes carry any labels, every remote falls into the **Unlabeled** cluster and the canvas looks similar to Hub mode. A hint banner above the canvas points to **Settings → Nodes** where labels are managed. Add at least one label to two or more remotes and reopen the topology to see the clusters split.
|
||||
<AccordionGroup>
|
||||
<Accordion title="A remote node's Current version reads 'unknown'">
|
||||
The control instance could not resolve the remote's `/api/meta` endpoint. Open **Settings → System → Nodes**, click **Test connection** on the row, and read the toast. The most common causes are a wrong API URL or scheme, a token that has been rotated on the remote (issue a new one and update the saved row), or a firewall or reverse proxy that is not forwarding to the remote's Sencho port. Once the remote is reachable, its version surfaces on the next fleet refresh and the **Update** button becomes available again.
|
||||
</Accordion>
|
||||
<Accordion title="'Update all' says no updates available">
|
||||
The bulk action only triggers updates on remotes that report a valid current version *and* a valid latest version. If a remote's Current reads `unknown`, it is excluded from **Update all** because the gateway cannot perform a safe version comparison. The per-row **Update** button is still available and will run a one-shot update with the latest version it could resolve.
|
||||
</Accordion>
|
||||
<Accordion title="Update fails with 'Remote node is unreachable'">
|
||||
The gateway could not connect to the remote's API while issuing the update. Check that the remote is powered on, that the API URL saved for the row is still correct, that the network allows traffic between the control instance and the remote on the configured port, and that `docker ps` on the remote shows the Sencho container running. Once the remote answers `/api/health`, retry from the row's **Update** button.
|
||||
</Accordion>
|
||||
<Accordion title="The Update button returns 'Admin access required'">
|
||||
Fleet update operations (per-row and **Update all**) are admin-only. Viewer and operator roles can see the table and the badges but the button returns a 403 when clicked. Ask a Sencho administrator either to run the update for you or to grant your account the admin role.
|
||||
</Accordion>
|
||||
<Accordion title="Free-mode topology positions don't persist">
|
||||
Free mode stores node positions in your browser's local storage. The arrangement is per-browser; it does not sync across devices or users. If positions reset on reload, the most common causes are a browser session that has local storage disabled (private or incognito windows, strict tracking-prevention modes, which fall back to in-memory state for the session only), storage cleared by browser cleanup tools or a profile reset, or simply a different browser or profile. To verify storage is writable, open DevTools → Application → Local Storage on the Sencho origin and confirm the `sencho-topology-preferences` key exists after switching modes or dragging a node.
|
||||
</Accordion>
|
||||
<Accordion title="Grouped topology shows everything in one Unlabeled cluster">
|
||||
Grouped mode clusters remote nodes by their primary node label. When no remotes carry any labels, every remote falls into the **Unlabeled** cluster and the canvas looks similar to Hub mode. A hint banner above the canvas points to **Settings · System · Nodes** where labels are managed. Add at least one label to two or more remotes and reopen the topology to see the clusters split.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
@@ -5,28 +5,43 @@ description: Link a stack to a Git repository and keep compose.yaml in sync via
|
||||
|
||||
Git Sources turn any stack into a GitOps target. Point Sencho at a repository, branch, and `compose.yaml` path; pull updates on demand or from CI; and review a diff before applying changes to disk. Optional sibling `.env` sync keeps configuration consistent too.
|
||||
|
||||
Git Sources are available to all Sencho users on the Community tier.
|
||||
<Note>
|
||||
Git Sources are available on every tier, including Community.
|
||||
</Note>
|
||||
|
||||
## How it works
|
||||
|
||||
1. Open a stack's editor and click **Git Source**.
|
||||
1. Open a stack and click the **Git Source** button in the editor toolbar.
|
||||
2. Fill in the repository URL, branch, and compose file path. Add a token if the repo is private.
|
||||
3. Click **Pull now** to fetch the latest compose content. Sencho shows a side-by-side diff.
|
||||
4. Click **Apply** to write the incoming content to disk. Optionally deploy immediately.
|
||||
3. Click **Pull now** to fetch the latest commit. Sencho opens a side-by-side diff between the on-disk files and the incoming version.
|
||||
4. Click **Apply** to write the incoming content to disk. Tick **Deploy after apply** in the same dialog to redeploy in one step.
|
||||
|
||||
Writes land in the stack's existing directory using the same storage Sencho uses for the in-browser editor. Existing history, alerts, and metrics are unaffected.
|
||||
Writes land in the stack's existing directory using the same storage Sencho uses for the in-browser editor.
|
||||
|
||||
## Anatomy of the panel
|
||||
|
||||
<Frame>
|
||||
<img src="/images/git-sources/panel.png" alt="Git Source panel for a stack already linked to a repository, showing populated Repository URL, Branch, Compose file path, the Authentication toggle, the Apply behavior radio group, and a Last applied commit row at the bottom" />
|
||||
</Frame>
|
||||
|
||||
The panel groups four regions:
|
||||
|
||||
- **Pending update banner.** Appears at the top when a webhook in **Review only** mode has fetched a new commit. Click **Review** to re-fetch the incoming commit and open the diff dialog.
|
||||
- **Form fields.** Repository URL, branch, compose file path, optional sibling `.env` sync, authentication toggle, and the apply behavior radio group.
|
||||
- **Last applied stat strip.** Shows the short SHA of the last commit Sencho applied to disk, plus the timestamp of the most recent successful save or pull.
|
||||
- **Footer actions.** **Remove** disconnects the source without touching the stack files; **Pull now** fetches the configured branch's HEAD; **Save** or **Update** persists form changes after a reachability check passes.
|
||||
|
||||
## Create a stack from a Git repository
|
||||
|
||||
Skip the "empty stack then link later" detour and point at a repo from the start. Click **Create Stack** in the sidebar, switch to the **From Git** tab, and fill in the same fields you would on the Git Source panel.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/git-sources/create-from-git-tab.png" alt="Create Stack dialog with the From Git tab selected, showing stack name, repository URL, branch, compose path, and a Deploy after create checkbox" />
|
||||
<img src="/images/git-sources/create-from-git-tab.png" alt="New stack dialog with the From Git tab selected, showing stack name, repository URL, branch, compose path, sibling .env toggle, authentication toggle, apply behavior radio group, a Deploy after create checkbox, and an HTTPS REPOS ONLY footer hint" />
|
||||
</Frame>
|
||||
|
||||
Sencho fetches the compose file, validates it with `docker compose config`, writes it to a fresh stack directory, and links the git source in one step. The last-applied commit sha is seeded from the fetch so the first pull produces a clean diff rather than a "local edits detected" warning.
|
||||
Sencho fetches the compose file, validates it with `docker compose config`, writes it to a fresh stack directory, and links the Git source in one step. The last-applied commit SHA is seeded from the fetch so the first manual pull produces a clean diff rather than a "local edits detected" warning.
|
||||
|
||||
Tick **Deploy after create** to run `docker compose up -d` immediately after the files land. If the deploy fails, the stack and git source are kept on disk so you can fix the underlying issue (missing image, port conflict, host resources) and retry the deploy from the editor.
|
||||
Tick **Deploy after create** to run `docker compose up -d` immediately after the files land. If the deploy fails, the stack and Git source are kept on disk so you can fix the underlying issue (missing image, port conflict, host resources) and retry the deploy from the editor.
|
||||
|
||||
### Failure modes on create
|
||||
|
||||
@@ -34,22 +49,18 @@ Tick **Deploy after create** to run `docker compose up -d` immediately after the
|
||||
|-----------|--------------|
|
||||
| Stack name already exists | Sencho returns **409** with "Stack already exists" and makes no changes on disk or in the database. Pick a different name or remove the existing stack. |
|
||||
| Repository unreachable or auth failed | Fetch fails before anything is created. The form stays open with an error toast describing the cause. |
|
||||
| Fetched compose fails validation | The stack directory is not created and no git source row is inserted. The error toast shows the `docker compose config` message. |
|
||||
| Fetch + validate succeed but optional deploy fails | The stack and git source are kept. The toast reads "Stack created, but deploy failed: ..." and you can retry the deploy from the editor. |
|
||||
| Fetched compose fails validation | The stack directory is not created and no Git source row is inserted. The error toast shows the `docker compose config` message. |
|
||||
| Fetch + validate succeed but optional deploy fails | The stack and Git source are kept. The toast reads "Stack created, but deploy failed: ..." and you can retry the deploy from the editor. |
|
||||
|
||||
## Configure a source
|
||||
|
||||
<Frame>
|
||||
<img src="/images/git-sources/panel.png" alt="Git Source panel with repository URL, branch, compose path, and apply mode" />
|
||||
</Frame>
|
||||
|
||||
| Field | Description |
|
||||
|-------|-------------|
|
||||
| **Repository URL** | `https://github.com/your-org/your-repo.git` (HTTPS only) |
|
||||
| **Branch** | Branch to track (e.g. `main`) |
|
||||
| **Compose file path** | Path within the repo (e.g. `deploy/compose.yaml`) |
|
||||
| **Also sync sibling `.env`** | When enabled, also pulls `.env` from the same directory as the compose file |
|
||||
| **Auth** | `None` for public repos, `Personal Access Token` for private repos |
|
||||
| **Also sync sibling `.env` file** | When enabled, also pulls the `.env` from the same directory as the compose file. The form shows the resolved path inline (e.g. `deploy/.env` for a compose at `deploy/compose.yaml`). |
|
||||
| **Authentication** | **Public (no auth)** for public repos, **Personal Access Token** for private repos |
|
||||
| **Apply behavior** | See the three modes below |
|
||||
|
||||
Saving runs a reachability check against the repository. If the URL is wrong, the token is invalid, the branch does not exist, or the file is missing, Sencho surfaces the error inline and nothing is persisted.
|
||||
@@ -58,10 +69,12 @@ Saving runs a reachability check against the repository. If the URL is wrong, th
|
||||
|
||||
| Mode | What happens when a webhook fires |
|
||||
|------|-----------------------------------|
|
||||
| **Review only** | Sencho fetches + validates the incoming commit and marks the stack as having a pending update. You review the diff and apply manually. |
|
||||
| **Auto-write** | Sencho writes the new compose + env to disk automatically but does not redeploy. Use this when another process handles rollout. |
|
||||
| **Review only** | Sencho fetches and validates the incoming commit and marks the stack as having a pending update. You review the diff and apply manually. |
|
||||
| **Auto-write files** | Sencho writes the new compose and env to disk automatically but does not redeploy. Use this when another process handles rollout. |
|
||||
| **Auto-deploy** | Sencho writes the files and immediately runs `docker compose up -d` so the stack picks up the new configuration. |
|
||||
|
||||
Auto-deploy implies Auto-write: you cannot deploy automatically without also writing the new files first.
|
||||
|
||||
You can always override on the spot: when you click **Apply** in the diff dialog, a **Deploy after apply** checkbox lets you deploy regardless of the configured mode.
|
||||
|
||||
## Pulling and reviewing changes
|
||||
@@ -69,42 +82,44 @@ You can always override on the spot: when you click **Apply** in the diff dialog
|
||||
Click **Pull now** on the Git Source panel to fetch the latest commit on the configured branch.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/git-sources/diff-dialog.png" alt="Diff dialog showing a side-by-side comparison between the on-disk compose.yaml and the incoming commit" />
|
||||
<img src="/images/git-sources/diff-dialog.png" alt="GIT · PULL PREVIEW dialog for the demo-app stack, showing a Local edits detected on disk warning above a Monaco side-by-side diff between the on-disk compose.yaml and the incoming commit, with a Deploy after apply checkbox and an Apply button in the footer" />
|
||||
</Frame>
|
||||
|
||||
The diff dialog shows:
|
||||
|
||||
- The commit sha being compared (short form, next to the stack name)
|
||||
- A side-by-side compare of the on-disk compose file and the incoming version
|
||||
- A `.env` tab when the source is configured to sync `.env`
|
||||
- A **validation** banner when the incoming compose fails `docker compose config` (you cannot apply an invalid file)
|
||||
- A **local edits detected** banner when the on-disk content differs from the last applied commit. Applying in this state overwrites those edits. The Apply button becomes a confirmation prompt.
|
||||
- The `GIT · PULL PREVIEW` kicker and the short SHA of the incoming commit at the top.
|
||||
- A side-by-side compare of the on-disk compose file and the incoming version.
|
||||
- A `.env` tab when the source is configured to sync `.env`.
|
||||
- An **Incoming compose failed validation** banner when the incoming compose fails `docker compose config`. The Apply button stays disabled until validation passes.
|
||||
- A **Local edits detected on disk** banner when the on-disk content differs from the last applied commit. Applying in this state opens an **Overwrite local edits?** confirmation modal whose primary button is **Overwrite and apply**.
|
||||
|
||||
### Pending updates
|
||||
|
||||
When a webhook fires in **Review only** mode, the stack gets a pending update badge in the sidebar and a dot on the **Git Source** button in the editor. Clicking either opens the diff dialog with the incoming content already loaded; there's no second network round-trip to apply.
|
||||
When a webhook fires in **Review only** mode, the stack gets a pending GitBranch icon next to its row in the sidebar and a pulsing dot on the **Git Source** button in the editor. Clicking either re-fetches the commit and opens the diff dialog; the panel also shows a **Pending update** banner with a **Review** button.
|
||||
|
||||
If the same stack also has an image update available, the image-update dot in the sidebar takes priority over the Git source icon, so only the update dot renders. The pending Git source is still surfaced inside the editor on the **Git Source** button.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/git-sources/sidebar-badge.png" alt="Sidebar stack entry with a small branded dot indicating a pending Git source update" />
|
||||
<img src="/images/git-sources/sidebar-badge.png" alt="Sidebar stack list with a small GitBranch icon next to the demo-app entry indicating a pending Git source update" />
|
||||
</Frame>
|
||||
|
||||
Click **Dismiss** on the Git Source panel to discard a pending update without applying.
|
||||
Click **Dismiss** in the diff dialog to discard a pending update without applying.
|
||||
|
||||
## Trigger from CI with a webhook
|
||||
|
||||
Git sources integrate with Sencho's existing webhook system. Create a webhook targeting the stack with the **Git source sync** action.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/git-sources/webhook-action.png" alt="Webhook creation form with Git source sync selected in the action dropdown" />
|
||||
<img src="/images/git-sources/webhook-action.png" alt="New webhook form with Name, Stack and Action fields. The Action select is open, showing Deploy, Restart, Stop, Start, Pull and Update, and Git source sync as options, with Git source sync highlighted at the bottom" />
|
||||
</Frame>
|
||||
|
||||
The webhook's behavior on trigger depends on the source's apply mode:
|
||||
|
||||
- **Review only**: fetch + validate + diff, mark pending.
|
||||
- **Auto-write**: fetch, validate, write to disk.
|
||||
- **Review only**: fetch, validate, diff, mark pending.
|
||||
- **Auto-write files**: fetch, validate, write to disk.
|
||||
- **Auto-deploy**: fetch, validate, write, deploy.
|
||||
|
||||
The Git source sync action is only selectable on webhooks whose target stack already has a Git source configured. Webhook triggers for a single source are debounced so a runaway pipeline cannot overwhelm Sencho (or your repository host's rate limits); the dashboard records the skipped trigger in the webhook's execution history.
|
||||
The Git source sync action is only selectable on webhooks whose target stack already has a Git source configured. Webhook triggers for a single source are debounced on a 10-second window so a runaway pipeline cannot overwhelm Sencho (or your repository host's rate limits); the dashboard records the skipped trigger in the webhook's execution history.
|
||||
|
||||
### GitHub Actions example
|
||||
|
||||
@@ -129,17 +144,21 @@ For private repositories, use a Personal Access Token scoped to read access on t
|
||||
- **GitLab**: a project or group access token with the `read_repository` scope.
|
||||
- **Bitbucket**: an app password with **Repositories: Read**.
|
||||
|
||||
Paste the token into the **Token** field and save. Sencho stores it encrypted at rest and never returns it in API responses or UI. When editing the source later, the token field shows a masked placeholder; leave it blank to keep the stored value, or type a new token to replace it. Switching the auth type to **None** clears the stored token.
|
||||
Paste the token into the **Token** field and save. Sencho stores it encrypted at rest and never returns it in API responses or UI. When editing the source later, the token field shows a masked placeholder; leave it blank to keep the stored value, or type a new token to replace it. Switching the auth type back to **Public (no auth)** clears the stored token.
|
||||
|
||||
The encryption boundary covers the pending update payload too: every pull caches the fetched compose and env content in the database so the diff dialog can reopen without a refetch, and that cached content is encrypted at rest in the same way as the token, since compose files routinely embed secrets via env interpolation.
|
||||
|
||||
## Local edits vs Git
|
||||
|
||||
Sencho tracks a hash of the compose + env contents at the moment of the last apply. When you pull, it compares that hash against the current on-disk content.
|
||||
Sencho tracks a hash of the compose and env contents at the moment of the last apply. When you pull, it compares that hash against the current on-disk content.
|
||||
|
||||
- Matching hash: applying overwrites content that Sencho itself last wrote.
|
||||
- Differing hash: someone edited the files outside Git. The diff dialog shows a warning, and Apply requires confirmation.
|
||||
- Differing hash: someone edited the files outside Git. The diff dialog shows the **Local edits detected on disk** banner, and Apply requires the **Overwrite local edits?** confirmation.
|
||||
|
||||
The in-browser editor and the Git Source panel both write to the same files, so you can always fall back to editing locally. The next pull will just flag the divergence rather than silently clobbering your edits.
|
||||
|
||||
Pulls, applies, and create-from-git operations on the same stack are serialized by a per-stack lock, so a webhook that fires in the middle of a manual apply waits for the apply to finish rather than racing it.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
<AccordionGroup>
|
||||
@@ -151,7 +170,7 @@ The in-browser editor and the Git Source panel both write to the same files, so
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Authentication failed">
|
||||
You supplied a token and the git host rejected it outright. The token is missing, expired, or lacks read access. Generate a new token and replace the value in the **Token** field. Sencho reports this as a form error, not a Sencho login problem, so you stay signed in.
|
||||
You supplied a token and the Git host rejected it outright. The token is missing, expired, or lacks read access. Generate a new token and replace the value in the **Token** field. Sencho returns this as a 400 form error rather than a 401, so an upstream auth failure does not sign you out of the dashboard.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Branch not found">
|
||||
@@ -163,19 +182,23 @@ The in-browser editor and the Git Source panel both write to the same files, so
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Compose validation failed">
|
||||
Sencho runs `docker compose config` against the incoming content before letting you apply. The error banner shows the exact message. Common causes: unresolved `${VAR}` interpolation (fix by enabling sibling `.env` sync and committing the file), invalid `include:` paths, or schema issues introduced by a recent compose change.
|
||||
Sencho runs `docker compose config` against the incoming content before letting you apply. The error banner shows the exact message. Common causes: unresolved `${VAR}` interpolation (commit a `.env` file next to the compose file and enable sibling `.env` sync), invalid `include:` paths, or schema issues introduced by a recent compose change. Validation has a 10-second budget; an unusually large compose with many services may need to be split.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Local edits detected">
|
||||
The on-disk files diverge from the last applied Git commit. Either apply anyway to overwrite the local edits (the confirmation prompt makes this explicit), or discard local work with a redeploy from the stack editor, or commit your local changes back to the repo so the diff becomes clean.
|
||||
The on-disk files diverge from the last applied Git commit. Either confirm **Overwrite and apply** to take the incoming content, or discard local work with a redeploy from the stack editor, or commit your local changes back to the repo so the diff becomes clean.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Webhook skipped (rate limited)">
|
||||
Sencho debounces rapid-fire triggers per source. Wait a few seconds and retry, or consolidate multiple CI triggers into a single call at the end of your pipeline.
|
||||
Sencho debounces rapid-fire triggers on a 10-second window per source. Wait at least 10 seconds and retry, or consolidate multiple CI triggers into a single call at the end of your pipeline. Skipped triggers appear in the webhook's execution history.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Network timeout">
|
||||
The clone did not finish in time. Fetches run with a bounded timeout to keep a slow or unreachable host from hanging the stack panel. Check that the Sencho host can reach the repository host (proxies, firewalls, DNS) and try again. If the repository is genuinely large, pin a smaller compose subpath or mirror it somewhere closer to the Sencho host.
|
||||
The clone did not finish in time. Fetches run with a 30-second timeout to keep a slow or unreachable host from hanging the stack panel. Check that the Sencho host can reach the repository host (proxies, firewalls, DNS) and try again. If the repository is genuinely large, pin a smaller compose subpath or mirror it somewhere closer to the Sencho host.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Pending commit has changed since this pull was fetched">
|
||||
You opened a diff dialog, then a webhook fired and replaced the pending commit before you clicked **Apply**. Close the dialog and reopen the panel to load the latest pending commit; the **Review** button will fetch the newer one.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Applied but deploy failed">
|
||||
|
||||
@@ -1,89 +1,183 @@
|
||||
---
|
||||
title: Global Observability
|
||||
description: A unified, searchable log stream from every container across all your stacks.
|
||||
description: A unified, searchable log stream from every running container on the active node.
|
||||
---
|
||||
|
||||
The **Logs** tab aggregates output from all running containers into a single scrollable view. Instead of tailing logs one container at a time, you see everything in one place, with contextual metrics and filtering to focus on what matters.
|
||||
The **Logs** tab aggregates output from every running container on the active node into a single live feed. Instead of tailing one container at a time, you scan everything in one place, with a live event-rate readout, stream / level / stack filters, and a downloadable replay of the current buffer.
|
||||
|
||||
<Note>
|
||||
Logs is hub-only and is hidden from the nav strip when a remote node is the active selection. See [Multi-Node Management](/features/multi-node#what-top-level-views-show-when-a-remote-node-is-active).
|
||||
</Note>
|
||||
|
||||
<Frame>
|
||||
<img src="/images/global-observability/global-observability-overview.png" alt="Global Observability view with live masthead, signal rail, and streaming log feed" />
|
||||
<img src="/images/global-observability/global-observability-overview.png" alt="Global Observability page with the live logs masthead, four-tile signal rail, filter strip, and a feed body showing day-banded rows in info green and error rose." />
|
||||
</Frame>
|
||||
|
||||
## Cockpit layout
|
||||
## Status masthead
|
||||
|
||||
The page is laid out as a vertical stack of surfaces, each with a distinct role:
|
||||
The masthead at the top of the page is the single place to read the stream's current condition.
|
||||
|
||||
| Surface | What it shows |
|
||||
|---------|---------------|
|
||||
| **Masthead** | Stream state (`Streaming`, `Idle`, or `Offline`) with a live status dot. Right side shows time since the last event and current session uptime. |
|
||||
| **Signal rail** | Four tiles: `EVENTS / MIN` with a 60-second sparkline, `ERRORS` count, `WARNINGS` count, and distinct `CONTAINERS` count. Values reflect the current buffer. |
|
||||
| **Filter strip** | Search box, stacks multi-select, stream pills (`All`, `Out`, `Err`), and level pills (`All`, `Info`, `Warn`, `Error`). |
|
||||
| **Feed** | Chronological rows grouped by day-banded headers (`NOW`, `2M AGO`, `TODAY 14:20`). Each row shows a severity dot, timestamp, container name, and message. Error rows are tinted rose; warning rows tinted amber. |
|
||||
| **Chip strip** | Floating bottom-right controls for `Pause`, `Clear`, and `Download`. |
|
||||
<Frame>
|
||||
<img src="/images/global-observability/masthead.png" alt="Logs masthead with the kicker 'LIVE LOGS · NODE · LOCAL', the state word 'Streaming' in editorial italics next to a pulsing brand-cyan dot, and the meta tiles 'LAST EVENT NOW' and 'SESSION 1M 43S' on the right." />
|
||||
</Frame>
|
||||
|
||||
## Streaming
|
||||
It carries:
|
||||
|
||||
Logs stream live via Server-Sent Events the moment you open the page. The masthead dot pulses cyan while the stream is active; it settles to grey when no events have arrived in the last ten seconds, and turns rose if the connection fails.
|
||||
- A **kicker** in uppercase mono tracking that reads `LIVE LOGS · NODE · <NAME>`. The local node renders as `LOCAL`; remote nodes render with their configured name uppercased.
|
||||
- A **pulsing dot** that mirrors the stream tone: brand cyan and pulsing while events arrive within the live window, gray and steady when the stream goes quiet, solid rose when the underlying SSE connection has errored.
|
||||
- A **state word** set in the editorial display face: `Streaming` while the dot is live, `Idle` when the stream has gone quiet, `Offline` when the connection has failed.
|
||||
- A **LAST EVENT** stat that names the band of the most recent event using the same vocabulary as the feed (`NOW`, `2M AGO`, `1H AGO`, then the calendar date), so you can see how stale the latest line is at a glance.
|
||||
- A **SESSION** stat counting how long this Logs tab has been open, formatted with uppercase letter suffixes: `1H 43M` once the session crosses an hour, `0M 12S` under it. The session resets when you switch the active node, since the stream re-connects against the new node.
|
||||
|
||||
If your browser cannot open an `EventSource` connection, Sencho falls back to polling every five seconds automatically. The transport choice is transparent, and no setting is needed to enable streaming.
|
||||
The state word flips to `Idle` after ten seconds without an event. That is a feature of the readout, not the stream: the SSE connection stays open in the background and the dot will pulse cyan again on the next event.
|
||||
|
||||
## Log format
|
||||
## Signal rail
|
||||
|
||||
Each row shows:
|
||||
A four-tile rail directly under the masthead summarizes the current buffer.
|
||||
|
||||
- **Severity dot** - green for `INFO`, amber for `WARN`, rose with a soft glow for `ERROR`
|
||||
- **Timestamp** - when the log line was emitted (24-hour local time)
|
||||
- **Container name** - the specific container that produced the line, rendered in brand cyan
|
||||
- **Message** - the raw log output, tinted rose for `STDERR` lines
|
||||
<Frame>
|
||||
<img src="/images/global-observability/signal-rail.png" alt="Four-tile signal rail: 'EVENTS / MIN' showing 20 with a 60-second sparkline, 'ERRORS' showing 418 in destructive red, 'WARNINGS' showing 9 in amber, 'CONTAINERS' showing 8 in stat-value white." />
|
||||
</Frame>
|
||||
|
||||
Rows for errors are tinted rose across the full width; warnings are tinted amber. The whole-row tint makes problems easy to scan without reading every line.
|
||||
| Tile | What it shows |
|
||||
|------|---------------|
|
||||
| **EVENTS / MIN** | Total events received in the last sixty seconds, with a 60-bucket sparkline (one bucket per second) under the number. The sparkline shifts left every tick so you read it as a live rhythm rather than a static chart. |
|
||||
| **ERRORS** | Count of `ERROR`-level lines in the buffer. Tints destructive red whenever the count is above zero; otherwise the number sits in subtitle gray. |
|
||||
| **WARNINGS** | Count of `WARN`-level lines in the buffer. Tints warning amber whenever the count is above zero; otherwise the number sits in subtitle gray. |
|
||||
| **CONTAINERS** | Count of distinct container names present in the in-memory buffer (after the **Clear** cutoff). Containers whose last entry has aged out of the buffer drop out of the count. |
|
||||
|
||||
## Day bands
|
||||
All four counters are scoped to what is currently in the buffer. **Clear** sweeps them along with the visible rows.
|
||||
|
||||
When consecutive log rows span a large time gap, a tracked-mono header is drawn between them showing how long ago the next block of rows was emitted. Labels roll up as time passes: recent events are `NOW`, then `2M AGO`, `12M AGO`, `1H AGO`, and eventually a full calendar date.
|
||||
## Filter strip
|
||||
|
||||
## Filtering logs
|
||||
A row of controls between the signal rail and the feed lets you narrow what you see without touching the underlying stream.
|
||||
|
||||
Use the filter strip above the feed to narrow what you see:
|
||||
<Frame>
|
||||
<img src="/images/global-observability/filter-strip.png" alt="Filter strip with a 'Search logs...' textbox, a 'STACKS · ALL' dropdown trigger, a stream segmented control with 'ALL OUT ERR' (ALL selected), and a level segmented control with 'ALL INFO WARN ERROR' (ALL selected)." />
|
||||
</Frame>
|
||||
|
||||
| Control | What it does |
|
||||
|---------|-------------|
|
||||
| **Search** | Full-text filter across the message, container name, and stack name (case-insensitive). |
|
||||
| **Stacks** | Checkbox dropdown; select one or more stacks to show only their containers' logs. Displays "All" when nothing is selected. |
|
||||
| **Stream pills** | Filter by output stream: `All`, `Out` (STDOUT), or `Err` (STDERR). |
|
||||
| **Level pills** | Filter by severity: `All`, `Info`, `Warn`, or `Error`. |
|
||||
|---------|--------------|
|
||||
| **Search** | Case-insensitive substring filter against the message body, container name, and stack name. |
|
||||
| **Stacks** | Dropdown of every stack discovered on the active node; ticking one or more boxes restricts the feed to those stacks. The trigger reads `Stacks · All` when nothing is selected and `Stacks · <n>` once at least one box is ticked. |
|
||||
| **Stream** | Segmented control with `All`, `Out`, `Err` for filtering by stream source. `Out` is `STDOUT`, `Err` is `STDERR`. |
|
||||
| **Level** | Segmented control with `All`, `Info`, `Warn`, `Error` for filtering by detected severity. |
|
||||
|
||||
Filters combine. For example, you can show only `Error` lines from a specific stack while searching for a keyword.
|
||||
Filters AND together and run in the browser, so there is no per-keystroke round-trip to the backend. Switching them does not lose any buffered data; rows reappear if you broaden the filter again.
|
||||
|
||||
## Pause and resume
|
||||
## Feed
|
||||
|
||||
Click **Pause** in the chip strip to freeze the feed while you inspect a run of events. The stream keeps filling in the background, and a cyan pill appears showing how many new events are waiting. Click the pill or **Resume** to flush them and return to live mode.
|
||||
The body of the page is a chronological list of every log line in the buffer that passes the current filter.
|
||||
|
||||
## Controls
|
||||
<Frame>
|
||||
<img src="/images/global-observability/feed-bands.png" alt="Feed body showing day band headers '3M AGO' and '2M AGO' between groups of rows. Each row shows a green INFO dot, a 24-hour timestamp, a cyan container name (seerr, profilarr), and the raw log message." />
|
||||
</Frame>
|
||||
|
||||
| Button | What it does |
|
||||
Each row carries:
|
||||
|
||||
- A small **severity dot** tinted by detected level: success green for `INFO`, warning amber for `WARN`, destructive rose with a soft glow for `ERROR`.
|
||||
- A **timestamp** in 24-hour local time (`HH:MM:SS`), set in tabular nums so columns line up across rows.
|
||||
- A **container name** in brand cyan. Hover to see the full `<stack>/<container>` path in a tooltip; the displayed name is the compose service name with the project prefix and replica suffix stripped.
|
||||
- A **message** containing the raw log output. `STDERR`-source lines render in destructive rose; `STDOUT`-source lines render in stat-value white.
|
||||
|
||||
A row's whole-row tint follows its detected level. Error rows pick up a subtle destructive wash, warning rows pick up an amber wash, info rows have no tint and gain a subtle accent on hover. The detected level is independent of the source, so an `STDOUT` line that contains the word `ERROR:` still gets the rose row tint.
|
||||
|
||||
When consecutive rows span more than a minute, a tracked-mono **band header** is inserted between them. Bands roll up over time: `NOW` for the last sixty seconds, then `2M AGO`, `12M AGO`, `1H AGO`, and eventually the calendar date.
|
||||
|
||||
The feed shows the most recent three hundred rows that pass the filter. If more rows match, a `Showing last 300 of <n>` notice sits at the top of the feed so you know lines have been clipped, even though the full ring buffer behind the scenes still holds them.
|
||||
|
||||
The feed uses two empty states, both rendered as a tracked-mono kicker over an italic caption. Before any event arrives, the body shows `Awaiting events` over `Logs will appear here as containers emit them.` When the current filter excludes every row in the buffer, the body shows `No matches` over `Try a broader filter to see logs again.`
|
||||
|
||||
The viewport autoscrolls to follow new rows. If you scroll up to read older entries, autoscroll pauses so the feed doesn't yank you back. Scrolling within fifty pixels of the bottom re-engages it.
|
||||
|
||||
## Pause, clear, download
|
||||
|
||||
A floating control strip is anchored to the bottom-right of the feed.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/global-observability/paused-resume-chip.png" alt="Floating control strip with a brand-cyan '22 NEW · RESUME' pill on the left and the Pause/Clear/Download chip group on the right. The Pause button shows a play icon and the label 'RESUME', indicating the stream is currently paused." />
|
||||
</Frame>
|
||||
|
||||
| Action | What it does |
|
||||
|--------|--------------|
|
||||
| **Pause / Resume** | Freezes and resumes the feed without disconnecting the stream. |
|
||||
| **Clear** | Clears the current log buffer in the UI. Does not delete logs from Docker. |
|
||||
| **Download** | Exports the currently filtered logs as a `.txt` file. The filename includes an ISO timestamp for easy identification. |
|
||||
| **Pause / Resume** | Freezes the feed without dropping the SSE connection. New events keep flowing into the in-memory buffer so the page does not lose ground. While paused, a brand-cyan **`<n> NEW · RESUME`** pill appears to the left of the strip and counts the events queued since you paused; clicking either the pill or the **Resume** button flushes them and returns to the live view. The pause toggle's icon and label flip between play / `RESUME` and pause / `PAUSE` in lockstep with the state. |
|
||||
| **Clear** | Hides every row that arrived before the click; counters reset to match. The underlying Docker logs are untouched. New events that arrive after the click reappear in the feed. |
|
||||
| **Download** | Exports the rows currently passing the filter as a `.txt` file. The filename is `sencho-logs-<ISO8601>.txt` (with `:` and `.` replaced by `-` so Windows accepts the name) and each line is formatted as `[<ISO timestamp>] [<stack>/<container>] <LEVEL>: <message>`. The button disables when the filter currently matches zero rows. |
|
||||
|
||||
## Auto-scroll
|
||||
## Filtering by level
|
||||
|
||||
The feed automatically scrolls to the bottom as new rows arrive. If you scroll up to inspect older entries, auto-scroll pauses so you are not pulled away. Scrolling back to the bottom re-enables it.
|
||||
The **Level** segmented control drops the feed to a single severity. Pairing it with **Stacks** is a fast way to investigate a misbehaving stack.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/global-observability/error-only-filter.png" alt="Logs page with the Level filter set to 'Error'. The feed shows only error-tinted rows from the sencho and radarr containers, with day bands '6M AGO' through 'NOW' visible." />
|
||||
</Frame>
|
||||
|
||||
The level on each row is detected from the message body, not the stream source. Sencho's parser scans the line for keywords (`info`, `debug`, `warn`, `error`) and structured fields (`level=...`, `[INFO]`, `Exception:`) before falling back to a source-based default of `STDERR → ERROR`, `STDOUT → INFO`. That means an `STDOUT` line containing `ERROR:` classifies as `ERROR` (and gets the rose tint), and an `STDERR` line containing `info` classifies as `INFO` (and gets the green dot).
|
||||
|
||||
## How streaming works
|
||||
|
||||
The page opens an `EventSource` against `/api/logs/global/stream?nodeId=<id>` as soon as it mounts. Each event is one log line emitted by `docker logs --follow` against every running container on the active node, demultiplexed and parsed before it leaves the server. When the connection opens, the server replays the last five hundred lines per container before switching to live tail, so you see immediate context instead of an empty feed while you wait for the next event.
|
||||
|
||||
A 30-second SSE keep-alive comment is written every tick to defeat reverse-proxy idle timeouts, so the stream stays open behind nginx, Cloudflare, Caddy, and so on without configuration.
|
||||
|
||||
If the browser cannot open the `EventSource` connection at all (a corporate proxy that strips SSE, an outbound block on the streaming endpoint), the page silently falls back to polling `/api/logs/global` every five seconds. The polling endpoint returns the last five hundred lines on each call as a single snapshot. The transport choice is automatic; you cannot pick one explicitly.
|
||||
|
||||
## Display limits
|
||||
|
||||
To keep the browser responsive, the log viewer enforces two capacity limits:
|
||||
Two budgets keep the page responsive even on a chatty fleet:
|
||||
|
||||
- **Memory buffer** - a maximum of 2,000 log entries are held in memory at a time. Older entries are dropped as new ones arrive.
|
||||
- **Rendered rows** - only the most recent 300 entries from the filtered set are rendered as DOM nodes. If more entries match, a notice appears at the top of the feed showing how many entries exist versus how many are displayed.
|
||||
- **Buffer** of two thousand entries in the browser's memory. As new events arrive past the cap, the oldest entries are dropped. The buffer keeps filling while the feed is paused, so a long pause does not stretch memory unbounded.
|
||||
- **Rendered rows** capped at three hundred. Beyond that, the `Showing last 300 of <n>` banner appears so you know more matches exist behind the slice.
|
||||
|
||||
For deeper investigation, use the [per-container log viewer](/features/editor#log-viewer) in the Editor tab, or `docker compose logs` directly from the [Host Console](/features/host-console).
|
||||
When you need to chase older lines than the buffer holds, jump to the [per-container log viewer](/features/editor#log-viewer) in the Editor tab or run `docker compose logs` from the [Host Console](/features/host-console).
|
||||
|
||||
## Active node
|
||||
|
||||
The Logs tab is scoped to one node at a time, the active node selected from the sidebar's node switcher.
|
||||
|
||||
- The kicker reads `LIVE LOGS · NODE · LOCAL` for the local node and `LIVE LOGS · NODE · <NAME>` for any registered remote node, so you always know which fleet member you are watching.
|
||||
- Switching the active node tears down the SSE connection, clears the buffer, and opens a fresh stream against the new node. The session timer in the masthead resets at the same moment.
|
||||
- There is no fleet-wide aggregation view. To compare two nodes, open the Logs tab in two browser tabs and switch the active node in each.
|
||||
|
||||
## Refresh cadence
|
||||
|
||||
The page is built from several independent loops so the masthead and the sparkline animate without flooding the network:
|
||||
|
||||
| Surface | Cadence |
|
||||
|---------|---------|
|
||||
| UI tick (uptime, band labels, live-window check) | 1 second |
|
||||
| Buffer flush from the SSE batch into the visible feed | 500 ms |
|
||||
| Polling fallback when SSE never opens | 5 seconds |
|
||||
| SSE keep-alive heartbeat from the server | 30 seconds |
|
||||
| Sparkline rolling window | 60 buckets × 1 second = 60 seconds |
|
||||
| `Idle` threshold (no events received) | 10 seconds |
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="The masthead reads 'Offline' and a 'Failed to fetch logs. Retrying...' banner is visible">
|
||||
The polling fallback got an HTTP error when it tried to read `/api/logs/global`. The most common cause is that the active remote node is unreachable; check the same node from **Settings · Nodes** and confirm the API token is still valid. If the active node is local, check the host's Docker daemon is reachable on the socket Sencho is configured to use. The page retries the request every five seconds; once it succeeds, the banner clears on its own.
|
||||
</Accordion>
|
||||
<Accordion title="The dot is gray and the state reads 'Idle' even though logs are appearing further down the page">
|
||||
`Idle` only checks whether an event has arrived in the last ten seconds. If your fleet is genuinely quiet (a maintenance window, a single low-traffic container), the dot rests gray and the band headers in the feed roll up to `2M AGO`, `5M AGO`, and so on. The next event flips the dot back to brand cyan and the state word back to `Streaming` on the next one-second tick.
|
||||
</Accordion>
|
||||
<Accordion title="A line that says 'ERROR' renders without the rose row tint">
|
||||
Severity is detected once per line by scanning the message body. Some loggers print `ERROR` as part of a stack trace header on a line that does not itself match the parser's keyword set, in which case the line is classified by its source: `STDOUT` falls back to `INFO`. To force the row to your preferred level, change the application's log format to something the parser recognises (`level=error`, `[ERROR]`, or a leading `Error:` / `Exception:`). The behavior is intentional: it keeps a noisy `STDOUT` line that quotes the word `ERROR` from polluting the error count.
|
||||
</Accordion>
|
||||
<Accordion title="The Resume pill keeps growing while I read a paused feed">
|
||||
The buffer caps at two thousand entries even while paused. Once you cross the cap, the oldest queued events drop, and the pill's count reflects the size of the queue the next flush will deliver, not the total number of events that arrived during the pause. Resume to flush, then re-pause if you need to keep reading; the page will not lose its place because autoscroll only re-engages when you scroll back to the bottom.
|
||||
</Accordion>
|
||||
<Accordion title="Clear hid the rows but a few old lines came back a moment later">
|
||||
**Clear** records the click time and hides every row whose timestamp is earlier than that moment, but the SSE stream keeps delivering events whose timestamps were generated by Docker before the click. A line emitted at `16:51:43.987` that arrives at the browser at `16:51:44.020` will still pass the filter even if you clicked **Clear** at `16:51:44.000`. The lag is normally a few hundred milliseconds; click **Clear** again to hide them.
|
||||
</Accordion>
|
||||
<Accordion title="Logs disappear when I switch the active node">
|
||||
The page is scoped to one node at a time. Switching the active node tears down the current SSE connection, drops the in-browser buffer, and connects against the new node. To compare two nodes side by side, open the Logs tab in two browser tabs and switch the active node in each.
|
||||
</Accordion>
|
||||
<Accordion title="I want to see logs from every node at once">
|
||||
The Logs tab does not aggregate across the fleet. Use the per-container log viewer in the Editor for a single container's history, or jump to the Host Console for ad-hoc `docker compose logs` against any node. Fleet-wide aggregation is not on the current roadmap; pipe container logs to your existing observability stack (Loki, ELK, Datadog) if you need cross-node search.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
<Note>
|
||||
Log data retention on the backend is controlled by **Settings → Developer → Notification Log Retention** (default: 30 days). This affects historical notification and alert logs stored in Sencho's database, not the live container stream shown here.
|
||||
Log retention on the backend is controlled by **Settings · Developer · Notification Log Retention** (default thirty days). That setting governs alert and notification history stored in Sencho's database. It does not apply to the live container stream this page renders, which is bound entirely by Docker's own log driver and the in-browser buffers above.
|
||||
</Note>
|
||||
|
||||
@@ -3,10 +3,10 @@ title: Global Search
|
||||
description: Jump to any page, node, or stack from anywhere in the app with a single keystroke.
|
||||
---
|
||||
|
||||
The **global search palette** lets you move around Sencho without reaching for the mouse. It covers top-level navigation, every configured node, and every stack on every online node in your fleet.
|
||||
The **global search palette** lets you move around Sencho without reaching for the mouse. It covers the destinations the top bar exposes for your tier and role, every configured node, and every stack on every online node in your fleet.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/global-search/palette-stacks.png" alt="Sencho global search palette filtering stacks across nodes" />
|
||||
<img src="/images/global-search/palette-pages.png" alt="Sencho global search palette open with no query, the Pages group listing Home, Fleet, Resources, App Store, Logs, and Auto-Update each with a leading icon." />
|
||||
</Frame>
|
||||
|
||||
## Opening the palette
|
||||
@@ -21,38 +21,48 @@ The shortcut works from the dashboard, editor, fleet view, resources, and every
|
||||
|
||||
## What you can find
|
||||
|
||||
The palette groups results into three sections:
|
||||
The palette groups results into three sections.
|
||||
|
||||
| Group | What it contains | What happens when you pick one |
|
||||
|-------|------------------|--------------------------------|
|
||||
| **Pages** | Every top-level navigation destination (Home, Fleet, Resources, App Store, Logs, Auto-Update, Console, Audit, Schedules) | Navigates to that page |
|
||||
| **Nodes** | Every node in your fleet, with a green dot for online and a grey dot for offline | Switches the active node without leaving the current page |
|
||||
| **Stacks** | Every compose stack on every online node, matched by filename | Switches to the stack's node and opens it in the editor |
|
||||
| **Pages** | The same set of destinations the top bar shows you. Home, Fleet, Resources, App Store, and Logs always appear; Auto-Update appears for admins on Skipper or higher; Console and Schedules appear for admins on Admiral; Audit appears for any role on Admiral with the audit permission. | Navigates to that page |
|
||||
| **Nodes** | Every node in your fleet, with a green dot for online and a grey dot for offline. The currently active node carries a small **ACTIVE** chip on the right. | Switches the active node without leaving the current page |
|
||||
| **Stacks** | Every compose stack on every online node, matched on the compose filename (extension included). | Switches to the stack's node and opens it in the editor |
|
||||
|
||||
Each stack row shows a status dot (running, exited, or unknown) on the left and the node name as a tag on the right. That way you always know which host a match lives on before you pick it.
|
||||
Each stack row shows a status dot on the left (green for running, grey for exited, faded grey for unknown) and the node name as a mono tag on the right. That way you always know which host a match lives on before you pick it.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/global-search/palette-nodes.png" alt="Global search palette scrolled to the Nodes group, with the active node showing a small ACTIVE chip on the right and a green online dot to the left of each node name." />
|
||||
</Frame>
|
||||
|
||||
## Typing to filter
|
||||
|
||||
Start typing and the results narrow in real time. Matching is substring-based, so you can type a fragment from anywhere in a stack's filename or a node's name. Examples:
|
||||
Start typing and the results narrow in real time. Pages and nodes filter instantly; stack search debounces for about 250 ms before fanning out across the fleet, so a fast typist sees the Stacks group fill in a beat after the rest. Matching is substring-based, so you can type a fragment from anywhere in a stack's filename or a node's name. Examples:
|
||||
|
||||
- `fleet` — jumps to the Fleet page.
|
||||
- `opsix` — selects the Opsix node and switches the active context to it.
|
||||
- `db` — lists every stack with "db" in its name, showing which node each one lives on.
|
||||
- `fleet` jumps to the Fleet page.
|
||||
- `local` selects the Local node and switches the active context to it.
|
||||
- `db` lists every stack with "db" in its filename, showing which node each one lives on.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/global-search/palette-stacks.png" alt="Global search palette with the query 'saelix' producing three Stacks rows, each with a green status dot on the left and a mono OPSIX node tag on the right." />
|
||||
</Frame>
|
||||
|
||||
When a query matches stacks on a remote node, selecting one switches the active node and opens the editor on that stack in a single action.
|
||||
|
||||
## Cross-node search
|
||||
|
||||
Stack search fans out across every online node in your fleet. Offline nodes are skipped so one unreachable host never slows the palette down. Results from the active node and remote nodes appear together, sorted by the order nodes were added.
|
||||
Stack search fans out across every online node in your fleet. Offline nodes are skipped so one unreachable host never slows the palette down. Results from the active node and remote nodes appear together, grouped in the order nodes appear in your fleet list.
|
||||
|
||||
The palette caps the Stacks group at 50 results. If you have more matches than that, keep typing to narrow the query and the extras will come into view.
|
||||
The Stacks group caps at 50 rows. When a query matches more than that, a `Showing first 50 of N` line renders below the list so you know there is more to find. Keep typing to narrow the query; the line disappears once fewer than 50 matches remain.
|
||||
|
||||
While the fleet is responding, the empty area reads `Searching...` until the first batch of stack hits arrives. If no group has a hit at all, the empty area reads `No results.`
|
||||
|
||||
## Keyboard navigation
|
||||
|
||||
Once the palette is open:
|
||||
|
||||
- <kbd>↑</kbd> / <kbd>↓</kbd> — move through the results.
|
||||
- <kbd>Enter</kbd> — pick the highlighted result.
|
||||
- <kbd>Esc</kbd> — close the palette.
|
||||
- <kbd>↑</kbd> / <kbd>↓</kbd> moves through the results.
|
||||
- <kbd>Enter</kbd> picks the highlighted result.
|
||||
- <kbd>Esc</kbd> closes the palette.
|
||||
|
||||
Offline nodes show as greyed-out and are not selectable.
|
||||
|
||||
@@ -6,18 +6,18 @@ description: An interactive terminal on your host OS, directly in the browser -
|
||||
The **Console** tab opens a full interactive terminal session on the machine running Sencho. It behaves exactly like an SSH session, but without needing an SSH server, client, or key management.
|
||||
|
||||
<Note>
|
||||
The Host Console requires a **Sencho Admiral** license. Community and Skipper users will not see the Console tab in the navigation bar. [Learn more about licensing](/features/licensing).
|
||||
The Host Console requires a **Sencho Admiral** license and the **admin** role. [Learn more about licensing](/features/licensing).
|
||||
</Note>
|
||||
|
||||
<Frame>
|
||||
<img src="/images/host-console/host-console-overview.png" alt="Host Console showing a PowerShell prompt in the COMPOSE_DIR working directory" />
|
||||
<img src="/images/host-console/host-console-overview.png" alt="Host Console showing a live bash session listing containers with the floating chip strip in the bottom right" />
|
||||
</Frame>
|
||||
|
||||
## What it is
|
||||
|
||||
The Host Console gives you a real terminal session on the Sencho host, streamed to your browser over a WebSocket. It supports:
|
||||
|
||||
- Full colour and cursor support
|
||||
- Full color and cursor support
|
||||
- Tab completion (provided by the host shell)
|
||||
- Automatic terminal resizing when you resize the browser window
|
||||
- Copy and paste
|
||||
@@ -25,53 +25,58 @@ The Host Console gives you a real terminal session on the Sencho host, streamed
|
||||
|
||||
## Opening the console
|
||||
|
||||
Click **Console** in the top navigation bar. The session starts immediately in the `COMPOSE_DIR` root. If a stack is selected in the sidebar, the terminal opens directly inside that stack's directory instead.
|
||||
Click **Console** in the top navigation bar. The session starts immediately in the `COMPOSE_DIR` root. If a stack is selected in the sidebar, the terminal opens directly inside that stack's directory instead, and a small back button appears in the masthead so you can return to the stack editor.
|
||||
|
||||
## Cockpit layout
|
||||
|
||||
The Console page is a vertical stack of two surfaces:
|
||||
The Console page is a vertical stack of two surfaces.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/host-console/host-console-masthead.png" alt="Masthead with cyan rail, Connected state, and shell, viewport, and session uptime metadata" />
|
||||
<img src="/images/host-console/host-console-masthead.png" alt="Masthead with brand-cyan accent on the left, Connected state in italic, and shell, viewport, and session uptime metadata on the right" />
|
||||
</Frame>
|
||||
|
||||
| Surface | What it shows |
|
||||
|---------|---------------|
|
||||
| **Masthead** | Connection state (`Connected`, `Reconnecting`, or `Disconnected`) with a live status dot that pulses cyan while the session is active. The kicker reads `HOST CONSOLE · {node}`. When the console was opened from a stack, a small `← {stack-name}` link sits beside the state word so you can return to the stack editor. |
|
||||
| **Metadata** | Right side of the masthead: the active **shell** (e.g. `BASH`), the current **viewport** (`{cols}×{rows}` as reported by the fit addon), and **session uptime** since the shell opened. |
|
||||
| **Terminal well** | The interactive xterm session, sized to fill the remaining height. |
|
||||
| **Chip strip** | Floating bottom-right controls: **Copy** (copies the current selection), **Clear** (wipes the visible scrollback without killing the shell), **Download** (exports the full scrollback as a timestamped `.txt`), and **Reconnect** (closes and reopens the WebSocket without leaving the page). |
|
||||
The **masthead** holds the live connection state. The state word recolors with the tone of the session:
|
||||
|
||||
- **Connected** in brand cyan with a pulsing dot whenever output is flowing
|
||||
- **Reconnecting** in amber when the WebSocket is being re-established
|
||||
- **Disconnected** in red when the session has ended
|
||||
|
||||
The kicker reads `HOST CONSOLE · {NODE}`, so it always tells you which node the session belongs to. The right side carries three metadata pills:
|
||||
|
||||
| Pill | Meaning |
|
||||
|------|---------|
|
||||
| **SHELL** | The pill reads `BASH`. In a Docker deployment, the spawned shell is bash, falling back to `sh` if bash is unavailable inside the image. |
|
||||
| **VIEWPORT** | Current terminal size in `{cols}×{rows}`, updated automatically when you resize the browser window. |
|
||||
| **SESSION** | Time elapsed since the shell was opened (`M:SS UP` while under an hour, then `1H 05M` once it crosses an hour). |
|
||||
|
||||
The **terminal well** sits below the masthead and fills the remaining height with an xterm.js session.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/host-console/host-console-chip-strip.png" alt="Floating chip strip with Copy, Clear, Download, and Reconnect buttons" />
|
||||
</Frame>
|
||||
|
||||
## Shell type
|
||||
The **chip strip** floats over the bottom right of the terminal:
|
||||
|
||||
The shell depends on the host operating system:
|
||||
- **Copy** copies the current xterm selection to your clipboard
|
||||
- **Clear** wipes the visible scrollback without killing the shell
|
||||
- **Download** exports the full scrollback as a timestamped `.txt` file
|
||||
- **Reconnect** closes and reopens the WebSocket without leaving the page
|
||||
|
||||
- **Linux / macOS hosts:** `bash` (falls back to `sh` if bash is not available)
|
||||
- **Windows hosts:** PowerShell
|
||||
## Sessions and limits
|
||||
|
||||
In a typical Docker deployment, the Sencho container runs Linux, so the console opens a `bash` shell regardless of the Docker host's OS. The PowerShell prompt shown in the screenshot above reflects a native Windows development environment.
|
||||
Sencho enforces a maximum of **5 concurrent console sessions** per instance. If you hit this limit, close an existing session before opening a new one. Each session runs a heartbeat; if the browser tab is closed or the network drops, the server detects the dead connection and cleans up the orphaned shell within 60 to 90 seconds.
|
||||
|
||||
## Availability
|
||||
## Environment variables
|
||||
|
||||
The Host Console is available exclusively to users with the **admin** role on the **Admiral** tier. Other roles (node-admin, deployer, viewer, auditor) cannot access the console, even on Admiral-licensed instances. If your instance is on the Community or Skipper tier, the Console tab does not appear in the navigation bar at all. Attempting to access the console endpoint directly without the correct license or role is rejected.
|
||||
|
||||
## Session limits
|
||||
|
||||
Sencho enforces a maximum of **5 concurrent console sessions** per instance. If you hit this limit, close an existing session before opening a new one. Each session also includes an automatic heartbeat; if the browser tab is closed or the network drops, the server detects the dead connection and cleans up the session within about a minute.
|
||||
Environment variables whose names suggest secrets (passwords, tokens, keys, credentials, auth-related variables, private keys, passphrases, encryption and signing secrets) and certain well-known connection strings (`DATABASE_URL`, `REDIS_URL`, `MONGO_URI`, `AMQP_URL`, `DSN`) are stripped from the console environment before the shell spawns. Running `env` or `printenv` inside the console will not reveal them.
|
||||
|
||||
## Security considerations
|
||||
|
||||
The Host Console has stricter access requirements than other features:
|
||||
The Host Console is one of the most powerful features in Sencho and is treated as such:
|
||||
|
||||
- **Admin role required.** Only users with the `admin` role can open a console session. Non-admin roles are blocked at both the UI and the API level.
|
||||
- **Browser sessions only.** API tokens used for automation or multi-node communication cannot open a console session.
|
||||
- **Admiral license required.** The license check is enforced at every layer, not just in the UI.
|
||||
- **Session invalidation.** Changing a user's password or role immediately invalidates any active session tokens. A user who is downgraded from admin loses console access instantly.
|
||||
- **Sensitive environment variables are stripped.** Environment variables whose names suggest secrets (passwords, tokens, keys, credentials, auth-related variables, private keys, passphrases, encryption and signing secrets) and certain well-known connection strings are automatically removed from the console environment. Running `env` or `printenv` inside the console will not reveal them.
|
||||
- **Admin role required.** Only users with the **admin** role can open a console session.
|
||||
- **Admiral license required.** The Console tab is not available on the Community or Skipper tiers.
|
||||
- **Browser sessions only.** Console sessions are only available from a signed-in browser session, not from API tokens.
|
||||
|
||||
<Warning>
|
||||
The Host Console provides unrestricted shell access to the machine running Sencho. Do not expose Sencho on a public network without HTTPS and strong authentication.
|
||||
@@ -88,18 +93,18 @@ The Host Console has stricter access requirements than other features:
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Maximum console sessions reached">
|
||||
Sencho limits concurrent console sessions to 5 per instance. Close one or more existing console sessions (either through the UI or by closing the browser tab) and try again. Stale sessions from disconnected browsers are cleaned up automatically within about a minute.
|
||||
Sencho limits concurrent console sessions to 5 per instance. Close one or more existing console sessions (either through the UI or by closing the browser tab) and try again. Stale sessions from disconnected browsers are cleaned up automatically within 60 to 90 seconds.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Session ends immediately after connecting">
|
||||
This typically means the shell could not be found on the host. In a Docker deployment, ensure `bash` or `sh` is available inside the container. You can check by running `docker exec <container> which bash` from the host. Also verify that your user has the `admin` role and your instance has an active Admiral license.
|
||||
This typically means the shell could not be found on the host. In a Docker deployment, ensure `bash` or `sh` is available inside the container. You can check by running `docker exec <container> which bash` from the host. Also verify that your user has the **admin** role and your instance has an active Admiral license.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Connection drops after a short time">
|
||||
If you are running Sencho behind a reverse proxy (e.g. Nginx, Caddy, Traefik), the proxy may be closing WebSocket connections after its default idle timeout. Increase the WebSocket timeout in your proxy configuration. For Nginx, set `proxy_read_timeout` and `proxy_send_timeout` to a higher value (e.g. `3600s`).
|
||||
If you are running Sencho behind a reverse proxy (Nginx, Caddy, Traefik), the proxy may be closing WebSocket connections after its default idle timeout. Increase the WebSocket timeout in your proxy configuration. For Nginx, set `proxy_read_timeout` and `proxy_send_timeout` to a higher value (for example `3600s`).
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Console tab is not visible">
|
||||
The Console tab only appears for users with the `admin` role on an Admiral-licensed instance. Verify both conditions are met. If you recently upgraded your license, you may need to refresh the page.
|
||||
The Console tab only appears for users with the **admin** role on an Admiral-licensed instance. Verify both conditions are met. If you recently activated an Admiral license, refresh the page.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
@@ -11,112 +11,158 @@ Sencho uses an open-core model. The **Community** tier is free forever with unli
|
||||
|
||||
## Plans
|
||||
|
||||
| Tier | Price | Accounts |
|
||||
|------|-------|----------|
|
||||
| **Community** | Free | 1 admin |
|
||||
| **Skipper** | $5.99/month billed annually, $7.99/month, or $249 lifetime | 1 admin + 3 viewers |
|
||||
| **Admiral** | $41.99/month billed annually, $49.99/month, or $1,499 lifetime | Unlimited |
|
||||
| Tier | Annual (per mo, billed yearly) | Monthly | Lifetime (EA) | Seats |
|
||||
|------|--------------------------------|---------|---------------|-------|
|
||||
| **Community** | Free | Free | Free | 1 admin |
|
||||
| **Skipper** | $11.99 | $14.99 | $449 | 1 admin + 3 viewers |
|
||||
| **Admiral** | $69.99 | $89.99 | $2,499 | Unlimited |
|
||||
|
||||
Lifetime pricing is an early-adopter offer available for a limited time only.
|
||||
For larger deployments, an **Enterprise** tier is available from $3,500 per year with SLA, priority support, security questionnaires, and custom contracts. See [the pricing page](https://sencho.io/pricing) for the full comparison.
|
||||
|
||||
Lifetime pricing is an Early Access offer available for a limited time.
|
||||
|
||||
### Feature breakdown
|
||||
|
||||
**Community** includes:
|
||||
- Unlimited nodes, compose editor, global logs, app store, alerts, and more
|
||||
- Fleet View with search, sort, filters, node-card expand, and topology
|
||||
- Stack labels
|
||||
- Network topology graph
|
||||
- Manual fleet snapshots (create, browse, restore, delete)
|
||||
- Per-node Sencho updates and the Check Updates view
|
||||
- Vulnerability scanning: install/update/uninstall Trivy, on-demand scans (vulnerabilities, secrets, misconfigurations), scan comparison, and CVE suppressions
|
||||
- Two-factor authentication (TOTP)
|
||||
- Custom OIDC single sign-on (works with Authelia, Keycloak, Authentik, Zitadel, Pocket ID, or any spec-compliant OIDC identity provider)
|
||||
|
||||
- Unlimited nodes, the Monaco compose editor, and the App Store with 199+ one-click templates
|
||||
- Real-time container stats, global logs, the interactive network topology graph, and stack labels
|
||||
- Git sources for compose stacks
|
||||
- Multi-node management in both Proxy and Pilot Agent modes
|
||||
- Manual fleet snapshots (create, browse, restore, delete) and per-node Sencho updates
|
||||
- Vulnerability scanning: install, update, and uninstall Trivy, on-demand scans for vulnerabilities, secrets, and misconfigurations, plus scan comparison
|
||||
- CVE suppressions
|
||||
- Alert rules with Discord, Slack, and webhook targets
|
||||
- Two-factor authentication (TOTP plus backup codes)
|
||||
- Custom OIDC single sign-on (works with Authelia, Keycloak, Authentik, Zitadel, Pocket ID, or any spec-compliant OIDC provider)
|
||||
|
||||
**Skipper** includes everything in Community, plus:
|
||||
- Webhooks
|
||||
- Bulk actions on a label (deploy, stop, or restart every stack tagged with it)
|
||||
- Atomic deployments
|
||||
- Bulk **Update All** across the fleet, scheduled scans, scheduled updates, and scheduled fleet snapshots
|
||||
- Scan policies with `block_on_deploy` enforcement, SBOM (SPDX, CycloneDX), and SARIF export
|
||||
|
||||
- Fleet View with search, sort, filter, and node-card drill-down
|
||||
- Webhooks (incoming, to trigger deploys from CI/CD)
|
||||
- Atomic deployments with rollback
|
||||
- Auto-update policies for stack images
|
||||
- One-click Google, GitHub, and Okta SSO presets
|
||||
- Auto-heal policies
|
||||
- Scheduled tasks for stack updates, vulnerability scans, and fleet snapshots
|
||||
- Scan policies with `block_on_deploy` deploy enforcement, SBOM (SPDX, CycloneDX), and SARIF export
|
||||
- Bulk actions on a label (deploy, stop, or restart every stack tagged with it)
|
||||
- Remote OTA node updates from the control instance
|
||||
- Fleet Actions (bulk update all, bulk restart Sencho)
|
||||
- Blueprints and Fleet Secrets
|
||||
- Preset SSO for Google, GitHub, and Okta
|
||||
- Viewer accounts (1 admin plus 3 viewers)
|
||||
|
||||
**Admiral** includes everything in Skipper, plus:
|
||||
- Unlimited admin and viewer accounts
|
||||
- Scoped RBAC (deployer, node-admin, auditor roles)
|
||||
- LDAP / Active Directory authentication
|
||||
- Audit log and host console
|
||||
- API tokens and private registries
|
||||
- Notification routing
|
||||
- Auto-update of the managed Trivy binary
|
||||
- All other scheduled operations (restart, prune, etc.)
|
||||
|
||||
<Tip>
|
||||
**SSO is available on every tier.** Community users can integrate any OIDC-compliant identity provider through the Custom OIDC option. Paid tiers add turnkey presets (Google, GitHub, Okta) and LDAP / Active Directory.
|
||||
</Tip>
|
||||
- Unlimited admin and viewer accounts with the full role set (deployer, node-admin, auditor)
|
||||
- LDAP / Active Directory authentication
|
||||
- Audit log with CSV export
|
||||
- Host Console (a browser-based terminal on the Sencho host)
|
||||
- API tokens for CI/CD scripts
|
||||
- Private and custom registry credentials
|
||||
- Notification routing
|
||||
- Sencho Mesh (cross-node container networking)
|
||||
- Sencho Cloud Backup
|
||||
- Auto-update of the managed Trivy binary
|
||||
- All other scheduled task actions (restart, prune, and more)
|
||||
|
||||
## Free trial
|
||||
|
||||
Sencho offers a **14-day Admiral trial** on monthly and annual Admiral plans so you can evaluate the flagship features (Host Console, Scheduled Operations, LDAP / Active Directory, audit log, unlimited accounts) with your real infrastructure before committing.
|
||||
Sencho offers a **14-day Admiral trial** so you can evaluate the flagship features (Host Console, Sencho Mesh, scheduled operations, LDAP / Active Directory, audit log, unlimited accounts) with your real infrastructure before committing. The trial is offered on the monthly and annual Admiral plans; the lifetime plan does not include a trial.
|
||||
|
||||
To start a trial:
|
||||
|
||||
1. In the Sencho dashboard, go to **Settings > License**, or visit the [pricing page](https://sencho.io/pricing).
|
||||
2. Click **Start monthly trial** or **Start annual trial**.
|
||||
3. Complete the Lemon Squeezy checkout. A valid card is required for verification; you are not charged until the trial ends.
|
||||
4. Lemon Squeezy will email your license key within a few minutes.
|
||||
5. Paste the key into the **Have a license key?** field on the License page and click **Activate**.
|
||||
1. Visit the [pricing page](https://sencho.io/pricing).
|
||||
2. Switch to the **Annual** or **Monthly** billing tab.
|
||||
3. Click **Start 14-day trial** on the **Admiral** card.
|
||||
4. Complete the checkout. A valid card is required for verification; you are not charged until the trial ends.
|
||||
5. Your license key is emailed to you within a few minutes.
|
||||
6. Activate the key in the Sencho dashboard as described in [Activating your license](#activating-your-license).
|
||||
|
||||
When your trial ends, Lemon Squeezy automatically charges the card you provided and your plan continues as a paid Admiral subscription. To avoid being charged, cancel from the **Manage Subscription** button (or from the Lemon Squeezy receipt email) any time before day 14.
|
||||
When your trial ends, the card you provided is automatically charged and your plan continues as a paid Admiral subscription. To avoid being charged, cancel from the **Manage subscription** button in **Settings → License** (or from your receipt email) any time before day 14.
|
||||
|
||||
<Tip>
|
||||
Fresh installs land on the **Community** tier until you activate a license key. All Community features (unlimited nodes, compose editor, global logs, alerts, Custom OIDC, and more) are available immediately.
|
||||
</Tip>
|
||||
|
||||
## Upgrading your plan
|
||||
|
||||
You can upgrade directly from **Settings > License** in your Sencho dashboard. The upgrade cards shown depend on your current tier:
|
||||
|
||||
- **Community users** see both **Skipper** and **Admiral** options with feature highlights and direct checkout links.
|
||||
- **Skipper users** see an **Admiral** upgrade card for when you need unlimited accounts and team features.
|
||||
- **Admiral users** are on the highest tier, so no upgrade cards are shown.
|
||||
|
||||
Clicking an upgrade button opens the Lemon Squeezy checkout in a new tab. After completing the purchase, you will receive a license key by email.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/licensing/license-community.png" alt="License settings showing Skipper and Admiral upgrade cards for Community users" />
|
||||
</Frame>
|
||||
|
||||
## Activating your license
|
||||
|
||||
1. Go to **Settings > License** in your Sencho dashboard.
|
||||
2. Paste your key into the **License Key** field at the bottom of the page.
|
||||
3. Click **Activate**.
|
||||
Open **Settings → License** in the Sencho dashboard. When the license is not active, the **Activate** section sits below the Plan card:
|
||||
|
||||
Sencho validates the key against Lemon Squeezy and unlocks your tier's features immediately.
|
||||
<Frame>
|
||||
<img src="/images/licensing/license-activate-section.png" alt="The Activate section of the License page on a Community-tier instance, with a License key input field and an Activate button on the right" />
|
||||
</Frame>
|
||||
|
||||
## License validation
|
||||
1. Paste your key into the **License key** field.
|
||||
2. Click **Activate**.
|
||||
|
||||
Active licenses are re-validated every **72 hours** against the Lemon Squeezy API. If your instance goes offline, there is a **30-day grace period** before it degrades to the Community tier.
|
||||
Sencho validates the key and unlocks your tier. If activation fails, the toast surfaces the verbatim error; common causes are an invalid key, a typo, or an activation limit reached on a previous instance.
|
||||
|
||||
## The Plan section
|
||||
|
||||
When a license is active, **Settings → License** opens on the **Plan** section. The page masthead at the top exposes three stat pills, and the section below lists the metadata for the active license.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/licensing/license-admiral-active.png" alt="License page showing an Admiral lifetime plan with the masthead stat strip listing SCOPE, PLAN, and DURATION, then the Plan section with the Customer, Product, and masked License key fields and the Deactivate button" />
|
||||
</Frame>
|
||||
|
||||
The masthead pills are:
|
||||
|
||||
| Pill | Meaning |
|
||||
|------|---------|
|
||||
| **SCOPE** | Reads `operator` when you are signed in as an admin. |
|
||||
| **PLAN** | The current tier: `community`, `skipper`, or `admiral`. Trial licenses show the trial tier (typically `admiral`). |
|
||||
| **DURATION** / **RENEWS** / **TRIAL** / **STATUS** | `DURATION: lifetime` for lifetime licenses, `RENEWS: <date>` for active subscriptions, `TRIAL: Xd left` for trials, and `STATUS: expired` for expired licenses. |
|
||||
|
||||
The **Plan** card lists:
|
||||
|
||||
- **Tier name** (e.g. `Sencho Admiral`) with a short status line, such as "Active license on this control plane", "Trial: X days remaining", or "Your license has expired."
|
||||
- **Customer**: the customer name on the purchase.
|
||||
- **Product**: the purchased product variant (e.g. `Sencho Admiral`).
|
||||
- **License key**: the last four characters of your key, displayed as `****-****-****-XXXX`. The full key is never re-displayed after activation.
|
||||
|
||||
## Managing your subscription
|
||||
|
||||
Once your license is active, you can manage your subscription in two ways:
|
||||
When the license is an active subscription (not lifetime), the Plan section action row exposes **Manage subscription** alongside **Deactivate**. **Manage subscription** opens the billing portal in a new tab, where you can update your payment method, view invoices, cancel, or switch plans.
|
||||
|
||||
- **From the License page:** Go to **Settings > License** and click **Manage Subscription**. This opens the Lemon Squeezy billing portal where you can update your payment method, view invoices, cancel, or switch plans.
|
||||
- **From the profile menu:** Click your profile icon in the top bar and select **Billing** for quick access to the same portal. This option is hidden for lifetime licenses.
|
||||
The same portal is reachable from the profile menu in the top-right corner of the app:
|
||||
|
||||
<Frame>
|
||||
<img src="/images/licensing/license-admiral-active.png" alt="License settings showing an active Admiral subscription with Manage Subscription button" />
|
||||
<img src="/images/licensing/profile-menu.png" alt="Profile dropdown popover showing an identity header with the admin role and Admiral tier badges, a navigation strip with Settings, Documentation, and Feedback entries, an Appearance theme picker, and a Log Out button" />
|
||||
</Frame>
|
||||
|
||||
Click your initials in the top-right corner to open the popover. For active subscription licenses, a **Billing** row appears between **Settings** and **Documentation** that opens the same portal as **Manage subscription**. Lifetime licenses have no recurring subscription to manage, so the **Billing** row is hidden.
|
||||
|
||||
### Lifetime licenses
|
||||
|
||||
Lifetime licenses do not have a recurring subscription, so the **Manage Subscription** button is not shown. Instead, the license card displays **Duration: Lifetime** to confirm your license never expires. All other features (deactivation, multi-node enforcement, validation) work the same as subscription licenses.
|
||||
Lifetime licenses have no recurring subscription, so the **Manage subscription** button and the profile menu **Billing** row are both hidden, and the masthead shows `DURATION: lifetime` instead of a renewal date. Deactivation, multi-node enforcement, and periodic validation work the same as for subscription licenses.
|
||||
|
||||
## License validation
|
||||
|
||||
Active licenses are re-validated every **72 hours**. If your instance goes offline, there is a **30-day grace period** before it degrades to the Community tier.
|
||||
|
||||
## License states
|
||||
|
||||
The License page renders differently depending on the license status:
|
||||
|
||||
| Status | Plan section | Activate section | Pricing section |
|
||||
|--------|--------------|------------------|-----------------|
|
||||
| **Community** | `Sencho Community` with "Free tier with the core experience." | Visible | Visible |
|
||||
| **Trial** | `Sencho Admiral (Trial)` with a countdown chip | Visible (so a paid key can replace the trial) | Hidden |
|
||||
| **Active subscription** | Tier name with Customer, Product, License key, plus `Manage subscription` and `Deactivate` | Hidden | Hidden |
|
||||
| **Active lifetime** | Tier name with Customer, Product, License key, plus `Deactivate` only | Hidden | Hidden |
|
||||
| **Expired** | `Sencho Community` with "Your license has expired. Renew to restore paid features." and a destructive **Status: Expired** field | Visible | Visible |
|
||||
| **Disabled** | `Sencho Community` with "Your license has been disabled. Contact support for assistance." | Visible | Visible |
|
||||
|
||||
The **Pricing** section, when visible, holds a single **See pricing** button that opens [the pricing page](https://sencho.io/pricing) in a new tab so you can pick a tier and billing cadence.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/licensing/license-community.png" alt="License page on a Community-tier instance, with the masthead showing SCOPE operator and PLAN community, then the Plan section reading 'Sencho Community · Free tier with the core experience.', the Activate section with the License key input, and the Pricing section with the See pricing button" />
|
||||
</Frame>
|
||||
|
||||
## Multi-node license enforcement
|
||||
|
||||
If you manage multiple servers through Sencho's [multi-node feature](/features/multi-node), you only need a license on your **primary instance**. All remote nodes automatically inherit the primary's license tier for proxied requests. No per-node activation is required.
|
||||
If you manage multiple servers through Sencho's [multi-node feature](/features/multi-node), you only need a license on your **control instance**. All remote nodes automatically inherit the control instance's license tier for proxied requests. No per-node activation is required.
|
||||
|
||||
For details on how this works, see [License enforcement across nodes](/features/multi-node#license-enforcement-across-nodes).
|
||||
|
||||
@@ -124,9 +170,9 @@ For details on how this works, see [License enforcement across nodes](/features/
|
||||
|
||||
To transfer your license to a different instance:
|
||||
|
||||
1. Go to **Settings > License**.
|
||||
2. Click **Deactivate License**.
|
||||
3. On your new instance, activate the same license key.
|
||||
1. Go to **Settings → License**.
|
||||
2. Click **Deactivate** in the Plan section.
|
||||
3. On your new instance, paste the same license key into **License key** and click **Activate**.
|
||||
|
||||
Deactivation reverts the current instance to the Community tier immediately.
|
||||
|
||||
|
||||
@@ -1,71 +1,140 @@
|
||||
---
|
||||
title: Multi-Node Management
|
||||
description: Connect multiple Sencho instances and manage all your servers from a single dashboard.
|
||||
description: Connect multiple Sencho instances and manage every server from one console with no central server, no SSH, and no shared Docker sockets.
|
||||
---
|
||||
|
||||
Sencho's multi-node feature lets you manage Docker Compose stacks on multiple servers, all from the same browser tab. Each server runs its own Sencho instance, and your primary instance acts as a transparent proxy to the others.
|
||||
Sencho's multi-node feature lets you operate Docker Compose stacks on every server you run from a single browser tab. Each server runs its own Sencho instance and the control instance acts as a transparent proxy to the others.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/multi-node/node-manager.png" alt="Node Manager showing a local and a remote node, both Online" />
|
||||
<img src="/images/multi-node/node-manager.png" alt="Settings · System · Nodes panel showing the masthead breadcrumb, SCOPE/NODES/REMOTE stat strip, Add node button, Generate Node Token card, and a populated table of one Local row plus seven remote rows with mixed Proxy and Pilot Agent modes." />
|
||||
</Frame>
|
||||
|
||||
## How it works
|
||||
|
||||
There is no central server. Each Sencho instance manages its own host independently. When you select a remote node, your browser's API calls are proxied through your local Sencho instance to the remote one, authenticated by a long-lived Bearer token. No SSH. No shared Docker sockets.
|
||||
There is no central server. Each Sencho instance manages its own host independently. When you select a remote node from the switcher, your browser keeps talking to your control instance; the control instance forwards the call to the remote one, authenticated by a per-node bearer token in Distributed API Proxy mode or by the enrolled outbound tunnel session in Pilot Agent mode. No SSH. No shared Docker sockets. No remote Docker TCP exposure.
|
||||
|
||||
Remote nodes connect in one of two modes. Pick the one that matches your network topology before you add the node, because the add-node flow branches on the choice.
|
||||
|
||||
## The local node
|
||||
|
||||
Your primary Sencho installation is always listed as **Local**. It is the default node, marked with a star icon, and cannot be deleted. All operations on the local node run directly against the host's Docker socket.
|
||||
Your control Sencho instance is always listed as **Local**. It is the default node, marked with a star icon, and cannot be deleted. All operations on the local node run directly against the host's Docker socket.
|
||||
|
||||
## Adding a remote node
|
||||
## Choose a remote mode
|
||||
|
||||
### Step 1: Generate a token on the remote machine
|
||||
When you add a remote node, the Mode picker offers two options. The form tells you the difference inline:
|
||||
|
||||
On the **remote** Sencho instance (the server you want to add), open **Settings → Nodes**. You will see a **Generate Node Token** section at the top of the page. Click **Generate Token**, then copy the token that appears. You will only see it once.
|
||||
> Pilot Agent requires only outbound HTTPS from the remote host. Distributed API Proxy requires the remote host to expose an inbound port.
|
||||
|
||||
<Note>
|
||||
The token is a long-lived credential that grants full control over that Sencho instance. Treat it like a password.
|
||||
</Note>
|
||||
| | Pilot Agent (default) | Distributed API Proxy |
|
||||
|---|---|---|
|
||||
| **Direction** | Remote dials the control plane (outbound only) | Control plane dials the remote (inbound to remote) |
|
||||
| **Network requirement** | Outbound HTTPS to the control instance | Inbound TCP port reachable from the control instance |
|
||||
| **What you run on the remote host** | One `docker run` command issued at enrollment | A full Sencho install reachable on a routable URL |
|
||||
| **Token model** | One-shot enrollment token, 15-minute expiry | Long-lived bearer token, rotates on regeneration |
|
||||
| **Best for** | Hosts behind NAT, dynamic IPs, restrictive ingress, edge boxes | Servers you already expose on a stable URL with TLS termination |
|
||||
|
||||
### Step 2: Add the node on your primary instance
|
||||
If you are not sure, start with **Pilot Agent**. It is the default because it works in more network shapes and the enrollment is a single command.
|
||||
|
||||
On your **primary** Sencho instance, open **Settings → Nodes** and click **+ Add Node**. Fill in:
|
||||
For a deep look at how the pilot tunnel works under the hood (credential lifecycle, security model, resource limits, environment variables, and pilot-specific troubleshooting), see [Pilot Agent](/features/pilot-agent).
|
||||
|
||||
## Add a remote node: Pilot Agent
|
||||
|
||||
### Step 1. Open the Add node form
|
||||
|
||||
On the control instance, click your avatar in the top-right and choose **Settings**. In the sidebar pick **System → Nodes**, then click **Add node**.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/multi-node/add-node-form.png" alt="Add Node form with Name, Type, API URL, Token, and Compose Directory fields" />
|
||||
<img src="/images/multi-node/add-node-pilot.png" alt="Add remote node modal with Type set to Remote and Mode set to Pilot Agent. The form shows Name, Type, Mode, Mode helper text, and Compose Directory fields. URL and token fields are hidden because Pilot Agent does not need them." />
|
||||
</Frame>
|
||||
|
||||
| Field | Description |
|
||||
|-------|-------------|
|
||||
| **Name** | A display name (e.g. `Production VPS`, `media-box`) |
|
||||
| **Type** | Select **Remote** (or **Local** for an additional local Docker socket) |
|
||||
| **Sencho API URL** | The full HTTP/HTTPS URL of the remote instance (e.g. `http://192.168.1.50:1852`) |
|
||||
| **API Token** | The token you generated in Step 1 |
|
||||
| **Compose Directory** | The root directory where compose stack folders live on the remote node (defaults to `/app/compose`) |
|
||||
| Field | Notes |
|
||||
|-------|-------|
|
||||
| **Name** | A display name (e.g. `staging-vps`, `media-box`). Visible in the switcher and across the UI. |
|
||||
| **Type** | Set to **Remote**. |
|
||||
| **Mode** | Set to **Pilot Agent**. |
|
||||
| **Compose Directory** | The root directory where compose stack folders live on the remote host. Defaults to `/app/compose`; change it if the remote uses a different volume mount. |
|
||||
|
||||
Click **Add Node**. Sencho immediately tests the connection and shows the result.
|
||||
Click **Add node**.
|
||||
|
||||
### Step 3: Verify connectivity
|
||||
### Step 2. Run the enrollment command on the remote host
|
||||
|
||||
A successful connection shows the remote node as **Online** with a green badge. If it shows **Offline** or **Unknown**, check:
|
||||
- The remote Sencho instance is running and reachable from your primary host
|
||||
- The API URL is correct (include the port if non-standard)
|
||||
- The token was copied correctly without extra whitespace
|
||||
A modal opens with a single `docker run` command. Copy it and run it on the remote host as the user that owns Docker.
|
||||
|
||||
Click the **wifi icon** (Test Connection) on any node row at any time to re-check status. A successful test shows a **Connection Details** panel below the table with information about the remote instance, including its OS, architecture, container count, and Sencho version.
|
||||
<Frame>
|
||||
<img src="/images/multi-node/pilot-enroll.png" alt="Enroll the pilot agent modal showing a docker run command with SENCHO_MODE=pilot, SENCHO_PRIMARY_URL pointing to the control instance, and a SENCHO_ENROLL_TOKEN. The footer reads Expires 15m from now and offers a Copy command button." />
|
||||
</Frame>
|
||||
|
||||
The command is one line. It pulls the `saelix/sencho:latest` image, mounts the host Docker socket, and starts a container named `sencho-agent` that opens an outbound tunnel back to the control plane.
|
||||
|
||||
The enrollment token in the command is **valid for 15 minutes and can only be used once**. If you wait too long, follow the regeneration steps below.
|
||||
|
||||
### Step 3. Verify the tunnel
|
||||
|
||||
Back on the Nodes table, the new row shows **Status: Online** once the agent dials home, and the Endpoint column reads `tunnel (seen Xs ago)`. Until the agent connects for the first time, the Endpoint reads `tunnel (waiting)`.
|
||||
|
||||
### Re-enrollment
|
||||
|
||||
If the enrollment token expires, the container is removed, or the host is rebuilt, click the **Edit Node** icon on the pilot row. The Edit modal shows a **Regenerate enrollment token** card.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/multi-node/edit-pilot-regenerate.png" alt="Edit node modal for a Pilot Agent node. Below the Compose Directory field a bordered card explains: 'Re-enroll the agent if the container was lost or the enrollment token expired. The previous tunnel is disconnected automatically.' A Regenerate enrollment token button sits below the explanation." />
|
||||
</Frame>
|
||||
|
||||
Click the button to mint a fresh enrollment command. The previous tunnel is disconnected automatically; run the new command on the remote host to reconnect.
|
||||
|
||||
## Add a remote node: Distributed API Proxy
|
||||
|
||||
### Step 1. Generate a long-lived token on the remote
|
||||
|
||||
On the **remote** Sencho instance (the server you want to add), open **Settings → System → Nodes**. Click **Generate Token** in the **Generate Node Token** card and copy the token that appears.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/multi-node/generate-token.png" alt="Generate Node Token card showing the explanation text and a freshly issued JWT-style token displayed in a monospace block with a copy button." />
|
||||
</Frame>
|
||||
|
||||
<Note>
|
||||
The token is a long-lived credential that grants full control over that Sencho instance. Treat it like a password. Generating a new token invalidates the previous one.
|
||||
</Note>
|
||||
|
||||
### Step 2. Add the node on your control instance
|
||||
|
||||
Open **Settings → System → Nodes** on the control instance and click **Add node**. Set Type to **Remote** and Mode to **Distributed API Proxy** to reveal the URL and token fields.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/multi-node/add-node-proxy.png" alt="Add remote node modal in Distributed API Proxy mode. The form shows Name, Type, Mode, Sencho API URL with a plain-HTTP warning banner underneath, API Token (masked), and Compose Directory. The HTTP warning recommends HTTPS or a VPN when the node is reachable over the public internet." />
|
||||
</Frame>
|
||||
|
||||
| Field | Notes |
|
||||
|-------|-------|
|
||||
| **Name** | A display name shown in the switcher. |
|
||||
| **Type** | **Remote**. |
|
||||
| **Mode** | **Distributed API Proxy**. |
|
||||
| **Sencho API URL** | The full URL of the remote Sencho instance, including scheme and port (e.g. `https://sencho.example.com` or `http://10.0.1.20:1852`). |
|
||||
| **API Token** | The token you copied in Step 1. |
|
||||
| **Compose Directory** | The root directory where compose stack folders live on the remote host. |
|
||||
|
||||
If the URL starts with plain `http://`, the form shows an inline warning. Plain HTTP is fine on a private network or VPN, but it leaks the bearer token over any path you do not control. See [Transport encryption](#transport-encryption) below.
|
||||
|
||||
Click **Add node**. Sencho immediately tests the connection and shows the result.
|
||||
|
||||
### Step 3. Verify connectivity
|
||||
|
||||
Click the **Test Connection** icon on any row to re-check status. A successful test shows a Connection Details panel below the table with the remote's OS, architecture, container count, and Sencho version.
|
||||
|
||||
If the row shows **Offline** or **Unknown**, see [Troubleshooting](#troubleshooting).
|
||||
|
||||
## Switching between nodes
|
||||
|
||||
The **node switcher** sits at the top of the sidebar and acts as a persistent identity anchor. It shows the active node's status dot, type (`LOCAL`, `REMOTE`, or `AGENT`), and name, so you always know which node you are viewing.
|
||||
The **node switcher** sits at the top of the sidebar and acts as a persistent identity anchor. It shows the active node's status dot, a `Node · LOCAL`, `Node · REMOTE`, or `Node · AGENT` kicker, and the node name, so you always know which node you are operating on.
|
||||
|
||||
When two or more nodes are registered, clicking the switcher opens a popover listing every connected node. Each row includes a status dot, name, and a metadata line with the node type, version, and (for pilot agents) the time since the agent last checked in. The active node is marked with a cyan accent rail, and a filled star calls out the default node.
|
||||
|
||||
Click any row to switch. All views (dashboard stats, stack list, editor, resources, logs) immediately reflect the selected node. A **Manage nodes** link at the bottom of the popover opens the Nodes settings page to add, edit, or remove nodes.
|
||||
When two or more nodes are registered, clicking the switcher opens a popover listing every connected node. Each row shows a status dot, the name, and a metadata line that joins the type chip, the last-seen timestamp (pilot agents only), and the node version with middle-dot separators (for example, `AGENT · SEEN 13H AGO · v0.76.3`). The active node is marked with a brand-coloured rail, and a filled star calls out the default node.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/multi-node/node-switcher-dropdown.png" alt="Node switcher popover listing connected nodes with type, version, and active-row accent" />
|
||||
<img src="/images/multi-node/node-switcher-dropdown.png" alt="Node switcher popover with a Connected header above an 8 NODES count. The Local row shows LOCAL · v0.76.3 with a star icon. Three Remote rows show the REMOTE chip. The sencho-pilot-test row shows AGENT · SEEN 13H AGO. A Manage nodes link sits at the bottom." />
|
||||
</Frame>
|
||||
|
||||
Click any row to switch. All views (dashboard, stack list, editor, resources, logs) immediately reflect the selected node. The **Manage nodes** link at the bottom of the popover opens the Settings → Nodes panel to add, edit, or remove nodes.
|
||||
|
||||
## What top-level views show when a remote node is active
|
||||
|
||||
Top-level views that manage fleet-wide state are scoped to the local hub and are hidden from the nav strip when a remote node is the active selection. They reappear automatically when you switch back to your local node.
|
||||
@@ -80,73 +149,81 @@ Top-level views that manage fleet-wide state are scoped to the local hub and are
|
||||
|
||||
Node-level views (Home, Resources, App Store, Console, the editor) work as expected against whichever node you have selected.
|
||||
|
||||
## Per-node scheduling and update indicators
|
||||
## The Nodes table
|
||||
|
||||
The Nodes table surfaces scheduling and update status at a glance for each node:
|
||||
The Nodes table surfaces routing, status, and per-node automation at a glance for every node:
|
||||
|
||||
<Frame>
|
||||
<img src="/images/per-node-scheduling/nodes-table-overview.png" alt="Nodes table showing Schedules and Updates columns for each node" />
|
||||
<img src="/images/multi-node/nodes-table-overview.png" alt="Nodes table close-up showing the ten columns: a default-marker column with a star on the Local row, Name, Type, Mode, Endpoint, Status, Labels, Schedules, Updates, and Actions. The pilot-test row has Mode 'Pilot Agent' and Endpoint 'tunnel (seen 13h ago)'. Other remotes show Mode 'Proxy' and full URLs in the Endpoint column." />
|
||||
</Frame>
|
||||
|
||||
| Column | What it shows |
|
||||
|--------|---------------|
|
||||
| **Schedules** | Number of active scheduled tasks targeting this node, plus a relative countdown to the next run (e.g. "next 2h") |
|
||||
| **Updates** | Whether auto-update policies are enabled ("Auto" badge), and how many stacks have pending image updates (fuchsia notification dot with expanding halo and a count) |
|
||||
| **Default** | A filled star marks the default (Local) node. The column has no header. |
|
||||
| **Name** | A globe or monitor icon followed by the display name. |
|
||||
| **Type** | `Local` or `Remote` badge. |
|
||||
| **Mode** | `-` for the local node; `Proxy` or `Pilot Agent` badge for remotes, with an icon matching the mode. |
|
||||
| **Endpoint** | `docker.sock` for local; the full Sencho API URL for proxy nodes; `tunnel (seen X ago)` or `tunnel (waiting)` for pilot agents. |
|
||||
| **Status** | `Online`, `Offline`, or `Unknown` badge. |
|
||||
| **Labels** | Per-node label palette. On Skipper and Admiral the cell shows the picker (an empty cell reads `No labels` with an Add label control); on Community the cell shows a single dash. |
|
||||
| **Schedules** | Number of active scheduled tasks targeting this node, plus a `next X` countdown to the next run. Click the count or the calendar icon in the Actions column to filter the Schedules view to that node. |
|
||||
| **Updates** | `Auto` if at least one auto-update policy is enabled on the node; `Off` otherwise. A pulsing dot and count appear when stacks have pending image updates. |
|
||||
| **Actions** | **View Schedules**, **Test Connection**, **Edit Node**, and **Delete Node** icon buttons. The local row hides Delete because the local node cannot be removed. |
|
||||
|
||||
Click the **calendar icon** on any node row to jump directly to the Schedules view filtered to that node. From there you can create, edit, or manage scheduled tasks scoped to the selected node. The filter bar shows which node you are viewing, with a **Clear filter** button to return to the full list.
|
||||
Clicking the schedules-link icon opens the Schedules view filtered to the selected node. From there you can create, edit, or manage scheduled tasks scoped to that node. The filter bar shows which node you are viewing, with a Clear filter control to return to the full list.
|
||||
|
||||
When a node is deleted, all scheduled tasks and update status data associated with it are automatically cleaned up.
|
||||
When a node is deleted, all scheduled tasks and update status data associated with it are cleaned up automatically.
|
||||
|
||||
## What Settings apply per node
|
||||
|
||||
When you select a remote node in the node picker, the Settings hub filters to the panels that control that specific instance. Values saved here never cross over to other nodes.
|
||||
When you select a remote node in the switcher, the Settings hub filters to the panels that control that specific instance. Values saved here never cross over to other nodes.
|
||||
|
||||
| Panel | Scope | Notes |
|
||||
|-------|:-----:|-------|
|
||||
| Appearance | Per browser | Density and theme preferences are stored in your browser, not on the node. |
|
||||
| Appearance | Per browser | Theme and density preferences are stored in your browser, not on the node. |
|
||||
| System Limits | Per node | Host CPU, RAM, disk, and crash-loop thresholds for the selected node. |
|
||||
| Notifications | Per node | Discord, Slack, and Webhook channels fire from the node that detects the event. |
|
||||
| Labels | Per node | Stack and container label palettes. |
|
||||
| Security | Per node | Trivy install, update, and scanner status. Scan policies and CVE suppressions are managed on the control node and apply fleet-wide. |
|
||||
| Security | Per node | Trivy install state and scanner readiness for the selected node. |
|
||||
| Developer | Per node | Retention windows for metrics and logs, plus Developer Mode. |
|
||||
| App Store | Per node | Template registry URL for the selected node's catalog. |
|
||||
| App Store | Per node | Template registry URL and featured-catalog source for the selected node's catalog. |
|
||||
| Support | Per browser | Diagnostics bundle, docs links, contact channels. |
|
||||
| About | Per browser | Build metadata for whichever instance the page is loaded from. |
|
||||
|
||||
Panels that manage control-plane concerns (Account, License, Users, SSO, API Tokens, Registries, Nodes, Routing, Webhooks) are hidden when a remote node is active.
|
||||
Panels that manage control-plane concerns (Account, License, Users, SSO, API Tokens, Registries, Cloud Backup, Nodes, Routing, Webhooks) are hidden when a remote node is active.
|
||||
|
||||
## License enforcement across nodes
|
||||
|
||||
When you have a paid license (Skipper or Admiral) on your primary instance, all remote nodes automatically inherit that license tier for proxied requests. You do not need to activate a license on each remote node separately.
|
||||
When the control instance has a paid license (Skipper or Admiral), all remote nodes inherit that license tier for proxied requests. You do not activate a license on each remote node separately.
|
||||
|
||||
### How it works
|
||||
|
||||
Your primary Sencho instance asserts its license tier to remote nodes on every proxied request. Remote nodes trust this assertion because it arrives alongside the valid node token you configured when adding the node. No additional configuration is required.
|
||||
The control instance asserts its license tier on every proxied request. Remote nodes trust the assertion because it arrives alongside the valid bearer token you configured when adding the node.
|
||||
|
||||
This means:
|
||||
|
||||
- **Paid primary → remote nodes**: Skipper and Admiral features work on all remote nodes, governed by the primary instance's license.
|
||||
- **Community primary → remote nodes**: Paid features are blocked on remote nodes, even if a remote node has its own paid license. The primary's tier is authoritative for proxied requests.
|
||||
- **Direct access to a node**: If you access a remote Sencho instance directly (not through the primary), it uses its own local license tier as usual.
|
||||
|
||||
### Why this matters
|
||||
|
||||
Without this trust chain, remote nodes would default to the Community tier and block paid features, even though the primary instance has a valid license. Distributed license enforcement eliminates this gap so your fleet behaves consistently regardless of which node you are operating on.
|
||||
- **Paid control plane → remote nodes**: Skipper and Admiral features work on every remote node, governed by the control instance's license.
|
||||
- **Community control plane → remote nodes**: Paid features are blocked on remote nodes, even if a remote node has its own paid license. The control instance's tier is authoritative for proxied requests.
|
||||
- **Direct access to a node**: If you load a remote Sencho instance in your browser directly (not through the control plane), it uses its own local license tier.
|
||||
|
||||
<Note>
|
||||
Remote nodes do not need their own license keys. A single license on the primary instance covers all nodes managed through it.
|
||||
Remote nodes do not need their own license keys. A single license on the control instance covers every node managed through it.
|
||||
</Note>
|
||||
|
||||
## Editing and deleting nodes
|
||||
|
||||
Click the **pencil icon** on any node row to edit its name, URL, token, or compose directory. Click the **trash icon** to remove a remote node. The local (default) node cannot be deleted.
|
||||
Click the pencil icon on any row to edit its name, URL, token, or compose directory. For a Pilot Agent row, the Edit modal also surfaces the **Regenerate enrollment token** card described earlier.
|
||||
|
||||
Click the trash icon to remove a remote node. The local row hides this icon because the default node cannot be deleted. Removing a node only deletes the routing entry on the control instance; the remote Sencho instance and its containers are not touched.
|
||||
|
||||
## Security
|
||||
|
||||
### Token security
|
||||
|
||||
Node tokens grant full control over the remote Sencho instance. Treat them like passwords:
|
||||
Bearer tokens grant full control over the remote Sencho instance. Treat them like passwords:
|
||||
|
||||
- **Rotate immediately** if a token is compromised: go to the remote instance's **Settings → Nodes** and click **Generate Token**. The old token is invalidated instantly.
|
||||
- Tokens are **encrypted at rest** in Sencho's database.
|
||||
- **Rotate immediately** if a token is compromised: open the remote instance's **Settings → System → Nodes** and click **Generate Token** to mint a new one. The previous token is invalidated instantly.
|
||||
- Tokens are **encrypted at rest** in the local SQLite database.
|
||||
- Tokens cannot be used to open interactive terminals (Host Console or container exec). Interactive shell access always requires a real browser session on that specific instance.
|
||||
|
||||
### Transport encryption
|
||||
@@ -154,20 +231,14 @@ Node tokens grant full control over the remote Sencho instance. Treat them like
|
||||
Sencho delegates transport encryption to your infrastructure rather than implementing TLS at the application layer. This is the same approach used by other self-hosted tools; it avoids certificate management burden while letting you use the encryption layer that best fits your environment.
|
||||
|
||||
<Warning>
|
||||
**Never send node tokens over plain HTTP across the public internet.** A token intercepted in transit grants full control of the remote instance. Always use one of the approaches below when nodes communicate over untrusted networks.
|
||||
**Never send bearer tokens over plain HTTP across the public internet.** A token intercepted in transit grants full control of the remote instance. Use one of the approaches below when nodes communicate over untrusted networks.
|
||||
</Warning>
|
||||
|
||||
Sencho shows an inline warning when you enter an HTTP URL in the Add Node form as a reminder:
|
||||
|
||||
<Frame>
|
||||
<img src="/images/multi-node/http-warning.png" alt="Add Node form showing an inline warning when an HTTP URL is entered" />
|
||||
</Frame>
|
||||
|
||||
There are three recommended approaches depending on your deployment:
|
||||
There are three recommended approaches.
|
||||
|
||||
#### Private network (LAN or VPC)
|
||||
|
||||
If all your Sencho instances are on the same local network, VPC, or subnet, HTTP is perfectly fine. The token never leaves the private network, so there is no interception risk.
|
||||
If all your Sencho instances are on the same local network, VPC, or subnet, HTTP is fine. The token never leaves the private network, so there is no interception risk.
|
||||
|
||||
```
|
||||
http://192.168.1.50:1852 ← safe on a private LAN
|
||||
@@ -227,6 +298,26 @@ https://sencho.example.com ← TLS terminated by Caddy/Nginx
|
||||
|
||||
### Why Sencho doesn't implement application-layer TLS
|
||||
|
||||
Some tools auto-generate self-signed TLS certificates between their server and agents. In practice, this provides minimal security benefit because the certificates are self-signed with no verification, meaning they do not protect against man-in-the-middle attacks.
|
||||
Some tools auto-generate self-signed TLS certificates between their server and agents. In practice that provides minimal security benefit because the certificates are self-signed with no verification, meaning they do not protect against man-in-the-middle attacks.
|
||||
|
||||
Sencho takes a deliberate approach: infrastructure-level encryption (VPN, reverse proxy, or private networking) is more robust, easier to manage, and does not impose certificate rotation burden on the user. This is the same model used by other self-hosted orchestrators.
|
||||
Sencho takes the opposite approach: infrastructure-level encryption (VPN, reverse proxy, or private networking) is more robust, easier to manage, and does not impose certificate rotation burden on you. This is the same model used by other self-hosted orchestrators.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="A remote node shows Offline">
|
||||
Click **Test connection** on the row to re-probe and read the toast. The most common causes are: a wrong URL (missing scheme, missing port, trailing slash typo), a token that was rotated on the remote (any new **Generate Token** invalidates the previous one, so re-issue and update the saved row), and a firewall or reverse proxy that is not forwarding to port 1852 on the remote. Distributed API Proxy mode requires the control instance to reach the remote on the URL you saved; Pilot Agent mode requires the remote to reach the control instance on outbound HTTPS.
|
||||
</Accordion>
|
||||
<Accordion title="A pilot agent stays on `tunnel (waiting)`">
|
||||
The agent has not dialled home yet. Three things can cause this. First, the enrollment token expired (15 minutes from generation): open **Edit node**, click **Regenerate enrollment token**, copy the new command, and run it on the remote host. Second, the `sencho-agent` container exited or never started; `docker ps -a` on the remote will show the exit code. Third, outbound HTTPS to the control instance is blocked, so check the remote's egress firewall or proxy.
|
||||
</Accordion>
|
||||
<Accordion title="Test Connection fails with 401 Unauthorized">
|
||||
The bearer token saved for the row no longer matches what the remote will accept. This usually means somebody clicked **Generate Token** on the remote (which invalidates the previous token) or the row was saved with a typo. Generate a fresh token on the remote, click the pencil icon on the row in the control instance, paste the new token into the API Token field, and save.
|
||||
</Accordion>
|
||||
<Accordion title="A paid feature is blocked when I switch to a remote node">
|
||||
The control instance's license tier is authoritative for proxied requests, so a Community control plane gates Skipper and Admiral features on every remote, even if the remote itself has its own paid license. Activate a paid license on the control instance to lift the gate fleet-wide. The reverse case (paid control plane, Community remote) works automatically because the control plane's tier is what the remote trusts.
|
||||
</Accordion>
|
||||
<Accordion title="Settings panels disappear when I switch to a remote">
|
||||
That is intentional. Account, License, Users, SSO, API Tokens, Registries, Cloud Backup, Nodes, Routing, and Webhooks are control-plane concerns and are hidden while a remote node is selected. Switch back to **Local** from the node switcher to manage them. The full list of which panels are per-node, per-browser, and control-plane-only lives in the [What Settings apply per node](#what-settings-apply-per-node) table above.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
@@ -1,54 +1,73 @@
|
||||
---
|
||||
title: Node Compatibility
|
||||
description: How Sencho handles version differences across nodes and gracefully degrades features that are not available on older instances.
|
||||
description: How Sencho handles version differences across nodes and gracefully gates features that are not available on the active node.
|
||||
---
|
||||
|
||||
When you manage [multiple nodes](/features/multi-node), each running its own Sencho instance, those instances may be on different versions. A node running an older version will not have features that shipped in a newer release. Sencho detects this automatically and disables features that the active node does not support, with no errors or broken pages.
|
||||
When you manage [multiple nodes](/features/multi-node), each running its own Sencho instance, those instances may be on different versions. A node running an older release will not have features that shipped in a newer one. Sencho detects this automatically and gates each feature on whether the active node advertises support for it, with no errors and no broken pages.
|
||||
|
||||
## How it works
|
||||
|
||||
Every Sencho instance advertises a set of **capabilities**: feature flags describing what it supports. When you switch to a node, your primary instance fetches this metadata and caches it briefly. Features that require a capability the active node does not have are visually dimmed with an explanation overlay.
|
||||
Every Sencho instance exposes a public `GET /api/meta` endpoint that advertises its version and a list of **capabilities**, one flag per feature it implements. When you switch to a node, your control instance fetches that node's metadata and caches it for 5 minutes on success or 30 seconds on failure, so the version pill and the per-feature gates stay in sync without re-hitting the remote on every navigation.
|
||||
|
||||
Core features like stack management, containers, and resource monitoring are always available on every Sencho version.
|
||||
Features whose capability is missing from the active node's list are replaced by a lock card naming the missing feature and prompting you to upgrade the node. Core features (stack management, containers, resources, logs) work on every Sencho version and are never gated this way.
|
||||
|
||||
<Note>
|
||||
Capabilities are flag-based, not version-based. Sencho never compares version numbers; it checks whether the active node explicitly advertises support for each feature. This keeps the system forward-compatible as new capabilities are added.
|
||||
</Note>
|
||||
|
||||
## What you will see
|
||||
|
||||
### Version in the node switcher
|
||||
|
||||
The sidebar node switcher shows each node's Sencho version next to its name. If a node is too old to report its version, no version is shown.
|
||||
The node switcher popover shows each node's Sencho version next to its type chip, separated by a middle dot.
|
||||
|
||||
### Disabled features
|
||||
<Frame>
|
||||
<img src="/images/node-compatibility/node-switcher-versions.png" alt="Node switcher popover listing seven nodes. The Local row reads LOCAL · v0.81.11 with a star icon for the default node. Three remote rows read REMOTE · v0.81.11. The sencho-pilot-test row reads AGENT · SEEN 20M AGO · v0.76.9, calling out that this node is on an older release than the rest of the fleet." />
|
||||
</Frame>
|
||||
|
||||
When you switch to a node that lacks a capability, the affected feature appears dimmed and blurred with a pill overlay explaining why. For example:
|
||||
A node's version pill appears once Sencho has fetched its metadata. The metadata cache populates the first time you visit a node, so freshly-enrolled remotes may show their type chip alone until you switch to them once. A node that does not return metadata at all shows no version pill on either visit.
|
||||
|
||||
> "Fleet Management is not available - prod-server is running v0.28.0"
|
||||
### Lock card on unsupported features
|
||||
|
||||
The feature content is still visible (blurred) so you can see what you are missing, but it is non-interactive. Switch to a node that supports the feature, or update the remote node to the latest Sencho version.
|
||||
When the active node does not advertise a capability that a feature needs, the feature panel is replaced by a centered lock card.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/node-compatibility/lock-card.png" alt="Capability lock card centered in the editor area. A circular icon badge holds an Unplug glyph at the top. Below it, a title reads Host Console is not available on this node and a body line reads sencho-pilot-test is running v0.76.9. Upgrade the node to use this feature." />
|
||||
</Frame>
|
||||
|
||||
The card carries the feature name, the node name, and (when known) the running version, so you can see at a glance what is missing and where. When the node has not reported a version, the body line drops the version and reads "*node-name* does not advertise this capability. Upgrade the node to use this feature." instead. The gated panel itself does not load; the lock card replaces it, which keeps slow capability-heavy views from briefly flashing into view before being hidden.
|
||||
|
||||
### Connection test details
|
||||
|
||||
When you test a remote node's connection in **Settings → Nodes**, the result includes the remote Sencho version alongside Docker info, container counts, and system stats.
|
||||
When you test a remote node from **Settings · Nodes**, the Connection Details panel that appears under the table lists the node's running Sencho version alongside its OS, architecture, container count, image count, and CPU count.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/node-compatibility/connection-test-version.png" alt="Connection Details panel for sencho-test-01 showing Instance: Remote Sencho, Sencho: v0.81.11, OS: Remote, Arch: Remote, Containers: 1, Images: 21, CPUs: 1, with a wifi-style icon in the heading." />
|
||||
</Frame>
|
||||
|
||||
This is the most direct way to confirm what version a remote node is on without switching to it first.
|
||||
|
||||
## Capability list
|
||||
|
||||
Every Sencho release includes a static list of capabilities. Here are the capabilities tracked as of the current version:
|
||||
Every Sencho release ships with a static list of capabilities. The current list is:
|
||||
|
||||
| Capability | Feature |
|
||||
|-----------|---------|
|
||||
| `stacks` | Stack management (always available) |
|
||||
| `containers` | Container operations (always available) |
|
||||
| `resources` | Resource monitoring (always available) |
|
||||
| `resources` | Resource browser (always available) |
|
||||
| `templates` | App Store (always available) |
|
||||
| `global-logs` | Global log viewer (always available) |
|
||||
| `system-stats` | System statistics (always available) |
|
||||
| `fleet` | Fleet management |
|
||||
| `auto-updates` | Auto-update policies |
|
||||
| `system-stats` | Host system statistics (always available) |
|
||||
| `fleet` | Fleet management view |
|
||||
| `auto-updates` | Auto-Update Readiness policies |
|
||||
| `labels` | Stack labels |
|
||||
| `webhooks` | Webhooks |
|
||||
| `webhooks` | Webhook triggers |
|
||||
| `network-topology` | Network topology view |
|
||||
| `notifications` | Alert notifications |
|
||||
| `notification-routing` | Notification routing rules |
|
||||
| `host-console` | Host console |
|
||||
| `host-console` | Host Console |
|
||||
| `container-exec` | Container exec terminal |
|
||||
| `audit-log` | Audit log |
|
||||
| `scheduled-ops` | Scheduled operations |
|
||||
| `sso` | SSO authentication |
|
||||
@@ -56,27 +75,28 @@ Every Sencho release includes a static list of capabilities. Here are the capabi
|
||||
| `users` | User management |
|
||||
| `registries` | Private registry management |
|
||||
| `self-update` | Self-update from the dashboard |
|
||||
| `vulnerability-scanning` | Image vulnerability scanning |
|
||||
|
||||
<Note>
|
||||
Capabilities are flag-based, not version-based. Sencho never compares version numbers; it checks whether the remote node explicitly advertises support for a feature. This means the system stays forward-compatible as new features are added.
|
||||
`vulnerability-scanning` is advertised only when the Trivy binary is present on the node. If a node does not have Trivy installed, it omits this capability and the scanning UI is replaced by a lock card.
|
||||
</Note>
|
||||
|
||||
## Handling older nodes
|
||||
## Handling nodes that do not advertise metadata
|
||||
|
||||
Nodes running a Sencho version from before the compatibility system was introduced do not have the metadata endpoint. Sencho handles this gracefully:
|
||||
If a remote node does not respond to `/api/meta` (for example, an unreachable instance, a slow handshake, or one that does not implement the endpoint), Sencho falls back to an offline metadata record with no version and an empty capability list. In that state:
|
||||
|
||||
- The version shows as blank in the node switcher
|
||||
- All gated features appear as unavailable on that node
|
||||
- Core features (stacks, containers, resources, logs) work normally
|
||||
- No errors are thrown; the UI degrades cleanly
|
||||
- The node's row in the switcher and the connection-test panel show no version pill.
|
||||
- Every capability-gated feature on that node shows the lock card.
|
||||
- Core features (stacks, containers, resources, logs) continue to work normally.
|
||||
- No errors are surfaced in the UI; the control instance retries the metadata fetch after a short backoff.
|
||||
|
||||
The fix is straightforward: update the remote node to the latest Sencho version. See the [upgrade guide](/operations/upgrade) for instructions.
|
||||
If you expect a node to support a feature that is being gated, the fastest fix is to update that node to the latest Sencho release. See the [upgrade guide](/operations/upgrade) for instructions.
|
||||
|
||||
## Interaction with license tiers
|
||||
|
||||
Some features require both a license tier (Skipper or Admiral) **and** node capability support. When both gates apply:
|
||||
Some features need both a license tier (Skipper or Admiral) **and** node capability support. The two gates evaluate in this order:
|
||||
|
||||
1. The license gate is checked first. If you are on the Community tier, you see the upgrade prompt.
|
||||
2. If your license covers the feature but the node does not support it, you see the compatibility overlay instead.
|
||||
1. The license gate is checked first. On the wrong tier, the feature's entry point (sidebar item, top-nav button, settings section) is hidden entirely, so you never reach the panel.
|
||||
2. If you are on the right tier but the active node does not advertise the capability, the entry point is visible but the panel is replaced by the capability lock card.
|
||||
|
||||
This means you will not see confusing "upgrade to unlock" messages for features that would not work on the active node anyway.
|
||||
The practical effect is that you only see the lock card for features your license already covers, so it is always actionable: upgrading the node will unlock the feature.
|
||||
|
||||
@@ -1,80 +0,0 @@
|
||||
---
|
||||
title: Notification Routing
|
||||
description: Route stack alerts to specific Discord, Slack, or webhook channels with per-stack routing rules.
|
||||
---
|
||||
|
||||
<Note>
|
||||
Notification Routing requires a **Sencho Admiral** license. Community and Skipper users can configure global notification channels in **Settings > Notifications**.
|
||||
</Note>
|
||||
|
||||
Notification Routing lets you direct alerts from specific stacks to dedicated channels. Instead of all alerts going to a single global endpoint, you can send production alerts to one Slack channel and staging alerts to another.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/notification-routing/notification-routing-overview.png" alt="Notification Routing settings showing the empty state with Add Route button" />
|
||||
</Frame>
|
||||
|
||||
## How routing works
|
||||
|
||||
When Sencho dispatches an alert (container crash, threshold breach, scheduled task failure), the routing engine:
|
||||
|
||||
1. **Checks routing rules** sorted by priority (lowest number first). If the alert's stack matches a rule's stack list, the notification is sent to that rule's channel.
|
||||
2. **Falls back to global channels** if no routing rule matches (or the alert has no stack context, such as host resource warnings). Global channels are configured in **Settings > Notifications**.
|
||||
|
||||
Routing rules and global channels are independent. A matched route **replaces** the global dispatch for that alert; it does not send to both.
|
||||
|
||||
## Creating a routing rule
|
||||
|
||||
1. Go to **Settings > Routing**
|
||||
2. Click **+ Add Route**
|
||||
3. Fill in the form:
|
||||
|
||||
<Frame>
|
||||
<img src="/images/notification-routing/notification-routing-dialog.png" alt="New Routing Rule dialog with Name, Stacks, Channel, Priority, and Enabled fields" />
|
||||
</Frame>
|
||||
|
||||
| Field | Description |
|
||||
|-------|-------------|
|
||||
| **Name** | A label for this rule (e.g. "Production to Discord") |
|
||||
| **Stacks** | One or more stacks this rule applies to. Use the searchable dropdown to find and add stacks. Selected stacks appear as removable pills below the dropdown. |
|
||||
| **Channel** | Choose Discord, Slack, or Webhook using the tab bar, then enter the endpoint URL. The URL must use HTTPS. |
|
||||
| **Priority** | Lower values are evaluated first. Default is 0. |
|
||||
| **Enabled** | Toggle the rule on or off without deleting it |
|
||||
|
||||
4. Click **Create** to save the rule
|
||||
|
||||
## Managing rules
|
||||
|
||||
Each routing rule appears as a card showing the rule name, channel type badge, assigned stack pills, the channel URL (truncated), and priority (when non-zero).
|
||||
|
||||
From the rule card, you can:
|
||||
|
||||
- **Toggle** a rule on/off with the switch
|
||||
- **Test** a rule by clicking the lightning bolt icon, which sends a test notification to the rule's channel
|
||||
- **Edit** a rule by clicking the pencil icon
|
||||
- **Delete** a rule via the trash icon (with a confirmation dialog)
|
||||
|
||||
## Priority and matching
|
||||
|
||||
Rules are evaluated in ascending priority order. If multiple rules match the same stack, **all matching rules fire**. This lets you send the same alert to multiple channels (e.g. both a Slack channel and a custom webhook).
|
||||
|
||||
If any rule matches, global channels are skipped for that alert.
|
||||
|
||||
## Fallback behavior
|
||||
|
||||
Alerts without a stack context always use global channels. These include:
|
||||
|
||||
- Host CPU, memory, and disk threshold warnings
|
||||
- Docker data accumulation notifications
|
||||
|
||||
Stack-scoped alerts that do not match any routing rule also fall back to global channels.
|
||||
|
||||
## Example setup
|
||||
|
||||
**Scenario:** You want production stack crashes in `#prod-alerts` on Slack, and all staging stacks in a Discord channel.
|
||||
|
||||
| Rule | Stacks | Channel | Priority |
|
||||
|------|--------|---------|----------|
|
||||
| Prod to Slack | `prod-api`, `prod-web` | Slack | 0 |
|
||||
| Staging to Discord | `staging-api`, `staging-web` | Discord | 10 |
|
||||
|
||||
Any other stack alerts (e.g. `dev-tools`) would fall back to your global notification channels.
|
||||
@@ -79,11 +79,11 @@ The **Logs** view aggregates output from all containers across all stacks into a
|
||||
|
||||
### Alerts & notifications
|
||||
|
||||
Configure threshold-based alerts (CPU, memory, network, restart count) per stack. Route notifications to Discord, Slack, or any generic webhook endpoint. Alerts are evaluated every minute with configurable duration and cooldown. [Learn more →](/features/alerts-notifications)
|
||||
Configure threshold-based alerts (CPU, memory, network, restart count) per stack. Route notifications to Discord, Slack, or any generic webhook endpoint. Alerts are evaluated every 30 seconds with configurable duration and cooldown. [Learn more →](/features/alerts-notifications)
|
||||
|
||||
### Notification routing
|
||||
|
||||
Route alerts to specific channels with per-stack routing rules. Send production alerts to a critical Slack channel while routing dev stack alerts to a less urgent Discord channel. Admiral only. [Learn more →](/features/notification-routing)
|
||||
Route alerts to specific channels with per-stack routing rules. Send production alerts to a critical Slack channel while routing dev stack alerts to a less urgent Discord channel. Admiral only. [Learn more →](/features/alerts-notifications#notification-routing)
|
||||
|
||||
### Audit log
|
||||
|
||||
@@ -173,7 +173,7 @@ Store credentials for private Docker registries: Docker Hub organizations, GHCR,
|
||||
|
||||
### Auto-Update Readiness
|
||||
|
||||
Review pending container updates across your fleet with risk tags (`patch`, `minor`, `major`, `digest rebuild`), one-line changelog previews, and rollback targets on a single board. The hero counts pending updates and tells you how many are ready to apply without human review; major version jumps and stacks with blocked registries are counted separately for review. Skipper or Admiral. [Learn more →](/features/auto-update-policies)
|
||||
Review pending container updates across your fleet with risk badges (`Safe · patch`, `Review · minor`, `Blocked · major`, `Digest rebuild`) and one-line changelog previews on a single board. The hero counts pending updates and tells you how many are ready to apply without human review; stacks with a major version bump are surfaced as a separate count for review. Skipper or Admiral. [Learn more →](/features/auto-update-policies)
|
||||
|
||||
### Auto-Heal Policies
|
||||
|
||||
|
||||
@@ -1,64 +1,113 @@
|
||||
---
|
||||
title: Pilot Agent
|
||||
description: Add remote nodes behind NAT, residential networks, or corporate firewalls without exposing any inbound port.
|
||||
description: Connect a remote Sencho node to your control instance through a single outbound WebSocket tunnel. No inbound ports, no TLS certificates, no public URL on the remote host.
|
||||
---
|
||||
|
||||
Pilot Agent mode connects a remote server to your primary Sencho instance through a single outbound WebSocket tunnel. Every request the primary sends to that node, HTTP or WebSocket, rides through this tunnel. The remote host never opens an inbound port and never needs a TLS certificate.
|
||||
Pilot Agent is one of Sencho's two remote-node modes. A small agent container runs on the remote host, dials your control instance over a single outbound WebSocket, and holds that connection open. Every request the control instance sends to that node rides through the same tunnel, whether it's HTTP, WebSocket, or Sencho Mesh TCP. The remote host never opens an inbound port and never needs a TLS certificate of its own.
|
||||
|
||||
This page is the architecture-and-operations reference for Pilot Agent. For the step-by-step UI walkthrough that adds a remote node, see [Multi-Node Management](/features/multi-node).
|
||||
|
||||
<Card title="Sencho Mesh" icon="link" href="/features/sencho-mesh">
|
||||
Once a node is enrolled, you can wire stacks across nodes by hostname using Sencho Mesh, which rides on this same tunnel.
|
||||
Once a node is enrolled, you can wire stacks across nodes by hostname using Sencho Mesh. Mesh traffic rides on the same pilot tunnel.
|
||||
</Card>
|
||||
|
||||
## When to use Pilot Agent
|
||||
## How Pilot Agent fits Sencho
|
||||
|
||||
Pick Pilot Agent when the remote host:
|
||||
- Sits behind NAT or a residential router.
|
||||
Sencho splits responsibilities between a **control instance** (the browser-facing Sencho your operators log into) and **nodes** (Docker hosts you manage from that control instance). The local node is always the control instance's own Docker socket. Remote nodes are other Sencho installations that the control instance manages on your behalf.
|
||||
|
||||
There are two remote modes, chosen per node when you register it:
|
||||
|
||||
<Frame>
|
||||
<img src="/images/pilot-agent/01-add-node-pilot.png" alt="Add remote node modal with Type set to Remote and the Mode dropdown open, showing the Pilot Agent option above the Distributed API Proxy option." />
|
||||
</Frame>
|
||||
|
||||
- **Pilot Agent** (this page). The remote dials the control instance. Outbound HTTPS from the remote is the only network requirement.
|
||||
- **Distributed API Proxy** ([documented under Multi-Node Management](/features/multi-node)). The control instance dials the remote. The remote must expose a reachable URL.
|
||||
|
||||
Both modes are first-class. They can coexist on the same control instance, one per node. Pilot Agent is part of Sencho's Community-tier core surface; there is no license gate on using it.
|
||||
|
||||
## When to choose Pilot Agent
|
||||
|
||||
Pilot Agent is the right choice when the remote host:
|
||||
|
||||
- Sits behind NAT or a residential router with no port forwarding.
|
||||
- Lives on a corporate network that blocks inbound connections.
|
||||
- Roams between networks (laptops, mobile hotspots, rotating cloud IPs).
|
||||
- Would otherwise need a reverse proxy, a dynamic DNS entry, or a self-signed TLS certificate just to be reachable.
|
||||
- Has a rotating cloud IP, a CG-NAT'd mobile uplink, or roams between networks.
|
||||
- Has only outbound internet access (a typical homelab or office workstation).
|
||||
- Would otherwise need a reverse proxy, dynamic DNS, or a self-signed TLS certificate purely to be reachable.
|
||||
|
||||
Pick **Distributed API Proxy** (documented in [Multi-Node Management](/features/multi-node)) when the remote host already has a stable, reachable URL, for example a VPS with a public IP or a LAN server on a home network. Both modes are supported side-by-side, one per node.
|
||||
Choose **Distributed API Proxy** instead when the remote already has a stable, reachable URL (a VPS with a public IP, a LAN server with a static address, or a Tailscale node you have already terminated TLS on). The [Multi-Node Management](/features/multi-node) page compares both modes side by side.
|
||||
|
||||
You do not have to pick one mode for the whole fleet. Mix and match per node.
|
||||
|
||||
## How it works
|
||||
|
||||
The agent runs inside a second container on the remote host, using the same `saelix/sencho:latest` image. Setting `SENCHO_MODE=pilot` plus a primary URL and a one-time enrollment token puts the container into agent mode. On boot it dials the primary at `wss://<primary>/api/pilot/tunnel` and holds that connection open. For every tunneled request the primary demultiplexes frames to an internal loopback server, which re-issues the request locally on the agent host against its Docker socket. License tier, role checks, and all other authorization continue to flow from the primary, exactly like proxy mode.
|
||||
Conceptually, the agent reverses the usual client/server direction.
|
||||
|
||||
Deploying with Docker Compose also lets the primary trigger over-the-air updates of the agent itself from the Fleet view, since Compose-managed containers carry the labels Sencho needs to recreate them in place.
|
||||
1. The agent container starts on the remote host with three environment variables: a mode flag, the control instance URL, and a short-lived enrollment token.
|
||||
2. It dials `wss://<control-instance>/api/pilot/tunnel` and holds the WebSocket open for as long as the container runs.
|
||||
3. For every request the control instance needs to make against that node (listing containers, deploying a stack, streaming logs, opening a console, forwarding mesh traffic), frames are multiplexed over that single connection.
|
||||
4. On the agent side, those frames are demultiplexed and replayed locally against the host's Docker socket and filesystem.
|
||||
|
||||
Only outbound HTTPS from the remote to the primary is required. Nothing else is exposed.
|
||||
Inside the control instance, the tunnel terminates at a **loopback bridge**: a tiny HTTP server on `127.0.0.1:<random-port>` per active tunnel. Every other Sencho feature (the existing HTTP proxy, WebSocket forwarder, mesh dialer, license-tier propagation) treats the bridge as just another remote URL. That is why a pilot-mode node behaves identically to a proxy-mode node from the operator's point of view. Once enrolled, it shows up in the same Nodes table, the same node switcher, the same dashboard.
|
||||
|
||||
## Setting the primary's public URL
|
||||
<Frame>
|
||||
<img src="/images/pilot-agent/03-nodes-table-tunnel.png" alt="Settings · System · Nodes table showing a Local row and a Pilot Agent row. The pilot row has Mode badge 'Pilot Agent', Endpoint 'tunnel (seen 45m ago)', and Status badge 'Online'." />
|
||||
</Frame>
|
||||
|
||||
The enrollment dialog bakes the primary's URL into the compose file under `SENCHO_PRIMARY_URL`. The primary infers that URL from the request the admin's browser sent, which works on a LAN but breaks when the pilot lives on a different network (a public cloud VPS, for example) and the admin happened to open the dialog at `http://127.0.0.1:1852` or a LAN address.
|
||||
What rides through the tunnel:
|
||||
|
||||
Set `SENCHO_PUBLIC_URL` on the primary to lock in a publicly-routable URL:
|
||||
- HTTP requests (every API call and stream).
|
||||
- WebSockets (live logs, host console, container exec, notifications).
|
||||
- Sencho Mesh TCP streams (cross-node service-to-service traffic).
|
||||
|
||||
What does NOT ride through the tunnel:
|
||||
|
||||
- Nothing inbound. The agent opens no listening port.
|
||||
- The agent serves no UI of its own. Operators always work from the control instance.
|
||||
- The agent's internal `/api/health` endpoint is reachable only from the tunnel's loopback bridge, not from outside the agent container.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
| Requirement | Detail |
|
||||
|---|---|
|
||||
| **Remote host** | Docker installed and a user that can run containers. About 1 GB free on the volume that backs `/app/data` and whatever you point `/app/compose` at. |
|
||||
| **Network from remote → control instance** | Outbound TCP to the control instance's HTTPS port (whatever your reverse proxy or Sencho exposes). Nothing inbound. |
|
||||
| **Network at the control instance** | The path from the remote to the control instance must allow WebSocket upgrades end-to-end. If you terminate TLS at a reverse proxy in front of Sencho, that proxy needs `Upgrade: websocket` passthrough. |
|
||||
| **Operator role** | Admin on the control instance. The Settings → Nodes panel is admin-only. |
|
||||
| **License tier** | None. Pilot Agent is Community-tier core surface. |
|
||||
|
||||
### The control instance's public URL
|
||||
|
||||
When the enrollment dialog generates the agent's compose file, it bakes the control instance's URL into `SENCHO_PRIMARY_URL` by reading the Host header on the admin's request. That works on a LAN, but it breaks when the pilot lives on a different network (a public cloud VPS, for example) and the admin happened to open the dialog at `http://127.0.0.1:1852` or a LAN address: the baked-in URL is then unreachable from the pilot.
|
||||
|
||||
Set `SENCHO_PUBLIC_URL` on the control instance to lock in a publicly-routable URL the dialog should always use:
|
||||
|
||||
```bash
|
||||
-e SENCHO_PUBLIC_URL=https://sencho.example.com
|
||||
```
|
||||
|
||||
Any HTTPS hostname reachable from the pilot works: a Cloudflare Tunnel, a reverse proxy, a port-forwarded home network with DDNS. The primary validates the value (must be `http(s)://`, no loopback) and falls back to the request host when unset.
|
||||
Any HTTPS hostname reachable from the pilot works: a Cloudflare Tunnel, a reverse proxy, a port-forwarded home network with DDNS. The control instance validates the value (must be `http(s)://`, no loopback) and falls back to the request Host when unset.
|
||||
|
||||
## Enrollment walkthrough
|
||||
## Enrollment and lifecycle
|
||||
|
||||
### 1. Add the node on the primary
|
||||
The end-to-end UI walkthrough for adding a Pilot Agent node (click paths, screenshots of the Add Node form, the Verify-the-tunnel step) lives in [Multi-Node Management](/features/multi-node#add-a-remote-node-pilot-agent). This section covers the underlying credential lifecycle, what happens during the connection, and how reconnects behave.
|
||||
|
||||
On your primary instance open **Settings → Nodes** and click **Add Node**.
|
||||
### Enrollment
|
||||
|
||||
- **Type:** Remote
|
||||
- **Mode:** Pilot Agent (the default for remote nodes)
|
||||
- **Name:** any label, for example `homelab-nuc`
|
||||
- **Compose Directory:** the folder on the remote host where stack folders will live
|
||||
When you submit the Add Node form with Mode set to Pilot Agent, the control instance:
|
||||
|
||||
Click **Add Node**. The form is replaced by an enrollment dialog with a Docker Compose snippet.
|
||||
1. Inserts the node into its registry with status `unknown` and no endpoint.
|
||||
2. Mints a one-time enrollment JWT scoped to that node. The token's hash is stored on the control instance so it cannot be replayed; the token itself is returned to the operator and embedded in a generated Docker Compose file.
|
||||
3. Shows the compose file in an enrollment dialog with a 15-minute countdown.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/pilot-agent/enrollment-dialog.png" alt="Pilot Agent enrollment dialog with compose file" />
|
||||
<img src="/images/pilot-agent/02-enrollment-dialog.png" alt="Enroll the pilot agent modal with kicker 'NODES · PILOT ENROLLMENT'. Step 1 saves a Docker Compose file as 'compose.yaml' with the env block showing SENCHO_MODE: pilot, SENCHO_PRIMARY_URL: https://sencho.example.com, and SENCHO_ENROLL_TOKEN: <short-lived token>. Step 2 shows the 'docker compose -f compose.yaml up -d' command. Footer reads 'Expires 14m from now' alongside a 'Copy compose file' button." />
|
||||
</Frame>
|
||||
|
||||
### 2. Deploy the agent on the remote host
|
||||
The enrollment token is one-time and short-lived by design. If a token is intercepted or leaked, the window of risk is small and the slot is consumed on first use.
|
||||
|
||||
Copy the compose file and save it as `compose.yaml` on the remote host. It looks like:
|
||||
The generated compose file looks like this, with the three pilot env vars already baked in:
|
||||
|
||||
```yaml
|
||||
name: sencho-agent
|
||||
@@ -80,35 +129,77 @@ volumes:
|
||||
sencho-agent-data:
|
||||
```
|
||||
|
||||
Bring the agent up from the same directory:
|
||||
Save it as `compose.yaml` on the remote host and bring the agent up from the same directory:
|
||||
|
||||
```bash
|
||||
docker compose -f compose.yaml up -d
|
||||
```
|
||||
|
||||
The enrollment token is single-use and expires after 15 minutes. On first connect the agent exchanges it for a long-lived tunnel credential, persisted inside the container volume at `/app/data/pilot.jwt`. Subsequent restarts reconnect automatically without needing a new token.
|
||||
Deploying through Compose also lets the control instance push over-the-air updates to the agent later from the Fleet view ([Remote Updates](/features/remote-updates)), without manual intervention on the remote host. Compose-managed containers carry the labels Sencho needs to recreate them in place.
|
||||
|
||||
Deploying the agent through Docker Compose also lets the primary push remote updates to it from the Fleet view ([Remote Updates](/features/remote-updates)) without manual intervention on the remote host.
|
||||
### First connect
|
||||
|
||||
### 3. Confirm the node is online
|
||||
The agent boots, reads the three env vars, and dials `wss://<control-instance>/api/pilot/tunnel` carrying the enrollment token in an `Authorization: Bearer` header.
|
||||
|
||||
The node flips to **Online** in the primary within a few seconds of the agent container starting. The **Endpoint** column shows `tunnel` alongside the time the primary last saw a frame from the agent.
|
||||
The control instance verifies the token, marks the enrollment slot consumed, and replies with a `hello` frame followed by a control frame that carries a **long-lived tunnel JWT** (365-day expiry). The agent writes that token to its data volume at `/app/data/pilot.jwt`. The original enrollment token is now useless: the agent does not need it again, and the control instance refuses to accept it a second time.
|
||||
|
||||
Open the node in the sidebar switcher and use it exactly like any other node: deploy stacks, tail logs, open container terminals, watch host stats.
|
||||
The tunnel is now active. The Endpoint column in the Nodes table flips from `tunnel (waiting)` to `tunnel (seen Xs ago)` and the node's status badge turns Online.
|
||||
|
||||
## Day-to-day operation
|
||||
### Subsequent reconnects
|
||||
|
||||
Once enrolled, a Pilot Agent node is indistinguishable from a proxy-mode node in the UI. The same switcher, the same stacks page, the same editor, the same host console, the same dashboard.
|
||||
On every later boot, restart, or transient network drop, the agent reads `pilot.jwt` from disk and dials with that. No new operator action is needed. The persisted token is good for a year from the original enrollment.
|
||||
|
||||
The agent reconnects automatically if the tunnel drops (transient network blip, primary restart, remote host reboot). It applies an exponential backoff between retries starting at 1 second and capping at 60 seconds, so reconnect traffic stays bounded even during long outages.
|
||||
The reconnect loop uses exponential backoff from 1 second to 60 seconds with up to 500 ms of jitter. The backoff resets only after a successful `hello` round-trip, so an agent that always fails the handshake (wrong URL, revoked token, version mismatch) cannot tight-loop and burn resources.
|
||||
|
||||
## Regenerating enrollment
|
||||
The control instance pings every connected agent every 30 seconds. Any stream with no activity for 10 minutes is closed and reclaimed.
|
||||
|
||||
If the agent container is destroyed before its first successful connect, or the enrollment token expires before you paste it on the remote, open the node's edit dialog on the primary and click **Regenerate enrollment token**. A fresh 15-minute token is issued and the previous tunnel (if any) is disconnected so the new agent replaces it cleanly.
|
||||
### Primary restart, agent restart
|
||||
|
||||
## Self-signed primary TLS certs
|
||||
- **Control instance restarts**: every agent's tunnel drops cleanly. Each agent reconnects within seconds. In-flight HTTP requests fail and the browser retries; long-lived WebSockets (logs, console) reconnect.
|
||||
- **Agent restarts**: the persisted token re-authenticates without operator intervention. Same reconnect-and-recover behavior.
|
||||
- **Operator deletes the node on the control instance**: the agent's existing tunnel is closed. The agent will keep trying to reconnect with its persisted token, which the control instance no longer accepts; the agent stays in a reconnect-backoff loop. To stop it, tear the agent down on the remote with `docker compose -f compose.yaml down -v`.
|
||||
|
||||
If your primary terminates TLS with an internal CA and the agent cannot validate the certificate against the system trust store, point the agent at the CA bundle with `SENCHO_PILOT_CA_FILE`. Add the bundle mount and the env var to the compose file:
|
||||
### Re-enrollment
|
||||
|
||||
If the agent container is destroyed before its first successful connect, the enrollment token expires unused, the control instance is restored from a backup with a different JWT secret, or you simply want to rotate the credential, open the node's Edit dialog and click **Regenerate enrollment token**.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/pilot-agent/04-regenerate-token.png" alt="Edit node modal for a pilot-agent node showing a Re-enroll the agent helper text and a Regenerate enrollment token button." />
|
||||
</Frame>
|
||||
|
||||
A fresh 15-minute token is issued and the previous tunnel (if any) is closed immediately so the new agent replaces it cleanly.
|
||||
|
||||
## Security and trust model
|
||||
|
||||
### Tokens
|
||||
|
||||
Two JWT scopes, both signed with the control instance's `auth_jwt_secret`:
|
||||
|
||||
| Token | Lifetime | Where it lives | Use |
|
||||
|---|---|---|---|
|
||||
| **Enrollment** | 15 min, one-time | Returned to the operator; SHA-256 hash stored on the control instance | Bootstraps the first connect |
|
||||
| **Tunnel** | 365 days | Agent's container volume at `/app/data/pilot.jwt` | Every reconnect after the first |
|
||||
|
||||
Enrollment is rate-limited to 10 requests per minute per operator (per source IP if unauthenticated) so a runaway script cannot exhaust the database with pending tokens.
|
||||
|
||||
### TLS
|
||||
|
||||
The agent dials `wss://`. Certificate validation is on by default against the system trust store. There is no flag to disable verification; that would defeat the credential model. For internal CAs, mount your bundle into the agent container and set `SENCHO_PILOT_CA_FILE` to the in-container path. See [Self-signed control-instance TLS certificates](#self-signed-control-instance-tls-certificates) below for a worked example.
|
||||
|
||||
### Trust boundaries
|
||||
|
||||
- **The control instance trusts the agent.** A connected agent's frames are believed. A compromised remote can flood frames; the per-tunnel resource limits documented below bound the blast radius.
|
||||
- **The agent trusts the control instance.** The control instance proves its identity by producing a valid JWT under the secret it shares with itself. A compromised control instance can ask the agent to do anything the agent could do directly on its host, and the agent holds the Docker socket, so that is effectively root on the remote.
|
||||
- **Tier-trust headers (`x-sencho-tier`, `x-sencho-variant`) are only honored for bearer tokens with `node_proxy` or `pilot_tunnel` scope.** Browser session cookies cannot spoof them.
|
||||
- **The agent exposes nothing inbound.** No port. No UI. No directly reachable health endpoint.
|
||||
|
||||
### What rotating the JWT secret does
|
||||
|
||||
The tunnel JWTs are signed with the control instance's `auth_jwt_secret`. If that secret rotates (because the control instance was rebuilt from scratch, restored into a different environment, or manually rotated), every existing tunnel JWT stops verifying. The agents will reconnect, fail authentication, and back off. Re-enroll each affected node (regenerate the enrollment token, then redeploy the agent on the remote with the refreshed compose file) to issue a new tunnel JWT signed by the current secret.
|
||||
|
||||
## Self-signed control-instance TLS certificates
|
||||
|
||||
If your control instance terminates TLS with an internal CA that the agent's system trust store does not recognise, mount the CA bundle into the agent container and point `SENCHO_PILOT_CA_FILE` at it:
|
||||
|
||||
```yaml
|
||||
name: sencho-agent
|
||||
@@ -132,53 +223,133 @@ volumes:
|
||||
sencho-agent-data:
|
||||
```
|
||||
|
||||
The agent uses the bundle as the only trust anchor for the tunnel WebSocket. TLS verification stays on; there is no env var to disable verification globally because that would defeat the credential trust model.
|
||||
The agent uses the bundle as the only trust anchor for the tunnel WebSocket. TLS verification stays on.
|
||||
|
||||
## Resource limits
|
||||
|
||||
Each tunnel has fixed protocol-level ceilings to keep one misbehaving agent from impacting the primary:
|
||||
Each pilot tunnel has fixed protocol-level ceilings so one misbehaving agent cannot impact the control instance or its peers.
|
||||
|
||||
- **Frame size:** individual WebSocket frames are capped at 8 MB. Single requests and responses larger than that are rejected and the tunnel reconnects.
|
||||
- **Concurrent streams:** at most 1024 multiplexed HTTP and WebSocket streams per tunnel (mesh TCP streams share the same pool). Above the cap the bridge returns 503 and the agent rejects new incoming streams with a typed error frame.
|
||||
- **Stream idle:** any stream with no activity for 10 minutes is closed and removed.
|
||||
- **System-wide tunnels:** a single primary accepts up to 256 concurrent pilot tunnels. Past the soft warning at 128 the primary logs a WARN; past the hard cap of 256 new tunnels are refused with WebSocket close code 1013 (Try Again Later) so the agent backs off rather than tight-looping.
|
||||
| Limit | Value | What it caps | Why it matters |
|
||||
|---|---|---|---|
|
||||
| **Frame size** | 8 MB | Individual WebSocket frame, including a single HTTP request or response body chunk | A request or response body larger than 8 MB is rejected and the tunnel reconnects. Stream HTTP responses are chunked, so normal traffic stays well under this. |
|
||||
| **Concurrent streams per tunnel** | 1024 | HTTP, WebSocket, and Sencho Mesh TCP streams share the same pool | New streams above the cap are rejected with 503 from the bridge and a typed error frame from the agent. A runaway script or a long-poll session that leaks streams can exhaust this; reload the affected tab or throttle the caller. |
|
||||
| **Stream idle timeout** | 10 minutes | Per-stream inactivity before teardown | Idle WebSockets and forgotten HTTP streams are reclaimed automatically. |
|
||||
| **System-wide tunnels** | 256, soft-warn at 128 | Concurrent pilot tunnels on a single control instance | Past the soft warning the control instance logs WARN. Past the hard cap, new tunnels are refused with WebSocket close code 1013 ("Try Again Later"); the agent backs off rather than tight-looping. |
|
||||
| **Tunnel write buffer pause** | 4 MB | High-water mark for the agent-bound write queue | When the control instance's queue is over the mark, inbound HTTP request bodies are paused and resumed on drain. Backpressure rather than memory growth. |
|
||||
| **Enrollment rate** | 10 per minute | Operator-issued enrollment tokens | Protects the database against a script accidentally minting tokens in a loop. |
|
||||
|
||||
An admin-only diagnostic endpoint, `GET /api/system/pilot-tunnels`, returns counters and a per-node breakdown of active tunnels (open stream count and `bufferedAmount` for each). Useful for triaging a single sticky tunnel inside an otherwise healthy fleet.
|
||||
|
||||
## Environment variables
|
||||
|
||||
These environment variables are read by the **agent** container at boot.
|
||||
|
||||
| Variable | Required? | Default | Notes |
|
||||
|---|---|---|---|
|
||||
| `SENCHO_MODE` | Yes | none | Must be `pilot`. Putting any other value here causes the container to start as a normal Sencho instance, not as an agent. |
|
||||
| `SENCHO_PRIMARY_URL` | Yes | none | The base URL of your control instance (e.g. `https://sencho.example.com`). The agent appends `/api/pilot/tunnel` and dials `wss://`. |
|
||||
| `SENCHO_ENROLL_TOKEN` | First boot only | none | The 15-minute enrollment token issued by the control instance. Ignored on subsequent boots once `pilot.jwt` is on disk. |
|
||||
| `SENCHO_PILOT_CA_FILE` | Optional | unset | Absolute path inside the container to a PEM bundle. Use when your control instance's TLS chain is rooted in a private CA. |
|
||||
| `DATA_DIR` | Optional | `/app/data` | Where the persisted `pilot.jwt` is stored. Override only if you are mounting a different volume layout. |
|
||||
| `COMPOSE_DIR` | Optional | `/app/compose` | Root directory on the agent host where compose stack folders live. The control instance reads and writes stacks under this path. |
|
||||
|
||||
One variable is read on the **control instance** side, not the agent:
|
||||
|
||||
| Variable | Required? | Default | Notes |
|
||||
|---|---|---|---|
|
||||
| `SENCHO_PUBLIC_URL` | Optional | request Host | Locks in the URL the enrollment dialog bakes into `SENCHO_PRIMARY_URL`. See [The control instance's public URL](#the-control-instances-public-url) above. Must be `http(s)://`, no loopback. |
|
||||
|
||||
The agent does not currently expose tunables for the ping interval (30 s), the reconnect backoff (1–60 s), or the per-stream idle timeout (10 min). These are deliberate fixed values; if your environment requires different ones, open an issue.
|
||||
|
||||
## Limitations and non-goals
|
||||
|
||||
These are the boundaries operators should know about before designing a fleet around Pilot Agent.
|
||||
|
||||
- **No mode conversion.** The Edit dialog shows a Mode field for an enrolled node, but switching a node between Pilot Agent and Distributed API Proxy after enrollment leaves the credentials and connection state inconsistent. To change modes, delete the node and re-create it in the desired mode.
|
||||
- **No audit log entries for enrollment lifecycle.** Node creation, enrollment regeneration, and node deletion do not write to the audit log today. This is on the roadmap.
|
||||
- **`pilot.jwt` is not cleaned up on node deletion.** When you delete a node from the control instance, the agent's persisted token stays on the remote's data volume. The agent will fail to reconnect on next restart, but the file persists. If you are repurposing the host, tear the agent down with `docker compose -f compose.yaml down -v` to remove the `sencho-agent-data` volume.
|
||||
- **JWT-secret rotation invalidates every tunnel.** Rebuilding the control instance from scratch or rotating `auth_jwt_secret` requires re-enrolling every agent. There is no out-of-band re-issuance flow.
|
||||
- **One tunnel per node.** Splitting a node's load across multiple control instances or running multiple agent containers against the same control instance for the same node is not supported.
|
||||
- **Mesh and pilot share the per-tunnel stream pool.** A node that runs heavy Sencho Mesh traffic counts those streams against the same 1024-stream cap as HTTP and WebSocket traffic.
|
||||
- **The agent has no UI of its own.** All operation flows through the control instance. The agent's container logs (`docker logs sencho-agent`) are the only direct visibility into agent-side behaviour.
|
||||
|
||||
## Sencho Mesh on Pilot Agent
|
||||
|
||||
Sencho Mesh works identically across both remote modes. On a Pilot Agent node, mesh traffic rides the existing long-lived tunnel instead of opening a separate connection. Stack opt-in to the mesh is per stack per node and is not affected by the transport choice. See [Sencho Mesh](/features/sencho-mesh) for the full mesh model.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
The generic node-connectivity issues (a node showing Offline, a pilot agent stuck on `tunnel (waiting)`, 401 Unauthorized on Test Connection) are covered in the [Multi-Node Management troubleshooting section](/features/multi-node#troubleshooting). The entries below cover failure modes that are specific to Pilot Agent.
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Node stays Offline after starting the agent container">
|
||||
Check the agent container's logs for the first `[Pilot]` line. It prints the primary URL it is dialing and the reason for any failed connect. Common causes: the primary URL is wrong or unreachable from the remote host, the enrollment token expired, or the primary is behind a reverse proxy that strips WebSocket upgrades.
|
||||
<Accordion title="Tunnel keeps disconnecting every ~30 seconds">
|
||||
The control instance pings the agent every 30 seconds; a reconnect that lines up with the ping interval almost always means an HTTP proxy or load balancer between the agent and the control instance is closing idle WebSockets faster than the ping cadence. Raise the proxy's WebSocket idle timeout above 30 seconds (a 90-second floor is safe). If you terminate TLS at a reverse proxy, also confirm that the proxy forwards the `Upgrade: websocket` header end-to-end.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Tunnel keeps disconnecting and reconnecting">
|
||||
Usually caused by an HTTP proxy or load balancer in front of the primary with a short idle timeout. The tunnel sends a ping every 30 seconds, so raise the proxy's WebSocket idle timeout above that (90 seconds is a safe floor). If you terminate TLS on the proxy, make sure WebSocket upgrade passthrough is enabled.
|
||||
<Accordion title="Tunnel rejected with WebSocket close code 1013 (Try Again Later)">
|
||||
The control instance has reached the system-wide hard cap of 256 concurrent pilot tunnels. A reconnect storm or a runaway enrollment is the usual cause. Inspect the control instance's logs for the `[Pilot] Active tunnel count at soft limit` warning that fires at 128 to spot the trend, and check `GET /api/system/pilot-tunnels` (admin-only) for the per-node breakdown to identify the flapping agent. The agent backs off and retries automatically once headroom returns.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Enrollment token expired before I ran the docker command">
|
||||
Open **Settings → Nodes** on the primary, click the pencil icon on the pending node, then **Regenerate enrollment token**. The dialog shows a fresh command with a new token.
|
||||
<Accordion title="Loopback bridge returns 503 'pilot tunnel: stream cap reached'">
|
||||
A single tunnel has hit the per-tunnel 1024-stream cap. The most common trigger is a browser session that opened many long-poll streams without closing them, or a script firing parallel API calls without bound. Reload the affected tab so its WebSocket and SSE streams reset. If the cap is being hit by automation, throttle the caller. Sencho Mesh streams share the same pool, so a runaway mesh workload can also push the count up.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Tunnel closes with WebSocket close code 1002 (protocol error)">
|
||||
The agent sent a frame larger than the 8 MB ceiling, or the wire decoder rejected a malformed frame. Run the control instance with developer mode enabled (Settings → Developer) to surface the diagnostic decode log, then check the logs for `[PilotBridge:diag] Malformed frame from agent`. The agent reconnects automatically; persistent failures usually mean a version mismatch between the control instance and the agent image.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="I ran the enrollment command on the wrong host">
|
||||
Tear the agent down on the wrong host (`docker compose -f compose.yaml down -v` from the directory you saved the file in), then regenerate the enrollment token on the primary and deploy the new compose file on the correct host.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Primary was restored from backup and the agent will not reconnect">
|
||||
The agent's persisted tunnel credential is signed with the primary's JWT secret. If that secret is rotated or the primary is rebuilt from scratch, existing tunnels stop verifying. Regenerate enrollment for the affected node and restart the agent container so it consumes the fresh token.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Tunnel rejected with 'pilot tunnel cap reached' (close code 1013)">
|
||||
The primary is at the system-wide concurrent tunnel cap of 256. A reconnect storm or runaway enrollment is the usual cause. Inspect the primary's logs for the `[Pilot] Active tunnel count at soft limit` warning that fires at 128 to find the trend, and check `GET /api/system/pilot-tunnels` (admin-only) for the per-node breakdown to identify a flapping agent. The agent backs off and retries automatically once headroom returns.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Bridge returns 503 'pilot tunnel: stream cap reached'">
|
||||
A single tunnel is at the 1024 concurrent stream cap. The most common trigger is a UI session that opened many long-poll streams and never closed them, or a script firing parallel API calls without bound. Reload the affected browser tab so its WebSocket and SSE streams reset. If the cap is being hit by automation, throttle the caller; the cap protects the primary's memory from a runaway agent or proxy.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Tunnel closes immediately with 'protocol error' (close code 1002)">
|
||||
The agent sent a frame larger than the 8 MB ceiling, or the wire decoder rejected a malformed frame. Run the agent with developer mode enabled on the primary (Settings → Developer) to surface the diagnostic decode log, then check the primary's logs for `[PilotBridge:diag] Malformed frame from agent`. The agent reconnects automatically; persistent failures usually mean a version mismatch between primary and agent images.
|
||||
Tear the agent down on the wrong host (`docker compose -f compose.yaml down -v` from the directory you saved the compose file in), then regenerate the enrollment token on the control instance and deploy the new compose file on the correct host.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="HTTP 429 'Too many enrollment requests'">
|
||||
Enrollment endpoints are rate-limited to 10 requests per minute per user (or per IP, when unauthenticated) so a script accidentally minting tokens in a loop cannot exhaust the database. Wait sixty seconds and retry. If you genuinely need to bulk-enroll many nodes, space the requests out or contact the operator who runs the primary about raising the limit.
|
||||
The enrollment endpoint is rate-limited to 10 requests per minute per operator (or per IP when unauthenticated). A script accidentally minting tokens in a loop trips this. Wait 60 seconds and retry. To bulk-enroll many nodes, space the requests out.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Control instance was restored from backup and the agent will not reconnect">
|
||||
The persisted tunnel credential is signed with the control instance's `auth_jwt_secret`. If that secret was not part of the backup (or has been rotated for any other reason), existing tunnels stop verifying. Regenerate enrollment for each affected node from Settings → Nodes and restart the agent containers so they consume the fresh tokens. There is no global re-issuance command.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="I deleted the node on the control instance but the agent keeps trying to reconnect">
|
||||
Deleting the node removes the routing entry on the control instance but does not touch the remote. The agent still holds `/app/data/pilot.jwt` and will keep dialing. Tear the agent down on the remote (`docker compose -f compose.yaml down -v` from the directory holding the compose file) so the `sencho-agent-data` named volume that stores the persisted token is removed.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Stacks tab on a pilot node shows nothing or fails to list">
|
||||
The agent's compose directory is missing or not writable. Verify the host path you mounted at `/app/compose` exists, is owned by a user the agent container can write as, and matches the Compose Directory configured for the node in Settings → Nodes. If the host path has moved, edit the node and update the Compose Directory field.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="I changed the Mode on a pilot node in Edit and now it is broken">
|
||||
The Edit dialog shows the Mode field, but switching from Pilot Agent to Distributed API Proxy (or back) after enrollment is not a clean flow; the persisted credentials and the new mode end up inconsistent. Delete the node, re-create it in the mode you want, and re-enroll or re-paste the API token.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Common questions
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Do I need port forwarding or a public IP on the remote?">
|
||||
No. The agent dials outbound; the remote needs no inbound port, no public IP, no DNS entry, and no TLS certificate of its own.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Is Pilot Agent the same as a VPN or Tailscale?">
|
||||
No. A VPN gives every host on the network a routable address for arbitrary peer-to-peer traffic. Pilot Agent is a single-purpose tunnel between one agent and one control instance, carrying only Sencho's own protocol. The two are complementary; you can run Sencho over a VPN (using Distributed API Proxy mode against the VPN address) if you prefer.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Does my data go through Sencho's servers?">
|
||||
No. The tunnel is a direct WebSocket between your remote host and your control instance. There is no Sencho-operated relay.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Can I run multiple agents on the same host?">
|
||||
Yes. Each agent needs its own container name and its own data volume so the persisted `pilot.jwt` files do not collide. Enroll each one against a distinct node entry on the control instance.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Does the control instance need a public domain?">
|
||||
No. The agent only needs to reach the control instance over whatever path your network provides: LAN, VPC, VPN, public internet, or a private tunnel. Set `SENCHO_PRIMARY_URL` to whatever URL works from the agent's vantage point.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Related
|
||||
|
||||
- [Multi-Node Management](/features/multi-node): the operator walkthrough for adding remote nodes in either mode, the Nodes table reference, and the cross-mode comparison.
|
||||
- [Sencho Mesh](/features/sencho-mesh): cross-node networking that rides on the pilot tunnel.
|
||||
- [Fleet view](/features/fleet-view): fleet-wide aggregation and node health across both modes.
|
||||
- [Security architecture](/security): credential model and trust posture across the rest of Sencho.
|
||||
|
||||
@@ -1,100 +1,123 @@
|
||||
---
|
||||
title: Private Registries
|
||||
description: Store credentials for private Docker registries so Sencho can automatically authenticate during deploy, pull, and image update checks.
|
||||
description: Store credentials for private Docker registries so Sencho can authenticate automatically during deploy, pull, and image-update checks.
|
||||
---
|
||||
|
||||
<Note>
|
||||
Private Registry Management requires a Sencho **Admiral** license. Skipper and Community Edition do not include this feature.
|
||||
Private Registries is an **Admiral** tier, admin-only feature. Credentials are managed centrally on the control instance and applied fleet-wide; the Registries section is hidden when you are viewing a remote node.
|
||||
</Note>
|
||||
|
||||
Sencho can store credentials for your private Docker registries and inject them automatically whenever it runs `docker compose pull` or `docker compose up`. This means your stacks can reference private images without needing to manually `docker login` on the host.
|
||||
Sencho stores credentials for your private Docker registries and injects them automatically whenever it runs `docker compose pull` or `docker compose up`. Stacks can reference private images without anyone having to run `docker login` on the host.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/private-registries/registries-overview.png" alt="Registries section in Settings showing one configured GHCR registry card with the global scope masthead stat" />
|
||||
</Frame>
|
||||
|
||||
## Supported registry types
|
||||
|
||||
| Type | Description | Credentials |
|
||||
|------|-------------|-------------|
|
||||
| **Docker Hub** | Private Docker Hub organizations and repositories | Username + access token |
|
||||
| **GHCR** | GitHub Container Registry (`ghcr.io`) | GitHub username + personal access token (PAT) |
|
||||
| **AWS ECR** | Amazon Elastic Container Registry | AWS Access Key ID + Secret Access Key |
|
||||
| **Custom** | Any self-hosted Docker V2 registry | Username + password or token |
|
||||
| **GitHub Container Registry (GHCR)** | `ghcr.io` images for users and organizations | GitHub username + personal access token (PAT) |
|
||||
| **AWS Elastic Container Registry (ECR)** | Amazon ECR private registries | AWS Access Key ID + Secret Access Key (+ region) |
|
||||
| **Custom / Self-hosted** | Any Docker V2 compatible registry | Username + password or token |
|
||||
|
||||
## Where to find it
|
||||
|
||||
Open **Settings → System → Registries**. The section appears only on the control instance and only to admin operators on an Admiral tier license.
|
||||
|
||||
## Adding a registry
|
||||
|
||||
1. Open **Settings Hub** and navigate to the **Registries** tab (visible to Admiral admins only).
|
||||
2. Click **Add Registry**.
|
||||
3. Select the registry type. When you switch types, the URL field auto-fills with the standard endpoint for Docker Hub and GHCR.
|
||||
4. Enter a descriptive name, the registry URL, and your credentials.
|
||||
5. For **ECR** registries, also provide the AWS region (e.g., `us-east-1`).
|
||||
6. Click **Test connection** to verify the credentials work before saving. Sencho talks to the registry's `/v2/` endpoint (or the AWS STS API for ECR) and reports success or the specific failure reason.
|
||||
7. Click **Add** to store the credentials.
|
||||
1. Click **Add registry**. The form expands inline above the registry list.
|
||||
2. Pick a **Registry Type**. The form rearranges to match: for ECR an **AWS Region** field appears and the credential fields are relabelled to **AWS Access Key ID** and **AWS Secret Access Key**; for GHCR the **Registry URL** is pre-filled with `ghcr.io`.
|
||||
3. Enter a descriptive **Name**.
|
||||
4. Enter the **Registry URL**. For Docker Hub the field is read-only and the canonical `https://index.docker.io/v1/` value is used automatically.
|
||||
5. Enter the credentials (username + secret, or AWS keys for ECR).
|
||||
6. For ECR, fill in the **AWS Region** (e.g., `us-east-1`).
|
||||
7. Click **Test connection** to probe the registry before saving. For Docker, GHCR, and self-hosted registries this hits the `/v2/` endpoint; for ECR it calls AWS STS `GetAuthorizationToken`.
|
||||
8. Click **Add** to store the credentials.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/private-registries/registries-with-entry.png" alt="Private Registries management view in Settings Hub showing a configured Docker Hub registry" />
|
||||
<img src="/images/private-registries/registries-empty.png" alt="Empty Registries section with the Add registry button and the 'No private registries configured' callout" />
|
||||
</Frame>
|
||||
|
||||
<Frame>
|
||||
<img src="/images/private-registries/registries-add-form.png" alt="Add Registry form with type selector, credentials, and URL fields" />
|
||||
<img src="/images/private-registries/registries-add-form.png" alt="Add registry form expanded with Registry Type set to Docker Hub, the Registry URL field locked to the canonical value, and credential fields below" />
|
||||
</Frame>
|
||||
|
||||
## Managing registries
|
||||
|
||||
Each configured registry appears as a card showing:
|
||||
Each saved registry renders as a card:
|
||||
|
||||
- **Name** and **type badge** (e.g., Docker Hub, GHCR, AWS ECR, Custom)
|
||||
- **Registry URL** (monospaced)
|
||||
- **Username**
|
||||
- **Secret status** (whether a secret is stored)
|
||||
- **AWS Region**, if the registry is an ECR type
|
||||
- **Created date**
|
||||
<Frame>
|
||||
<img src="/images/private-registries/registries-card-detail.png" alt="Registry card showing a GHCR entry with type badge, URL, username, Secret stored pill, created date, and Test connection, Edit, and Delete icon buttons" />
|
||||
</Frame>
|
||||
|
||||
Each card has three action buttons:
|
||||
The card surfaces:
|
||||
|
||||
- **Test connection** (checkmark icon) to verify credentials
|
||||
- **Edit** (pencil icon) to modify the registry name, URL, credentials, or type
|
||||
- **Delete** (trash icon) to remove the registry after confirmation
|
||||
- **Name** and **type badge** (`Docker Hub`, `GitHub (GHCR)`, `AWS ECR`, or `Custom`).
|
||||
- **Registry URL** in a monospaced font.
|
||||
- **Username**.
|
||||
- **Secret stored** (or **No secret**) with a status icon.
|
||||
- **Region**, when the registry is an ECR type.
|
||||
- **Created** date.
|
||||
|
||||
When editing a registry, you can leave the secret field blank to keep the existing stored secret. Only fill it in if you want to replace it.
|
||||
Three icon actions sit on the right of the card:
|
||||
|
||||
- **Test connection** (checkmark icon) re-runs the live probe against the saved credentials.
|
||||
- **Edit** (pencil icon) re-opens the form with the existing values. The **Secret / Token** field shows `(leave blank to keep current)` as its placeholder; type a new value only if you want to rotate the credential.
|
||||
- **Delete** (trash icon) opens a destructive confirmation dialog.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/private-registries/registries-delete-confirm.png" alt="Delete registry confirmation dialog warning that stacks using images from this registry will fail to pull until credentials are re-added" />
|
||||
</Frame>
|
||||
|
||||
Deleting a registry removes the credential record immediately. Any stack that references images from that registry will start failing its pull step on the next deploy until you add the registry back.
|
||||
|
||||
### Testing connectivity
|
||||
|
||||
You can test credentials at two points:
|
||||
You can verify credentials at two points:
|
||||
|
||||
- **Before saving**, using the **Test connection** button inside the add or edit form. This is useful for confirming a token or password works without committing it to storage first.
|
||||
- **After saving**, by clicking the checkmark icon on any registry card. This re-decrypts the stored secret and re-runs the same probe.
|
||||
- **Before saving**, from the **Test connection** button inside the add or edit form. The probe runs against the values currently in the form and never persists them, so you can confirm a token or password works before committing it.
|
||||
- **After saving**, from the **Test connection** icon on any registry card. Sencho decrypts the stored secret and re-runs the same probe.
|
||||
|
||||
For standard registries, Sencho authenticates against `/v2/` using Basic auth and falls back to the registry's token endpoint if it receives a `401` with a `WWW-Authenticate` challenge. For ECR registries, the test verifies that the AWS credentials can successfully obtain an authorization token.
|
||||
For Docker Hub, GHCR, and self-hosted registries, the probe authenticates against `/v2/` with Basic auth and follows a token-endpoint redirect on a 401 `WWW-Authenticate` challenge. For ECR, the probe asks AWS for an authorization token and returns success once the call succeeds.
|
||||
|
||||
## How credentials are applied
|
||||
|
||||
### Deploy and pull operations
|
||||
|
||||
When you deploy or update a stack, Sencho:
|
||||
For every deploy or update Sencho:
|
||||
|
||||
1. Resolves credentials for all configured registries.
|
||||
2. For ECR registries, reuses a cached authorization token when one is still valid; otherwise fetches a fresh token from AWS. Cached tokens are refreshed a few minutes before AWS reports them as expired, so deploys never fail on a borderline token.
|
||||
3. Writes a temporary Docker config file with all registry auth entries.
|
||||
4. Sets the `DOCKER_CONFIG` environment variable so `docker compose` uses the temporary config.
|
||||
1. Resolves credentials for every configured registry.
|
||||
2. For ECR, reuses a cached authorization token when one is still valid; otherwise fetches a fresh one from AWS. Cached tokens are refreshed a few minutes before the AWS-reported expiry, so deploys do not fail on a borderline token.
|
||||
3. Writes a temporary Docker config file containing all the resolved auth entries.
|
||||
4. Sets the `DOCKER_CONFIG` environment variable so `docker compose` reads from the temporary config instead of the host's `~/.docker/config.json`.
|
||||
5. Runs the compose operation (pull and/or up).
|
||||
6. Cleans up the temporary config file immediately after.
|
||||
6. Deletes the temporary config immediately after the operation finishes.
|
||||
|
||||
This approach ensures credentials are never persisted on disk beyond the duration of the operation.
|
||||
Credentials are never persisted to disk on the host beyond the duration of a single operation.
|
||||
|
||||
If a stored secret cannot be decrypted at deploy time (for example, because the encryption key file was replaced), Sencho skips that one registry, records a warning in the deploy log stream prefixed with `[Sencho] Warning:`, and continues. This lets public-image deploys succeed even when one private registry is misconfigured, while making the failure visible so you can re-save the credentials.
|
||||
If a stored secret cannot be decrypted at deploy time (for example because the encryption key file was replaced), Sencho skips that one registry, writes a warning to the deploy log stream prefixed with `[Sencho] Warning:`, and continues. Public-image deploys still succeed when one private registry is misconfigured, and the failure is visible so the operator can re-save the credentials.
|
||||
|
||||
### Image update checks
|
||||
|
||||
Sencho's background image update checker also uses stored registry credentials. When checking for newer image versions, it passes your credentials to the registry's authentication endpoint so it can compare local and remote digests for private images.
|
||||
The background image-update checker uses the same stored credentials. When it polls a registry to compare local and remote digests for a private image, it authenticates with the registry-specific credentials configured here.
|
||||
|
||||
## ECR setup
|
||||
|
||||
AWS ECR uses short-lived authentication tokens (valid for 12 hours) derived from your IAM credentials. Sencho handles this automatically:
|
||||
AWS ECR uses short-lived authentication tokens (valid for 12 hours) derived from IAM credentials. Sencho handles the token lifecycle automatically:
|
||||
|
||||
1. Store your **AWS Access Key ID** and **Secret Access Key** as the username and secret.
|
||||
2. Specify the **AWS Region** where your ECR registry lives.
|
||||
3. On every deploy or pull, Sencho reuses a cached authorization token when one is still valid, or calls the AWS `GetAuthorizationToken` API to obtain a fresh one.
|
||||
1. Store the **AWS Access Key ID** and **Secret Access Key** as the credential pair.
|
||||
2. Set the **AWS Region** where the registry lives.
|
||||
3. On every deploy or pull, Sencho reuses a cached authorization token while it is valid, or calls the AWS `GetAuthorizationToken` API to obtain a fresh one.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/private-registries/registries-ecr-form.png" alt="Add registry form with type set to AWS Elastic Container Registry (ECR), showing the relabelled AWS Access Key ID and AWS Secret Access Key inputs and the AWS Region field" />
|
||||
</Frame>
|
||||
|
||||
<Warning>
|
||||
Use an IAM user or role with only the `ecr:GetAuthorizationToken` and `ecr:BatchGetImage` permissions. Avoid using root account credentials.
|
||||
Use an IAM user or role limited to the ECR read permissions below. Do not paste root account credentials.
|
||||
</Warning>
|
||||
|
||||
### IAM policy example
|
||||
@@ -121,49 +144,54 @@ AWS ECR uses short-lived authentication tokens (valid for 12 hours) derived from
|
||||
|
||||
| Registry | URL to use |
|
||||
|----------|-----------|
|
||||
| Docker Hub | `https://index.docker.io/v1/` |
|
||||
| Docker Hub | `https://index.docker.io/v1/` (set automatically) |
|
||||
| GHCR | `ghcr.io` |
|
||||
| AWS ECR | `{account_id}.dkr.ecr.{region}.amazonaws.com` |
|
||||
| Self-hosted | Your registry hostname, e.g. `registry.example.com` |
|
||||
|
||||
## Security
|
||||
|
||||
- **Encrypted storage** - Registry secrets are encrypted at rest, using the same encryption layer as remote node tokens and SSO secrets.
|
||||
- **No persistent Docker login** - Credentials are written to a temporary file for the duration of each operation and immediately deleted afterward.
|
||||
- **Secrets never exposed** - The API never returns decrypted secrets. The UI shows only whether a secret is stored.
|
||||
- **Audit trail** - Registry create, update, and delete operations are recorded in the [Audit Log](/features/audit-log).
|
||||
- **Admin-only access** - Only admin users with an Admiral license can manage registry credentials.
|
||||
- **Encrypted storage.** Registry secrets are encrypted at rest with the same encryption layer used for remote-node tokens and SSO secrets.
|
||||
- **No persistent Docker login.** Credentials are written to a temporary file for the duration of each compose operation and deleted afterward.
|
||||
- **Secrets never exposed.** The API never returns decrypted secrets. The UI shows only whether a secret is stored.
|
||||
- **Admin role required.** Registry management is restricted to admin operators, even on Admiral. Viewers and operators with non-admin roles cannot see the section.
|
||||
- **API tokens cannot manage registries.** Registry credentials can only be created, edited, or deleted from an admin browser session. Automation tokens are scoped away from this surface so a leaked CI key cannot rewrite pull credentials.
|
||||
- **Audit trail.** Registry create, update, and delete operations are recorded in the [Audit Log](/features/audit-log).
|
||||
|
||||
## Multi-node behavior
|
||||
## Fleet-wide application
|
||||
|
||||
Registry credentials are stored per Sencho instance. When managing remote nodes, each node runs its own Sencho instance with its own registry credentials. Configure private registries on each node that needs access to private images.
|
||||
Registry credentials live on the control instance and are applied to every node it manages. When the control deploys to a remote node, the resolved Docker config is injected through the same secure channel used for the deploy itself; remote nodes never store registry credentials locally, and the Registries section is hidden from their own Settings pane.
|
||||
|
||||
That means you configure each private registry once on the control instance and every stack across the fleet picks up the credentials automatically.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### Pull fails with a 401 even though the credentials saved successfully
|
||||
<AccordionGroup>
|
||||
<Accordion title="Pull fails with a 401 even though the credentials saved successfully">
|
||||
The most common cause is a URL format mismatch. Use the values from the **Registry URL reference** table above. For Docker Hub the value is set automatically; for GHCR use `ghcr.io` exactly. If the URL is correct, open the registry, re-enter the secret, and click **Test connection** inside the form. Some providers rotate tokens silently, and the card view cannot tell you a stored token was revoked without an active probe.
|
||||
</Accordion>
|
||||
|
||||
The most common cause is a URL format mismatch. Use the forms in the **Registry URL reference** table above, and note that the URL is case-sensitive for most registries. For Docker Hub, any URL you enter is stored as the canonical `https://index.docker.io/v1/` because that is the key the Docker CLI expects.
|
||||
<Accordion title="An ECR registry works for a while and then stops pulling">
|
||||
AWS ECR authorization tokens are valid for 12 hours. Sencho caches the token in memory and refreshes it automatically a few minutes before the AWS-reported expiry, so no operator action is needed in steady state. If pulls still fail, confirm the IAM user still has `ecr:GetAuthorizationToken` and that the AWS Access Key ID has not been rotated. Click **Test connection** on the card to force a fresh token fetch.
|
||||
</Accordion>
|
||||
|
||||
If the URL is correct, re-open the registry, re-enter the secret, and click **Test connection** inside the form. Some providers rotate tokens silently, and the card view cannot tell you that the stored token was revoked without attempting a live probe.
|
||||
<Accordion title='The deploy log shows `[Sencho] Warning: Registry "X" credentials unavailable`'>
|
||||
Sencho could not decrypt the stored secret for that registry. The deploy continues without that registry's credentials, which is fine if the stack pulls only public images. To fix it, open the registry, re-enter the secret, and save. If the same warning appears for every registry at once, the encryption key file in your data directory has been replaced or lost and every stored secret needs to be re-entered.
|
||||
</Accordion>
|
||||
|
||||
### An ECR registry works for a while and then stops pulling
|
||||
<Accordion title="Test connection says it failed but deploys still pull images successfully">
|
||||
Some registries, notably certain self-hosted mirrors and proxy caches, restrict access to the `/v2/` discovery endpoint while still allowing pulls. Sencho uses `/v2/` as the probe target, so a failure there does not always indicate broken credentials. If your deploys succeed, the test result can be ignored. The probe is a best-effort check, not a save-time gate.
|
||||
</Accordion>
|
||||
|
||||
AWS ECR authorization tokens are valid for 12 hours. Sencho caches the token in memory and refreshes it automatically a few minutes before the AWS-reported expiry, so no action is needed on your part.
|
||||
<Accordion title="The Registries section is not visible">
|
||||
The section is shown only on the control instance, and only when the control's license is Admiral and the signed-in operator has the admin role. On a remote node viewed through the node switcher, the section is hidden by design; manage registries on the control instead. If a non-admin operator should be able to manage registries, change their role under **Settings → Identity → Users** first.
|
||||
</Accordion>
|
||||
|
||||
If pulls still fail, confirm the IAM user still has the `ecr:GetAuthorizationToken` permission and that the AWS Access Key ID has not been rotated. Click **Test connection** on the registry card to force a fresh token fetch.
|
||||
<Accordion title="A node displays a 'Private Registries is not available on this node' lock card">
|
||||
The selected node is on a Sencho version too old to surface this feature. Upgrade the node to a current Sencho release. Until then, credentials configured on the control still apply to that node's deploys through the injection path; only the management UI is hidden.
|
||||
</Accordion>
|
||||
|
||||
### The deploy log shows `[Sencho] Warning: Registry "X" credentials unavailable`
|
||||
|
||||
Sencho could not decrypt the stored secret for that registry. The deploy continues without that registry's credentials, which is fine if the stack pulls only public images.
|
||||
|
||||
To fix the registry, open it in the Registries tab, re-enter the secret, and save. If the same warning appears for every registry, the encryption key file at your data directory has been replaced or lost, and every stored secret needs to be re-entered.
|
||||
|
||||
### Test connection says it failed, but deploys still pull images successfully
|
||||
|
||||
Some registries (notably certain self-hosted mirrors and proxy caches) restrict access to the `/v2/` discovery endpoint while still allowing pulls. Sencho uses `/v2/` as the probe target, so a failure there does not always indicate broken credentials.
|
||||
|
||||
If your deploys succeed, you can safely ignore the test result. The test is a best-effort check, not a gate on saving.
|
||||
|
||||
### Credentials work on the local node but not on a remote node
|
||||
|
||||
Registry credentials are stored per Sencho instance, not shared across nodes. Open the remote node from the node switcher and configure the same registry on that node's Registries tab.
|
||||
<Accordion title="An API token cannot create or update registries">
|
||||
Registry management is restricted to admin browser sessions; API tokens are not scoped for this surface. Use a signed-in browser for the create, edit, and delete flow and reserve API tokens for the operations they are designed for.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
@@ -1,129 +1,245 @@
|
||||
---
|
||||
title: RBAC & User Management
|
||||
description: Role-based access control for Sencho - manage admin, viewer, deployer, node admin, and auditor accounts with scoped permissions.
|
||||
description: Role-based access control for Sencho. Manage admin, viewer, deployer, node admin, and auditor accounts with scoped permissions, MFA, and SSO.
|
||||
---
|
||||
|
||||
<Note>
|
||||
Multi-user support requires a **Sencho Skipper** or **Admiral** license. Community Edition supports a single admin account only. Intermediate roles (Deployer, Node Admin, Auditor) and scoped permissions require **Admiral**.
|
||||
Multi-user support requires a Sencho **Skipper** or **Admiral** license. Community Edition runs as a single admin account. The **Deployer**, **Node Admin**, and **Auditor** roles, plus scoped permissions, require **Admiral**.
|
||||
</Note>
|
||||
|
||||
Sencho supports role-based access control with five distinct roles. **Admin** and **Viewer** are available on all paid tiers, while **Deployer**, **Node Admin**, and **Auditor** are exclusive to Admiral.
|
||||
<Note>
|
||||
The Users panel is hub-only. It does not appear under Settings on a remote node, since accounts and roles are managed from the gateway that proxies the fleet.
|
||||
</Note>
|
||||
|
||||
Sencho ships with five built-in roles that map to the permissions most operators reach for: full operator access, read-only observation, day-to-day deploys, scoped fleet management, and audit-only compliance. Admiral adds **scoped permissions** so you can grant a viewer the right to deploy one specific stack without elevating them anywhere else.
|
||||
|
||||
## Roles
|
||||
|
||||
| Role | Description | Tier |
|
||||
|------|-------------|------|
|
||||
| **Admin** | Full access to all features: deploy, edit, manage users, configure nodes, and system settings | Skipper+ |
|
||||
| **Viewer** | Read-only access to dashboards, logs, stats, and file contents | Skipper+ |
|
||||
| **Deployer** | Can deploy, restart, stop, and start stacks, but cannot edit compose files, delete stacks, or access system settings | Admiral |
|
||||
| **Node Admin** | Full stack and node management within their scope, but no access to system settings, users, or license management | Admiral |
|
||||
| **Auditor** | Read-only access like Viewer, plus access to the audit log | Admiral |
|
||||
| Role | What it grants | Tier |
|
||||
|------|----------------|------|
|
||||
| **Admin** | Full operator access: deploy, edit compose, manage users, configure nodes, view audit log, every system setting | Skipper+ |
|
||||
| **Viewer** | Read-only access to stacks, logs, stats, file contents, and node listings | Skipper+ |
|
||||
| **Deployer** | Deploy, restart, stop, and start stacks. Cannot edit compose files, create or delete stacks, or view nodes | Admiral |
|
||||
| **Node Admin** | Full stack and node management across the fleet. No access to system settings, users, or license | Admiral |
|
||||
| **Auditor** | Read-only access to stacks, nodes, and the audit log. No write access anywhere | Admiral |
|
||||
|
||||
### Permission matrix
|
||||
|
||||
| Action | Admin | Node Admin | Deployer | Auditor | Viewer |
|
||||
|--------|-------|------------|----------|---------|--------|
|
||||
| View stacks, logs, stats | Yes | Yes | Yes | Yes | Yes |
|
||||
| Deploy / restart / stop / start stacks | Yes | Yes | Yes | No | No |
|
||||
| Edit compose and `.env` files | Yes | Yes | No | No | No |
|
||||
| Create and delete stacks | Yes | Yes | No | No | No |
|
||||
| View nodes | Yes | Yes | Yes | Yes | Yes |
|
||||
| Add / edit / delete nodes | Yes | Yes | No | No | No |
|
||||
| Audit log | Yes | No | No | Yes | No |
|
||||
| System settings | Yes | No | No | No | No |
|
||||
| User management | Yes | No | No | No | No |
|
||||
| License management | Yes | No | No | No | No |
|
||||
| Webhooks | Yes | No | No | No | No |
|
||||
| API tokens | Yes | No | No | No | No |
|
||||
| Host console | Yes | No | No | No | No |
|
||||
Each row is one of the permission keys the backend checks. The matrix below is the source of truth for what each role can do globally. Admin always grants every permission; the other roles are explicit.
|
||||
|
||||
## Managing users
|
||||
| Permission | Admin | Node Admin | Deployer | Auditor | Viewer |
|
||||
|------------|:-----:|:----------:|:--------:|:-------:|:------:|
|
||||
| View stacks, logs, stats (`stack:read`) | Yes | Yes | Yes | Yes | Yes |
|
||||
| Deploy, restart, stop, start (`stack:deploy`) | Yes | Yes | Yes | No | No |
|
||||
| Edit compose and `.env` files (`stack:edit`) | Yes | Yes | No | No | No |
|
||||
| Create stacks (`stack:create`) | Yes | Yes | No | No | No |
|
||||
| Delete stacks (`stack:delete`) | Yes | Yes | No | No | No |
|
||||
| View nodes (`node:read`) | Yes | Yes | No | Yes | Yes |
|
||||
| Add, edit, delete, cordon nodes (`node:manage`) | Yes | Yes | No | No | No |
|
||||
| Audit log (`system:audit`) | Yes | No | No | Yes | No |
|
||||
| System settings (`system:settings`) | Yes | No | No | No | No |
|
||||
| User management (`system:users`) | Yes | No | No | No | No |
|
||||
| License management (`system:license`) | Yes | No | No | No | No |
|
||||
| Webhooks (`system:webhooks`) | Yes | No | No | No | No |
|
||||
| API tokens (`system:tokens`) | Yes | No | No | No | No |
|
||||
| Host console (`system:console`) | Yes | No | No | No | No |
|
||||
| Container registries (`system:registries`) | Yes | No | No | No | No |
|
||||
|
||||
Go to **Settings > Users** to view the user list. The table shows each user's username, role badge, creation date, and action buttons.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/rbac/users-list.png" alt="User Management showing the users table with username, role, and creation date" />
|
||||
</Frame>
|
||||
|
||||
From the user table you can:
|
||||
|
||||
- **Edit** a user by clicking the pencil icon
|
||||
- **Delete** a user by clicking the trash icon (with a confirmation dialog)
|
||||
|
||||
You cannot delete your own account. The delete button is disabled for the currently logged-in user.
|
||||
|
||||
### Creating a user
|
||||
|
||||
Click **Add User** to open the creation form.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/rbac/user-create-form.png" alt="New User form with Username, Role, Password, and Confirm Password fields" />
|
||||
</Frame>
|
||||
|
||||
| Field | Description |
|
||||
|-------|-------------|
|
||||
| **Username** | At least 3 characters. Letters, numbers, underscores, and hyphens only. |
|
||||
| **Role** | Select from the available roles (see below) |
|
||||
| **Password** | At least 8 characters |
|
||||
| **Confirm Password** | Must match the password field |
|
||||
|
||||
On Admiral, all five roles appear in the role selector. On Skipper, only Admin and Viewer are available.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/rbac/role-selector-dropdown.png" alt="Role selector dropdown showing all five roles on Admiral" />
|
||||
</Frame>
|
||||
|
||||
### Editing a user
|
||||
|
||||
Click the pencil icon on a user row to edit. You can change the username, role, and optionally set a new password. Leave the password fields blank to keep the existing password.
|
||||
|
||||
## Scoped permissions
|
||||
|
||||
<Note>
|
||||
Scoped permissions require an **Admiral** license.
|
||||
</Note>
|
||||
|
||||
Roles can be scoped to specific stacks or nodes. This lets you grant a user elevated permissions on particular resources without giving them broad access.
|
||||
|
||||
Scoped permissions **add to** the user's global role. They never reduce it. A user with a global Viewer role plus a scoped Deployer assignment on the "my-app" stack can deploy "my-app" but has read-only access to everything else.
|
||||
|
||||
When editing a user on Admiral, a **Scoped Permissions** section appears below the user form. From there you can:
|
||||
|
||||
1. **View** the user's current scoped assignments, each showing the role badge, resource type, and resource name
|
||||
2. **Add** a new scope by selecting a role (Deployer, Node Admin, or Admin), resource type (Stack or Node), and the specific resource
|
||||
3. **Remove** an existing scope with the trash icon
|
||||
|
||||
### Example scenarios
|
||||
|
||||
- A **Viewer** with a scoped **Deployer** assignment on the `frontend` stack can deploy, restart, and stop only that stack.
|
||||
- A **Deployer** with a scoped **Node Admin** assignment on node "staging-server" can manage stacks and nodes on that server, plus deploy globally.
|
||||
- A **Node Admin** without any scoped assignments can manage all stacks and nodes but has no access to system settings.
|
||||
|
||||
## SSO auto-provisioning
|
||||
|
||||
With an Admiral license, users can also be created automatically when they log in via SSO (LDAP, Google, GitHub, or Okta). SSO users appear in the Users list alongside local accounts and are assigned a role based on identity provider group membership or claim mapping.
|
||||
|
||||
SSO users cannot log in with a password; they must always authenticate through their identity provider. After SSO provisioning, an admin can add scoped permissions to SSO users just like local accounts.
|
||||
|
||||
To set up identity provider authentication, see [SSO Authentication](/features/sso).
|
||||
|
||||
## Session security
|
||||
|
||||
Sencho enforces user changes immediately:
|
||||
|
||||
- **Account deletion**: A deleted user's active sessions are rejected on the next request. There is no delay or grace period.
|
||||
- **Role changes**: When an admin changes a user's role, the new permissions take effect immediately for all of that user's active sessions.
|
||||
- **Password changes**: Changing a password (either your own or as an admin reset) invalidates all other active sessions for that user. The session that performed the change remains valid.
|
||||
- **SSO users**: Password fields are hidden for SSO-provisioned users. SSO accounts always authenticate through their identity provider.
|
||||
|
||||
## Migration from single-admin setup
|
||||
|
||||
When you upgrade to a Skipper or Admiral license, your existing single-admin credentials are automatically migrated. No manual action is required; your login continues to work as before, and your account is assigned the Admin role.
|
||||
On Admiral, a user with a lower global role can still hold extra permissions on specific stacks or nodes through scoped assignments (covered below). Scoped permissions are additive: they grant more, never less.
|
||||
|
||||
## Account limits by tier
|
||||
|
||||
| Tier | Admin accounts | Non-admin accounts | Intermediate roles | Scoped permissions |
|
||||
|------|---------------|-------------------|-------------------|-------------------|
|
||||
|------|---------------|--------------------|--------------------|--------------------|
|
||||
| **Community** | 1 | 0 | No | No |
|
||||
| **Skipper** | 1 | 3 | No | No |
|
||||
| **Admiral** | Unlimited | Unlimited | Yes | Yes |
|
||||
|
||||
Quotas are enforced when you click **Create user**. Hitting a cap returns a `403` with a clear message, and the form keeps your input so you can adjust the role.
|
||||
|
||||
## Managing users
|
||||
|
||||
The Users panel lives at **Settings · Users**, under the **Identity** group of the settings sidebar. It is visible only to users with the Admin role on Skipper or Admiral, and is hidden when a remote node is the active selection.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/rbac/users-list.png" alt="Settings Users panel showing the Identity sidebar selection, a SCOPE operator chip, an OPERATORS 2 counter, an Add user button, and a four-column table (Username, Role, Created, Actions) with two rows: admin marked (you) with a disabled trash icon and viewer with active edit and trash icons." />
|
||||
</Frame>
|
||||
|
||||
The table shows one row per user with their **Username**, **Role** badge, account **Created** date, and per-row action icons. The signed-in admin's row carries a small `(you)` marker after the username and the delete icon is disabled, so you cannot lock yourself out by deleting your own account.
|
||||
|
||||
Per-row actions, left to right:
|
||||
|
||||
- **Edit** (pencil icon): open the user in the inline edit form.
|
||||
- **Reset 2FA** (shield icon, warning color): appears only on users that have enrolled in two-factor authentication. Opens the reset confirmation dialog described in [Two-factor reset](#two-factor-reset).
|
||||
- **Delete** (trash icon, destructive color): open the delete confirmation. Disabled on your own row.
|
||||
|
||||
### Creating a user
|
||||
|
||||
Click **Add user**. The form opens inline below the button (it is not a modal), with the heading **New User** and four fields laid out in two columns.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/rbac/add-user-form.png" alt="Users panel with the New User form expanded inline above the table. The form shows Username and Role fields on the top row (Username placeholder username; Role combobox defaulting to Viewer), Password and Confirm Password fields below (Password placeholder min. 8 characters), and Cancel and CREATE USER buttons on the right." />
|
||||
</Frame>
|
||||
|
||||
| Field | Rules |
|
||||
|-------|-------|
|
||||
| **Username** | At least 3 characters. Letters, numbers, underscores, and hyphens only. Submitting with `.` or whitespace returns `Username can only contain letters, numbers, underscores, and hyphens.` |
|
||||
| **Role** | Combobox. On Admiral you see all five roles; on Skipper you see Admin and Viewer only. |
|
||||
| **Password** | At least 8 characters. The placeholder reads `min. 8 characters`. |
|
||||
| **Confirm Password** | Must match the password field, validated on submit. |
|
||||
|
||||
The role combobox on Admiral exposes the full set:
|
||||
|
||||
<Frame>
|
||||
<img src="/images/rbac/role-selector.png" alt="Role combobox open inside the New User form on an Admiral instance, listing Admin, Viewer (checkmark), Deployer, Node Admin, and Auditor as selectable options." />
|
||||
</Frame>
|
||||
|
||||
Click **Create user** to submit. The form clears, the table refreshes, and an audit-log entry is written with the actor, target username, and assigned role.
|
||||
|
||||
### Editing a user
|
||||
|
||||
Click the pencil icon on a row to switch the form into **Edit User** mode. The same four fields render with the existing values pre-filled, plus a separate **Scoped Permissions** box below for Admiral instances (see [Scoped permissions](#scoped-permissions)).
|
||||
|
||||
The password fields change subtly in edit mode:
|
||||
|
||||
- The **Password** label becomes **New Password (optional)** with the placeholder `Leave blank to keep`. Submit without filling them in and the current password is preserved.
|
||||
- If the user was provisioned via SSO, the password fields are replaced with an inline line that reads `Password is managed by the identity provider (<provider>).` Sencho never stores or rotates passwords for SSO accounts.
|
||||
|
||||
Click **Update user** to save. Changing the role takes effect on the next API request from any of that user's active sessions; see [Session security](#session-security) below.
|
||||
|
||||
## Scoped permissions
|
||||
|
||||
<Note>
|
||||
Scoped permissions require **Admiral**.
|
||||
</Note>
|
||||
|
||||
Scoped permissions let you grant a user a higher role on a specific stack or node without elevating them globally. A Viewer can be granted Deployer on one stack; a Deployer can be granted Node Admin on one server.
|
||||
|
||||
The box appears below the user form whenever you are editing a user on Admiral.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/rbac/scoped-permissions.png" alt="Edit User form for the viewer account with a Scoped Permissions box below. The box contains an existing assignment row (a Deployer badge with the text on Stack: bazarr and a destructive trash icon on the right) and a three-column add-scope row underneath (Role combobox set to Deployer, Resource Type combobox set to Stack, Resource combobox showing Select..., and a disabled Add button)." />
|
||||
</Frame>
|
||||
|
||||
The add-scope form has three controls and an **Add** button:
|
||||
|
||||
| Control | Options |
|
||||
|---------|---------|
|
||||
| **Role** | Deployer, Node Admin, or Admin. The scoped role picker is narrower than the global role picker. Viewer and Auditor cannot be scoped (they are floor-only roles). |
|
||||
| **Resource Type** | `Stack` or `Node`. |
|
||||
| **Resource** | The picker shows stacks (when type is `Stack`) or remote nodes (when type is `Node`) the gateway knows about. Resource names match what you see in the sidebar. |
|
||||
|
||||
Click **Add** to save the assignment. Existing scopes render as a row with the role badge, the line `on <type>: <resource>`, and a trash icon for removal. Removing an assignment is instant; the user's effective permissions are recomputed on their next request.
|
||||
|
||||
Scoped assignments are **additive only**. A Viewer with a scoped Deployer on `frontend` can deploy `frontend` but stays read-only on every other resource. Scopes never reduce the global role.
|
||||
|
||||
### Example scenarios
|
||||
|
||||
- A **Viewer** with a scoped **Deployer** assignment on the `frontend` stack can deploy, restart, and stop only that stack. They cannot edit compose or delete it.
|
||||
- A **Deployer** with a scoped **Node Admin** assignment on node `staging-server` can manage every stack and node operation on that server, while keeping plain Deployer rights on the rest of the fleet.
|
||||
- A **Node Admin** without any scoped assignments has full stack and node management across every node, but still cannot reach system settings, the user list, or the audit log.
|
||||
|
||||
## Two-factor reset
|
||||
|
||||
Admins can reset a user's two-factor authentication enrollment. Use this only when the user has lost access to their authenticator app or backup codes.
|
||||
|
||||
The shield icon next to the pencil icon appears for any user who has finished TOTP enrollment. Clicking it opens a confirmation modal with the kicker `USERS · RESET 2FA`, the title `Reset 2FA for <username>`, and the body:
|
||||
|
||||
> Removes the user's authenticator enrolment and backup codes. They will sign in with just their password on their next login and can re-enrol from their account settings. Use this when a user has lost access to their authenticator.
|
||||
|
||||
Click **Reset 2FA** to confirm or **Cancel** to dismiss. Confirming bumps the user's session token version, which signs out every other active session of theirs on the next request, and writes an audit-log entry.
|
||||
|
||||
### Lockout after failed two-factor attempts
|
||||
|
||||
If a user submits five wrong codes during sign-in (TOTP or backup), the account is locked for **15 minutes**. The lockout applies to sign-in only; sessions already authenticated continue to work. An admin **Reset 2FA** clears the lockout by removing the enrollment.
|
||||
|
||||
If you need to track repeated lockouts, the audit log records each failed login attempt and the lockout state changes.
|
||||
|
||||
## Deleting a user
|
||||
|
||||
Click the trash icon on a user row. A confirmation modal with the kicker `USERS · DELETE · IRREVERSIBLE` opens.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/rbac/delete-confirm.png" alt="Delete confirmation modal with the kicker USERS DELETE IRREVERSIBLE, the title Delete user viewer in italic display type, the body Removes the user immediately. They lose access right away., and Cancel and Delete buttons on the lower right." />
|
||||
</Frame>
|
||||
|
||||
The body reads:
|
||||
|
||||
> Removes the user immediately. They lose access right away.
|
||||
|
||||
This is literal. Deletion bumps the user's token version, so every active session of theirs returns `401` on their next API request. There is no grace period and no soft-delete table. The deletion is also written to the audit log.
|
||||
|
||||
You cannot delete your own account; the trash icon on the `(you)` row is disabled.
|
||||
|
||||
## Session security
|
||||
|
||||
Sencho enforces user changes immediately by versioning JWT tokens at the user record. Each token carries the user's `tv` (token-version) claim; mismatches return `401` on the next request without consulting any session table.
|
||||
|
||||
| Event | Effect on active sessions |
|
||||
|-------|---------------------------|
|
||||
| **Account deletion** | Token version bump. Every existing session of the deleted user returns `401` on its next request. |
|
||||
| **Role change** | The database role wins. The user's existing JWTs continue to validate, but the new role is read from the user record on every request, so the new permissions take effect immediately. |
|
||||
| **Password change (self)** | Token version bump. All **other** sessions of the user are invalidated. The session that performed the change is reissued a fresh token, so the user stays signed in where they made the change. |
|
||||
| **Two-factor reset (admin)** | Token version bump. Every session of the target user is invalidated. |
|
||||
|
||||
Cookies and Bearer tokens go through the same auth middleware, so the same rules apply to API-token-based sessions where a token is bound to a user.
|
||||
|
||||
## SSO auto-provisioning
|
||||
|
||||
With SSO configured on Admiral, users authenticate through an identity provider (LDAP, Custom OIDC, Google, GitHub, Okta). On their first successful sign-in, Sencho auto-creates a user record. SSO accounts appear in the Users list alongside local accounts and can be edited the same way; only the password and (optionally) the role differ.
|
||||
|
||||
Two SSO-specific behaviors to keep in mind:
|
||||
|
||||
- **Password fields are hidden when editing an SSO user.** The form shows `Password is managed by the identity provider (<provider>)` in place of the password inputs. SSO users always authenticate through their IdP.
|
||||
- **Optional MFA enforcement.** Each SSO provider config exposes a `Require MFA` toggle. Off (default), SSO users are not required to enroll in TOTP. On, every SSO-provisioned user must enroll TOTP after their first successful sign-in before they can use the rest of the console.
|
||||
|
||||
The role assigned at provisioning is the role configured on the SSO provider (or, for LDAP, derived from group membership). After provisioning, an admin can adjust the role and add scoped permissions just like any local account.
|
||||
|
||||
To configure a provider, see [SSO Authentication](/features/sso). The tier split for provider configuration (Custom OIDC at Community, preset providers at Skipper, LDAP at Admiral) is enforced separately from the rest of the user-management surface.
|
||||
|
||||
## API tokens for automation
|
||||
|
||||
Users are for humans; **API tokens** are for machines. Sencho issues opaque tokens with the `sen_sk_` prefix and three scopes (Read Only, Deploy Only, Full Admin) that can authenticate CI pipelines, monitoring loops, or deployment scripts without going through the user table at all. Tokens cannot reach the user-management endpoints regardless of scope.
|
||||
|
||||
See [API Tokens](/features/api-tokens) for the full reference.
|
||||
|
||||
## Audit logging
|
||||
|
||||
Every user-management mutation is recorded in the audit log:
|
||||
|
||||
- `Create user <username>` with the assigned role
|
||||
- `Update user <username>` with the changed fields
|
||||
- `Delete user <username>`
|
||||
- `Reset 2FA for <username>`
|
||||
- Scoped role assignments (`POST /api/users/:id/roles`) and removals (`DELETE /api/users/:id/roles/:assignmentId`)
|
||||
|
||||
Entries include the acting user, IP address, HTTP method and path, response status, and a human-readable summary. Sign-in events and failed-MFA attempts are recorded in the same trail. See [Audit Log](/features/audit-log) for the full schema and filter options.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="The Users entry is missing from the Settings sidebar">
|
||||
The Users entry is hidden in three cases. **One,** the active license is Community: the Users panel is gated to Skipper+ and does not render on Community. Activate a Skipper or Admiral license under **Settings · License** to expose it. **Two,** you are signed in as a non-admin (Viewer, Deployer, Auditor): the entry is admin-only. **Three,** you have a remote node selected: the panel is hub-only and is hidden in the sidebar when any remote node is active. Switch back to the local node via the node switcher in the masthead.
|
||||
</Accordion>
|
||||
<Accordion title="The role I want is greyed out in the role combobox">
|
||||
The combobox only shows roles available on your tier. On Skipper, the combobox lists Admin and Viewer only. **Deployer**, **Node Admin**, and **Auditor** are Admiral-only roles and do not appear on Skipper. Upgrade to Admiral, or use scoped permissions equivalents once you do.
|
||||
</Accordion>
|
||||
<Accordion title="Creating a user fails with `Your license allows a maximum of N account(s)`">
|
||||
You have hit the seat limit for your tier. Skipper allows one admin and three non-admin users; Admiral has no cap. Either delete an unused account or upgrade. The exact remaining capacity is visible on the OPERATORS counter in the panel header.
|
||||
</Accordion>
|
||||
<Accordion title="A user complains they were signed out unexpectedly">
|
||||
Token-version bumps invalidate sessions. Two events do this: an admin changed the user's password, or an admin reset their 2FA. Both rotate the user's token version, so every JWT issued before the rotation is rejected on the next request. The user can sign in again with their (possibly new) password. Role changes do **not** sign the user out; they take effect on the next request without rotating the token version.
|
||||
</Accordion>
|
||||
<Accordion title="A scoped Deployer cannot deploy a stack they were granted">
|
||||
Two causes. **One,** the assignment was created on Admiral but the license has since dropped to Skipper. The permission resolver only consults scoped assignments when the effective tier is Admiral; on Skipper the scope is ignored and the user falls back to their global role. **Two,** the resource type or name on the assignment does not match the request's resource. Re-open the user in the edit form and check the existing-scope row matches the stack name (case-sensitive) exactly.
|
||||
</Accordion>
|
||||
<Accordion title="The shield (Reset 2FA) icon is missing on a user I expected to see it on">
|
||||
The icon only appears for users with a finished TOTP enrollment. If the user started enrollment but never confirmed their first code, the enrollment is incomplete and the icon stays hidden. Ask the user to finish enrollment from their account settings, or, if they cannot, leave the row alone: there is nothing to reset.
|
||||
</Accordion>
|
||||
<Accordion title="A locked-out user keeps re-locking after waiting 15 minutes">
|
||||
The 15-minute window expires on the clock, but the failure counter only resets on a successful sign-in. If the user retries with another wrong code after the window expires, the counter is still at five and the lockout re-engages immediately. Reset the user's 2FA from the row action to clear both the enrollment and the failure counter, then ask them to sign in with their password and re-enroll TOTP from their account settings.
|
||||
</Accordion>
|
||||
<Accordion title="An SSO user has the wrong role assigned at provisioning">
|
||||
The role assigned at first sign-in comes from the SSO provider configuration (group mapping for LDAP, claim mapping for OIDC). The user record already exists, so edit the role from **Settings · Users** for an immediate fix, and update the provider config under **Settings · SSO** to prevent the same drift on the next provisioning.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
@@ -1,117 +1,127 @@
|
||||
---
|
||||
title: Remote Updates
|
||||
description: Check for outdated nodes and trigger over-the-air Sencho updates from the Fleet View.
|
||||
description: Pull the latest Sencho image and recreate the container on any node in the fleet, including the gateway, from the Fleet view.
|
||||
---
|
||||
|
||||
Sencho can update remote nodes directly from the dashboard. When a node is running an older version than the latest available release, a one-click update pulls the latest image and recreates the container automatically. This includes the local (gateway) node itself.
|
||||
Sencho can update any node in the fleet to the latest published release without ever opening an SSH session. The control instance opens the **Node updates** sheet, dispatches the update to one or more nodes, and watches each one come back online with the new version.
|
||||
|
||||
This page covers the mechanism: prerequisites, what happens on a node during an update, how completion and failure are detected, and how to recover. The full UI tour for the Node updates sheet itself lives in [Fleet View](/features/fleet-view#node-updates).
|
||||
|
||||
<Note>
|
||||
Per-node remote updates and the Check Updates view are available on every tier (admin role required). The bulk **Update All** action is a Skipper or Admiral feature.
|
||||
Per-node updates and the Check Updates view are available on every tier. The bulk **Update all** action is a Skipper or Admiral feature. Triggering an update (per-row or bulk) requires the **admin** role; viewer and operator roles can read update status but cannot dispatch updates.
|
||||
</Note>
|
||||
|
||||
## Prerequisites
|
||||
|
||||
Remote updates work when each node meets these conditions:
|
||||
A node can self-update when all of the following are true:
|
||||
|
||||
- Deployed via **Docker Compose** (the canonical install for both primary instances and pilot agents)
|
||||
- The Docker socket (`/var/run/docker.sock`) is mounted into the container
|
||||
- Running a Sencho version that supports the `self-update` [capability](/features/node-compatibility)
|
||||
- It was deployed via **Docker Compose** with the published `saelix/sencho` image (the canonical install for both control instances and pilot agents).
|
||||
- The Docker socket (`/var/run/docker.sock`) is mounted into the container so Sencho can drive Compose from inside.
|
||||
- It runs a Sencho version that advertises the `self-update` capability. See [Node Compatibility](/features/node-compatibility) for the full capability list.
|
||||
|
||||
Nodes deployed by orchestrators like Kubernetes do not support self-update and need to be updated through the orchestrator's own deployment flow.
|
||||
Nodes deployed with `docker run`, with hand-rolled systemd units, or with orchestrators like Kubernetes do not advertise `self-update`. Their **Update** button returns an error and they are skipped by **Update all**.
|
||||
|
||||
## Checking for updates
|
||||
## Anatomy at a glance
|
||||
|
||||
Open the **Fleet** tab and click **Check Updates** in the header. This opens the **Node Updates** dialog, which lists every node with its current version and update status.
|
||||
Open the **Fleet** tab and click **Check Updates** in the page header to open the **Node updates** sheet. The sheet lists every registered node with its current and latest Sencho version, the cluster-wide summary, and the action buttons.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/fleet-view/fleet-node-updates.png" alt="Node Updates dialog showing update status for each node" />
|
||||
<img src="/images/fleet-view/fleet-node-updates.png" alt="Node updates sheet titled 'Node updates · 8 nodes · 8 updates available'. Recheck and 'Update all (7)' buttons sit under the header. Four summary cards read '0 Up to date', '8 Available', '0 Updating', '0 Failed'. A 'Filter nodes…' search box sits above a table with columns Node / Type / Current / Latest / Status. Eight rows show every node at v0.76.6 with an Update button on the right. The sencho-pilot-test row reports Current as 'unknown'. The footer reads 'LATEST VERSION v0.76.7'." />
|
||||
</Frame>
|
||||
|
||||
The dialog includes:
|
||||
For the per-control breakdown of the sheet (header actions, summary cards, table columns, footer), see [Fleet View · Node Updates](/features/fleet-view#node-updates).
|
||||
|
||||
- **Summary cards** at the top showing counts of nodes that are Up to date, have updates Available, are currently Updating, or have Failed
|
||||
- **Latest version** label showing the newest available Sencho release
|
||||
- **Filter** search box to find specific nodes by name or type
|
||||
- **Node table** with columns for name, type, current version, latest version, and status
|
||||
## Triggering a remote update
|
||||
|
||||
Each node's status column shows one of:
|
||||
There are two surfaces that initiate an update on a remote node:
|
||||
|
||||
| Status | Meaning |
|
||||
|--------|---------|
|
||||
| **Up to date** badge | Node is running the latest available version |
|
||||
| **Update** button | A newer version is available; click to update |
|
||||
| **Updating** badge | The node is pulling the new image and restarting |
|
||||
| **Updated** badge | The node came back online with the new version |
|
||||
| **Timed out** badge | The node did not come back within 5 minutes |
|
||||
| **Failed** badge | The update was rejected or the image pull failed on the remote host |
|
||||
|
||||
Nodes that are too old to report their version show "unknown" in the current version column. These nodes are treated as outdated.
|
||||
|
||||
### Updating state
|
||||
|
||||
When an update is in progress, the dialog shows a spinning "Updating" badge and the summary card count changes in real time. The Fleet View polls every 5 seconds while an update is active.
|
||||
- The **Update** button on a row inside the Node updates sheet.
|
||||
- The **Update to vX.Y.Z** outline button along the bottom of any online node card on the Fleet grid.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/fleet-view/fleet-node-updating.png" alt="Node Updates dialog showing a node in the updating state" />
|
||||
<img src="/images/fleet-view/node-card-update-available.png" alt="Node card for the Opsix remote showing an Online badge, the version chip 'v0.76.6', a warning 'Update available' pill, the Running / Stopped / Stacks counters (4 / 0 / 3), CPU / RAM / Disk usage bars, and an outline 'Update to v0.76.7' button across the bottom of the card." />
|
||||
</Frame>
|
||||
|
||||
## Updating a single node
|
||||
Either surface dispatches the same backend call. While the call is in flight, the button reads **Triggering...** with a spinning icon, then settles into the **Updating** badge once the gateway has accepted the request.
|
||||
|
||||
You can trigger an update in two ways:
|
||||
The gateway switches to a fast 5-second polling loop while any node is in the `Updating` state, so the badge advances in near real time without waiting for the next 30-second fleet refresh.
|
||||
|
||||
- Click the **Update** button next to a node in the Node Updates dialog
|
||||
- Click the **Update to vX.Y.Z** button directly on a node card in Fleet View
|
||||
## Updating the local (gateway) node
|
||||
|
||||
For remote nodes, the update happens in the background. The Fleet View automatically polls at a faster rate (every 5 seconds) while an update is in progress, so you can watch the status change in near real-time.
|
||||
Updating the gateway is special because the dashboard is hosted by the very container that is about to restart. Clicking **Update** on the local row, or **Update to vX.Y.Z** on the Local card, opens a confirmation step before anything happens on disk.
|
||||
|
||||
## Updating all nodes
|
||||
<Frame>
|
||||
<img src="/images/fleet-view/local-update-confirm.png" alt="Alert dialog with kicker 'LOCAL · UPDATE', heading 'Update local node', body 'Pulls the latest Sencho image and restarts the server. The dashboard briefly disconnects and reconnects automatically when the update completes.', and two buttons: Cancel and 'Update & restart'." />
|
||||
</Frame>
|
||||
|
||||
Click **Update All (N)** in the footer of the Node Updates dialog to trigger updates on all outdated remote nodes simultaneously. The local node is excluded from bulk updates to avoid losing dashboard connectivity. Nodes that previously failed or timed out are automatically retried in a bulk update.
|
||||
After **Update & restart** is confirmed:
|
||||
|
||||
## Local node updates
|
||||
1. The browser captures the gateway's current boot timestamp from `/api/health`.
|
||||
2. The server pulls the latest image and spawns a short-lived helper container that runs `docker compose up -d --force-recreate` against the host's compose working directory.
|
||||
3. A full-screen **Updating Sencho...** overlay takes over the browser tab. The overlay polls `/api/health` every 3 seconds and keeps the page from reloading until a *new* boot timestamp comes back, even if the API briefly responds during the pull.
|
||||
4. Once a fresh boot timestamp is reported, the overlay reloads the page on the new version.
|
||||
|
||||
When you update the local (gateway) node:
|
||||
|
||||
1. A confirmation dialog appears explaining that the dashboard will briefly disconnect.
|
||||
2. After confirming, the server pulls the latest image and restarts.
|
||||
3. A reconnecting overlay appears and polls the server every few seconds. It stays visible for the entire pull and restart cycle, even if the API briefly responds in between, so the page does not reload until the new container has actually booted.
|
||||
4. Once the new container reports a fresh boot timestamp, the page automatically reloads with the new version.
|
||||
|
||||
If the new container does not come up within 5 minutes, a timeout message appears with a manual reload option. If the gateway can detect that the update did not complete (for example, the image pull failed or the restart helper container could not spawn), the local node card surfaces a "Failed" badge with the underlying error within about 3 minutes instead of waiting for the full timeout.
|
||||
If the new container does not come up within 5 minutes, the overlay surfaces an **Update timed out** message with a *Try Reloading* button so the page is never left waiting indefinitely. If the gateway can detect that the update did not even start (for example, the image pull failed before the helper container could spawn) it surfaces a **Failed** badge with the underlying error on the Local card. The badge appears as soon as the helper writes its error file, or by the 3-minute mark at the latest, instead of waiting for the full 5-minute timeout.
|
||||
|
||||
<Note>
|
||||
The self-update helper container inherits all bind mounts from the main Sencho container. If your `docker-compose.yml` references `env_file`, `configs`, or `secrets` outside the compose working directory, those paths must be mounted into the Sencho container using the same host and container path (1:1 rule). See [Troubleshooting](/operations/troubleshooting#local-self-update-fails-with-env-file-not-found) if you encounter "env file not found" errors during a local update.
|
||||
The self-update helper container inherits all bind mounts from the main Sencho container 1:1. If your `docker-compose.yml` references `env_file`, `configs`, or `secrets` outside the compose working directory, those host paths must be mounted into the Sencho container at the *same container path* as on the host. See [Troubleshooting](/operations/troubleshooting#local-self-update-fails-with-env-file-not-found) if you encounter `env file not found` errors during a local update.
|
||||
</Note>
|
||||
|
||||
## What happens during an update
|
||||
|
||||
When an update is triggered on a node, Sencho:
|
||||
For both local and remote nodes, an update goes through the same three steps:
|
||||
|
||||
1. Pulls the latest `saelix/sencho` image using Docker Compose
|
||||
2. Recreates the container with the new image
|
||||
3. The node goes briefly offline during the restart
|
||||
1. **Pull** the latest `saelix/sencho` image from the registry.
|
||||
2. **Recreate** the container with the new image via `docker compose up -d --force-recreate`.
|
||||
3. **Restart** the Sencho process. The node is briefly offline during the swap.
|
||||
|
||||
The gateway monitors the remote node until it comes back online, then marks it as **Updated**. Completion is detected by three signals: a version change, a change in the node's process start time, or detecting that the node went briefly offline and came back (indicating a container restart). The "Updated" badge clears automatically after about 60 seconds and the node returns to "Up to date" status.
|
||||
The gateway tracks each in-flight update in memory and watches the target node for one of four completion signals, in this order of precedence:
|
||||
|
||||
If the node is still reachable and unchanged after about 90 seconds, the gateway marks the update as **Failed**. This usually means the image pull failed on the remote host. Check the Docker logs on the remote node for details.
|
||||
- The remote reports a *new* `version` value over its `/api/meta` endpoint.
|
||||
- The remote's process `startedAt` timestamp moves forward, indicating a fresh container start.
|
||||
- The remote went offline at any point during the watch window and has come back online.
|
||||
- More than 15 seconds have elapsed and the remote's reported version is at or above the comparison target (the gateway's own version, or the published latest, whichever is appropriate).
|
||||
|
||||
## Handling failures
|
||||
When any signal fires, the row flips from **Updating** to a green **Updated** badge. The **Updated** state stays visible for 60 seconds, then auto-clears so the row settles back to **Up to date**.
|
||||
|
||||
If an update times out or fails, the badge shows **Timed out** or **Failed** with two action buttons:
|
||||
## Failure detection and recovery
|
||||
|
||||
- **Retry** (circular arrow icon) clears the failed state and re-triggers the update
|
||||
- **Dismiss** (X icon) clears the failed state without retrying, returning the node to "Update available"
|
||||
The gateway uses two thresholds to decide when an update has failed:
|
||||
|
||||
Hovering over a failed or timed-out badge reveals the error message with details about what went wrong.
|
||||
| Threshold | What it means |
|
||||
|-----------|---------------|
|
||||
| **About 3 minutes** | If the remote is still reachable but no completion signal has fired, the gateway flips the row to **Failed**. This is the most common outcome of a registry pull error or a misconfigured Compose file: the remote answered the dispatch, started the pull, but never restarted. |
|
||||
| **5 minutes** | If the remote has been unreachable and has not come back, the gateway flips the row to **Timed out**. The image was almost certainly pulled, but the new container did not stay up. |
|
||||
|
||||
<Frame>
|
||||
<img src="/images/fleet-view/fleet-node-failed.png" alt="Node Updates dialog showing a failed node with retry and dismiss buttons" />
|
||||
</Frame>
|
||||
Both **Failed** and **Timed out** badges expose two inline actions:
|
||||
|
||||
You can also click **Recheck** in the dialog footer to clear all failed and timed-out states at once, refresh the cached latest version from GitHub, and fetch fresh version information from every node.
|
||||
- A circular-arrow **Retry** button that clears the failed state and re-dispatches the update against the same node.
|
||||
- An **X** **Dismiss** button that clears the failed state without retrying, returning the row to **Update available** so the operator can investigate before trying again.
|
||||
|
||||
Hovering either badge reveals the underlying error message reported by the remote (or by the helper container, for the gateway itself).
|
||||
|
||||
The header of the sheet exposes a **Recheck** button that does three things in one click: it flushes the cached "latest published version" lookup (resolved against the GitHub Releases API, with a Docker Hub fallback), it clears every terminal `Failed` and `Timed out` badge across the fleet, and it re-fetches the version metadata from every node.
|
||||
|
||||
<Note>
|
||||
Update tracking is stored in memory on the gateway. Restarting the gateway clears all update states, so any stuck "Timed out" or "Failed" badges will resolve on their own after a restart.
|
||||
Update tracking is held entirely in the gateway's memory. Restarting the gateway clears all in-flight, failed, and timed-out states automatically, so a stuck row always resolves itself on the next gateway restart even without a manual **Dismiss**.
|
||||
</Note>
|
||||
|
||||
If you run into issues with remote updates, see the [Troubleshooting](/operations/troubleshooting#remote-update-button-does-not-appear) page.
|
||||
## Troubleshooting
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Local update fails with 'env file not found'">
|
||||
The self-update helper container is launched with the same bind-mount set as the main Sencho container, but only paths that are explicitly mounted are visible inside the helper. If your `docker-compose.yml` references an `env_file` (or `configs`, or `secrets`) at a host path that is *not* mounted into the Sencho container, the helper cannot read it and the recreate step fails. Mount the referenced host directory into the Sencho container at the same container path (the 1:1 rule), redeploy Sencho once, and retry the update. The full walk-through with example mounts is in [Troubleshooting](/operations/troubleshooting#local-self-update-fails-with-env-file-not-found).
|
||||
</Accordion>
|
||||
<Accordion title="The reconnecting overlay never clears">
|
||||
The overlay polls `/api/health` every 3 seconds for up to 5 minutes. If it has not cleared by then, the new container is either still pulling a very large image, refusing to start, or crashing on boot. Open a host shell on the gateway and run `docker logs <sencho-container>` to inspect the boot output, then `docker ps` to confirm a new container ID is running. The overlay's *Try Reloading* button bypasses the timer once the container is healthy again.
|
||||
</Accordion>
|
||||
<Accordion title="A remote shows 'Failed' after about 3 minutes">
|
||||
The remote answered the dispatch but did not restart in time. The two most common causes are a registry pull error (rate limit, private-registry auth missing, network egress blocked) or a Compose file that fails validation under the new version. Open a shell on the remote and run `docker compose pull` followed by `docker compose up -d` against the Sencho working directory; the error printed there is the same one Sencho captured. Click **Retry** on the row once the cause is fixed.
|
||||
</Accordion>
|
||||
<Accordion title="A remote shows 'Timed out' after 5 minutes">
|
||||
The remote went offline during the swap and has not come back within the 5-minute watch window. The image almost certainly pulled (otherwise the row would have surfaced **Failed** earlier), but the new container is crashing on boot or is bound to a port that another process is now holding. Inspect `docker ps -a` on the remote for a recently exited Sencho container and `docker logs` it to see the crash. Once the remote answers `/api/health` again, click **Retry** to re-dispatch.
|
||||
</Accordion>
|
||||
<Accordion title="A node reports 'unknown' for its current version">
|
||||
The remote is reachable but its `/api/meta` response either does not include a `version` field or it cannot be parsed. The node is treated as outdated for the per-row **Update** button, so you can still trigger an update on it, but it is excluded from **Update all** because the gateway has no safe way to compare versions. After the update completes the version field comes back populated and the row joins **Update all** on the next dispatch.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
@@ -1,140 +1,149 @@
|
||||
---
|
||||
title: Scheduled Operations
|
||||
description: Automate recurring Docker operations like stack restarts, lifecycle management, fleet snapshots, and system prunes on a cron schedule.
|
||||
description: Automate stack lifecycle, image updates, vulnerability scans, fleet snapshots, and system prunes on a cron schedule.
|
||||
---
|
||||
|
||||
Schedules is a unified surface for every recurring maintenance operation Sencho knows how to run: stack restarts, per-node and fleet-wide image updates, lifecycle events (stop, take down, start, backup), system prunes, and vulnerability scans. The default view is a next-24-hour Timeline of upcoming runs across five lanes; an All tasks table view lists every schedule regardless of when it next fires.
|
||||
|
||||
<Note>
|
||||
Scheduled Operations is available to admins on Skipper and Admiral. Skipper unlocks **Auto-update Stack**, **Vulnerability Scan**, and **Fleet Snapshot**. All other actions (Restart Stack, System Prune, Backup Stack Files, Stop / Take Down / Start Stack) remain Admiral.
|
||||
Available to admins on Skipper and Admiral. Skipper unlocks **Auto-update Stack**, **Auto-update All Stacks**, **Vulnerability Scan**, and **Fleet Snapshot**. The remaining actions (Restart Stack, System Prune, Backup Stack Files, Stop / Take Down / Start Stack) require Admiral. The action picker hides operations your tier cannot run.
|
||||
</Note>
|
||||
|
||||
<Note>
|
||||
Schedules is hub-only and is hidden from the nav strip when a remote node is the active selection. See [Multi-Node Management](/features/multi-node#what-top-level-views-show-when-a-remote-node-is-active).
|
||||
</Note>
|
||||
|
||||
## Overview
|
||||
## Anatomy at a glance
|
||||
|
||||
Scheduled Operations lets you automate recurring maintenance tasks across your infrastructure. Define a cron schedule, choose an action, and Sencho handles the rest, including a full execution history log so you always know what ran and when.
|
||||
|
||||
The view opens on a **Timeline** that plots the next 24 hours of scheduled work across five lanes (Restart, Update, Scan, Prune, Lifecycle) so you can see, at a glance, what is about to fire and when. Toggle to **All tasks** for the full CRUD table.
|
||||
Open the **Schedules** tab from the top navigation bar. The page opens on the Timeline view with the masthead, the five lane track, and a bottom time axis.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/scheduled-operations/timeline.png" alt="Schedules timeline showing the next 24 hours across four lanes with a cyan now rail" />
|
||||
<img src="/images/scheduled-operations/timeline.png" alt="Schedules page in Timeline view. Header reads 'Scheduled Operations' on the left with a clock icon; on the right are a Timeline / All tasks toggle, a Refresh button, and a 'New Schedule' primary button. Below, a kicker 'NEXT 24 HOURS' sits above an italic display heading 'Next 24 hours' and a monospace date range 'Thu, May 14 11:21 → Fri, May 15 11:21'. A right-aligned 'Next' pill shows '03:00' with the subtitle 'Vul Scan · in 15h 38m'. Five lanes (Restart, Update, Scan, Prune, Lifecycle) run horizontally; a purple '03:00 Vul Scan' pill sits on the Scan lane and an amber '04:00 Nightly Snapshot' pill sits on the Prune lane. A cyan vertical 'now' rail glows at the left edge. Six time ticks (11:21, 16:09, 20:57, 01:45, 06:33, 11:21) run along the bottom axis." />
|
||||
</Frame>
|
||||
|
||||
## Timeline view
|
||||
|
||||
The timeline is the default view. It shows:
|
||||
The Timeline plots every firing of every enabled task across a rolling 24-hour window starting from the current minute.
|
||||
|
||||
- A hero with the current 24-hour window as a date range, and the **next firing** on the right (time, task name, and relative countdown).
|
||||
- Five color-coded lanes: **Restart** (cyan), **Update** (green), **Scan** (purple), **Prune** (amber), and **Lifecycle** (blue). Snapshot tasks share the Prune lane; stop, down, start, and backup tasks share the Lifecycle lane.
|
||||
- One pill per firing within the window, positioned proportionally to the task's next run time. Click any pill to open that task's execution history.
|
||||
- A vertical cyan **now rail** at the left edge, and six mono time ticks along the bottom axis.
|
||||
- **Masthead.** A `NEXT 24 HOURS` kicker, an italic display heading, the window's start and end timestamps in a monospace range, and a right-anchored **Next** pill that reads out the time and task name of the next firing and a relative countdown.
|
||||
- **Five lanes.** Restart (brand cyan), Update (success green), Scan (label purple), Prune (warning amber), and Lifecycle (label blue). Snapshot tasks share the Prune lane; the four stack-lifecycle actions (Backup Stack Files, Stop Stack, Take Stack Down, Start Stack) share the Lifecycle lane.
|
||||
- **Pills.** One pill per firing within the window, positioned proportionally to the firing's time. Each pill shows the firing time and the task name. Pills are color-matched to their lane. Click a pill to open the run history sheet for that task.
|
||||
- **Now rail.** A glowing vertical rail at the current minute, anchored to the left of the track at page open and drifting right as time passes (the page recomputes positions periodically).
|
||||
- **Axis.** Six monospace time ticks run along the bottom, evenly spaced through the window.
|
||||
|
||||
Tasks that fire more than once in the window (e.g. an hourly cron) render a pill for each firing. Disabled tasks do not appear on the timeline.
|
||||
Multi-fire crons (for example, `0 */6 * * *`) render a pill for each firing inside the window. Disabled tasks do not appear on the Timeline. If nothing fires in the next 24 hours, the lanes still render, and an empty-state line points you at the All tasks view.
|
||||
|
||||
Toggle to **All tasks** from the header to see every schedule in a table, regardless of whether it fires in the next 24 hours.
|
||||
Toggle to **All tasks** from the masthead to see every schedule in a table regardless of when it next fires.
|
||||
|
||||
## Supported Actions
|
||||
## All tasks view
|
||||
|
||||
| Action | Target | Description |
|
||||
|--------|--------|-------------|
|
||||
| **Restart Stack** | A specific stack (or specific services within it) on a specific node | Restarts all or selected containers in the stack |
|
||||
| **Auto-update Stack** | A specific stack on a specific node | Checks each image for updates and recreates the stack if any image has a newer version. See [Auto-Update Readiness](/features/auto-update-policies) for the companion board. Available on Skipper and Admiral. |
|
||||
| **Fleet Snapshot** | All nodes | Creates a fleet-wide backup of all compose files and `.env` files. Available on Skipper and Admiral. |
|
||||
| **System Prune** | The default node | Prunes selected resources, optionally filtered by Docker label |
|
||||
| **Vulnerability Scan** | All images on a specific node | Runs Trivy against every image on the target node and records the results. Requires Trivy to be installed, see [Installing Trivy](/operations/trivy-setup). Available on Skipper and Admiral. |
|
||||
| **Backup Stack Files** | A specific stack on a specific node | Backs up the stack's compose file and `.env` to `<DATA_DIR>/backups/<stackName>/`. The most recent backup per stack is kept; each run overwrites the previous one. |
|
||||
| **Stop Stack** | A specific stack on a specific node | Runs `docker compose stop` — containers are stopped but preserved. Use this for off-hours power saving when you want a fast start later. |
|
||||
| **Take Stack Down** | A specific stack on a specific node | Runs `docker compose down` — containers are removed. Use this to fully release resources when the stack is not needed for an extended period. |
|
||||
| **Start Stack** | A specific stack on a specific node | Runs `docker compose up -d`. Works whether the stack was stopped or taken down: if containers exist they are started; if not, they are created and started from the compose file. |
|
||||
|
||||
## Creating a Scheduled Task
|
||||
|
||||
1. Navigate to the **Schedules** tab in the top navigation bar (visible to Skipper and Admiral admins).
|
||||
2. Click **New Schedule**.
|
||||
3. Fill in the form:
|
||||
- **Name**: A descriptive label (e.g. "Nightly staging restart").
|
||||
- **Action**: Choose from the full action list. The form fields below change based on your selection.
|
||||
- **Node**: (Restart Stack, Auto-update Stack, and Vulnerability Scan) Select the node to run against. For stack actions it determines where the target stack lives; for Vulnerability Scan it determines which node's images are scanned.
|
||||
- **Stack**: (Restart Stack and Auto-update Stack) Select the stack to target. Becomes available after choosing a node.
|
||||
- **Services**: (Restart Stack only) Optionally select specific services within the stack to restart. Leave empty to restart all services.
|
||||
- **Prune Targets**: (System Prune only) Select which resources to prune: containers, images, networks, volumes. All are selected by default.
|
||||
- **Label Filter**: (System Prune only) Optionally filter prune operations to resources matching a specific Docker label (e.g. `com.docker.compose.project=mystack`).
|
||||
- **Cron Expression**: Standard 5-field cron format. A human-readable preview appears below the input.
|
||||
- **Enabled**: Toggle the task on or off.
|
||||
- **Delete after successful run**: When checked, the task automatically removes itself after its first successful execution. This turns the task into a one-time operation. Failures keep the task so you can inspect the error, adjust the target if needed, and trigger a retry with **Run Now** or by re-enabling the schedule.
|
||||
4. Click **Create**.
|
||||
The All tasks toggle swaps the lane track for a sortable table.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/scheduled-operations/create-dialog.png" alt="Create scheduled task dialog with action, cron expression, and prune target options" />
|
||||
<img src="/images/scheduled-operations/all-tasks.png" alt="All tasks table view of Scheduled Operations. Columns are Name, Action, Target, Schedule, Status, Next Run, Enabled, Actions. Two rows are visible: 'Vul Scan' with a Vulnerability Scan badge, target 'system', schedule 'At 03:00 AM / 0 3 * * *', Success status, next run '5/15/2026, 3:00:00 AM', ON toggle, and an action group of Run now / Execution history / Edit / Delete (Delete in destructive red). 'Nightly Snapshot' with a Fleet Snapshot badge, target 'fleet', schedule 'At 04:00 AM / 0 4 * * *', Success, next run '5/15/2026, 4:00:00 AM', ON, same action group." />
|
||||
</Frame>
|
||||
|
||||
## Task List
|
||||
| Column | What it shows |
|
||||
|---|---|
|
||||
| **Name** | The task name. |
|
||||
| **Action** | A badge labelling the operation (e.g. Restart Stack, Vulnerability Scan, Fleet Snapshot). |
|
||||
| **Target** | The stack the task targets (with a service list in parentheses when restart is scoped to specific services), or the target type for non-stack actions (`system`, `fleet`). |
|
||||
| **Schedule** | A human-readable description of the cron with the raw expression on a second line. |
|
||||
| **Status** | The last run result: **Success** (green), **Failed** (red), or `Never run` if the task has not fired yet. |
|
||||
| **Next Run** | The timestamp of the next firing, or a dash if the task is disabled or has no upcoming runs. |
|
||||
| **Enabled** | An on/off toggle that pauses or resumes the schedule without deleting it. |
|
||||
| **Actions** | Per-row controls: see [Managing tasks](#managing-tasks). |
|
||||
|
||||
The task list is displayed as a table with the following columns:
|
||||
## Supported actions
|
||||
|
||||
| Column | Description |
|
||||
|--------|-------------|
|
||||
| **Name** | The task name |
|
||||
| **Action** | Task type badge: Restart Stack, Fleet Snapshot, System Prune, or Vulnerability Scan |
|
||||
| **Target** | Stack name (with selected services, if any), or the target type for non-stack actions |
|
||||
| **Schedule** | Human-readable description with the raw cron expression below |
|
||||
| **Status** | Last run result: **Success** (green), **Failed** (red), or "Never run" |
|
||||
| **Next Run** | When the task will next execute |
|
||||
| **Enabled** | Toggle switch to enable or disable the task |
|
||||
| **Actions** | Action buttons (see [Managing Tasks](#managing-tasks)) |
|
||||
| Action | Tier | Target | What it does |
|
||||
|---|---|---|---|
|
||||
| **Restart Stack** | Admiral | A specific stack (optionally specific services) on a specific node | Restarts all or selected containers in the stack. |
|
||||
| **Auto-update Stack** | Skipper | A specific stack on a specific node | Checks each image in the stack for a newer tag and recreates the stack if any image has an update. See [Auto-Update Readiness](/features/auto-update-policies) for the companion review board. |
|
||||
| **Auto-update All Stacks** | Skipper | A specific node | Runs the auto-update check across every stack on the node that has auto-updates enabled. Stacks with auto-updates turned off are skipped. |
|
||||
| **Fleet Snapshot** | Skipper | The whole fleet | Creates a versioned, fleet-wide snapshot of every node's compose files and `.env` files. See [Fleet Backups](/features/fleet-backups). |
|
||||
| **System Prune** | Admiral | The selected node | Prunes containers, images, networks, and volumes (any subset), optionally filtered by a Docker label. |
|
||||
| **Vulnerability Scan** | Skipper | A specific node | Runs Trivy against every image on the node and persists the results. Requires Trivy to be installed on the target node ([Installing Trivy](/operations/trivy-setup)). |
|
||||
| **Backup Stack Files** | Admiral | A specific stack on a specific node | Copies the stack's compose file and `.env` to `<DATA_DIR>/backups/<stack>/`. One slot per stack; each run overwrites the previous backup. For a versioned archive use a Fleet Snapshot instead. |
|
||||
| **Stop Stack** | Admiral | A specific stack on a specific node | Runs `docker compose stop`. Containers are stopped but preserved. Use for off-hours power saving when you want a fast restart later. |
|
||||
| **Take Stack Down** | Admiral | A specific stack on a specific node | Runs `docker compose down`. Containers are removed. Use to fully release resources when the stack is not needed for an extended period. |
|
||||
| **Start Stack** | Admiral | A specific stack on a specific node | Runs `docker compose up -d`. Works for both stopped and removed containers: if they exist they are started, if not they are created from the compose file. |
|
||||
|
||||
## Granular Targeting
|
||||
## Creating a scheduled task
|
||||
|
||||
### Per-Service Restart
|
||||
|
||||
When creating a Restart Stack schedule, you can target individual services instead of restarting the entire stack. After selecting a stack, Sencho reads the compose file and displays checkboxes for each defined service. Select the services you want to restart, or leave all unchecked to restart every service in the stack.
|
||||
Click **New Schedule** in the header. The form opens in a centered modal. The Action picker only shows operations your tier can run, so Skipper admins see four options and Admiral admins see all ten.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/scheduled-operations/per-service-restart.png" alt="Service checkboxes displayed when creating a per-service restart schedule" />
|
||||
<img src="/images/scheduled-operations/action-picker.png" alt="The New scheduled task modal with the Action combobox expanded. The dropdown lists Restart Stack, Auto-update Stack, Auto-update All Stacks, Fleet Snapshot, System Prune, and Vulnerability Scan as the first six entries. Below the picker, partly visible, sit a Services row with 'echo' and 'prober' checkboxes, the Cron Expression input, the Enabled toggle, and the Delete after successful run checkbox." />
|
||||
</Frame>
|
||||
|
||||
### Scheduled Vulnerability Scans
|
||||
Common fields:
|
||||
|
||||
A Vulnerability Scan task runs Trivy against every image on the selected node and persists the results. The scan uses the same digest-based 24-hour cache as manual scans, so unchanged images are not rescanned on every run. See [Vulnerability Scanning](/features/vulnerability-scanning) for how results are surfaced in the UI and [Installing Trivy](/operations/trivy-setup) for setup on each node.
|
||||
- **Name.** A descriptive label.
|
||||
- **Action.** Pick the operation. The form below changes based on your selection.
|
||||
- **Cron Expression.** A standard 5-field cron. A human-readable preview appears below the input as you type (`At 03:00 AM`, `At 04:00 AM, only on Sunday`, and so on). See [Cron expression reference](#cron-expression-reference).
|
||||
- **Enabled.** Toggle the task on or off without deleting it.
|
||||
- **Delete after successful run.** When enabled, the task removes itself from the schedule after its first successful execution. Failures keep the task so you can inspect the error and retry. See [Delete after successful run](#delete-after-successful-run).
|
||||
|
||||
When a scheduled scan finishes, Sencho dispatches a completion notification with a summary of what was scanned and a breakdown of findings by severity. The full message format is documented in [Alerts & Notifications → Scheduled scan completion](/features/alerts-notifications#scheduled-scan-completion).
|
||||
Conditional fields per action:
|
||||
|
||||
### Prune Label Filter
|
||||
|
||||
When creating a System Prune schedule, you can scope the prune to resources matching a specific Docker label. This lets you target resources from a particular stack or project without affecting unrelated containers, images, or volumes.
|
||||
|
||||
Enter a label in `key=value` format (e.g. `com.docker.compose.project=mystack`). Leave the field empty to prune all unused resources of the selected types.
|
||||
- **Stack actions** (Restart Stack, Auto-update Stack, Backup Stack Files, Stop / Take Down / Start Stack) add a **Node** combobox and a **Stack** combobox. Restart Stack additionally renders a **Services** checkbox grid sourced from the stack's compose services, so you can scope the restart to a subset instead of restarting the entire stack.
|
||||
- **Auto-update All Stacks** adds a **Node** combobox with the helper text "Only stacks with auto-updates enabled on this node will be updated."
|
||||
- **Vulnerability Scan** adds a **Node** combobox with the helper text "Every image on the selected node will be scanned."
|
||||
- **System Prune** adds a **Prune Targets** group (Containers, Images, Networks, Volumes; all selected by default) and a **Label Filter** input for scoping the prune to resources matching a Docker label.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/scheduled-operations/prune-label-filter.png" alt="Label filter input for scoping prune operations to specific Docker labels" />
|
||||
<img src="/images/scheduled-operations/create-restart.png" alt="The New scheduled task modal configured for a Restart Stack task. Name reads 'Nightly staging restart'. Action is Restart Stack. Node is Local. Stack is 'audit-mesh-prod'. A Services group with helper text '(leave empty for all)' shows two unchecked checkboxes labelled 'echo' and 'prober'. Cron Expression is '0 3 * * *' with the preview 'At 03:00 AM'. The Enabled toggle reads ON. The Delete after successful run checkbox is unchecked. Cancel and Create buttons sit at the bottom right." />
|
||||
</Frame>
|
||||
|
||||
### Stack Lifecycle Scheduling
|
||||
## Granular targeting
|
||||
|
||||
The Stop Stack, Take Stack Down, Start Stack, and Backup Stack Files actions let you schedule lifecycle events for individual stacks on a per-node basis.
|
||||
### Per-service restart
|
||||
|
||||
A common pattern for development or staging stacks is:
|
||||
- A **Stop Stack** or **Take Stack Down** task scheduled for the end of the working day (e.g. `0 19 * * 1-5` — 7 PM on weekdays).
|
||||
- A **Start Stack** task scheduled for the start of the working day (e.g. `0 8 * * 1-5` — 8 AM on weekdays).
|
||||
When creating a Restart Stack task and a stack is selected, Sencho reads the compose file and renders a checkbox per defined service under **Services (leave empty for all)**. Check one or more boxes to restart only those services. Leave them all unchecked to restart the entire stack.
|
||||
|
||||
**Stop vs Down:** Use Stop Stack when you want containers ready to resume quickly; Docker keeps the container filesystem in place and `Start Stack` simply restarts the existing containers. Use Take Stack Down when you want to fully release resources; `Start Stack` recreates the containers from the compose file on the next run.
|
||||
### Scheduled vulnerability scans
|
||||
|
||||
**Backup Stack Files** keeps only the most recent backup per stack under `<DATA_DIR>/backups/<stackName>/`. For a point-in-time archive across all nodes, use a [Fleet Snapshot](/features/fleet-backups) schedule instead.
|
||||
A Vulnerability Scan task runs Trivy against every image on the selected node and persists the findings. Trivy must be installed on that node ([Installing Trivy](/operations/trivy-setup)). Manual and scheduled scans share the same digest-based 24-hour cache, so unchanged images that were scanned recently are reused instead of rescanned on every run, which keeps execution time low.
|
||||
|
||||
All four actions run against the local Sencho instance only. To schedule lifecycle operations on a remote node, manage the schedule from within that node's own Sencho UI.
|
||||
When a scheduled scan finishes, Sencho dispatches a completion notification. The severity reflects the outcome (info on a clean run, warning when findings are present); the category is `scan_finding`. The full message format and how it surfaces in the bell is documented in [Alerts & Notifications · Vulnerability scanning](/features/alerts-notifications#vulnerability-scanning).
|
||||
|
||||
### Delete after Successful Run
|
||||
<Frame>
|
||||
<img src="/images/scheduled-operations/create-scan.png" alt="The New scheduled task modal configured for a Vulnerability Scan. Name reads 'Nightly vulnerability scan'. Action is Vulnerability Scan. Node is Local with the helper text 'Every image on the selected node will be scanned.' Cron Expression is '0 3 * * *' with the preview 'At 03:00 AM'. The Enabled toggle reads ON. Cancel and Create buttons sit at the bottom right." />
|
||||
</Frame>
|
||||
|
||||
When **Delete after successful run** is enabled, the task removes itself from the schedule after its first successful execution. The entire task record, including run history, is removed. Use this for one-time preparatory or clean-up tasks: pre-scaling a stack before a deployment, a one-off backup before a config change, or stopping a stack once a migration is complete.
|
||||
### Prune label filter
|
||||
|
||||
If the run fails, the task stays in the schedule unchanged so you can inspect the error and retry. Disabling the option at any time before the successful run prevents auto-deletion.
|
||||
When creating a System Prune task, the **Label Filter (optional)** input scopes the prune to resources matching a specific Docker label. Use `key=value` format (for example, `com.docker.compose.project=staging`). Leave the field empty to prune every unused resource of the selected types.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/scheduled-operations/create-prune.png" alt="The New scheduled task modal configured for a System Prune. Name reads 'Weekly cleanup'. Action is System Prune. The Prune Targets group shows four checked boxes: Containers, Images, Networks, Volumes. The Label Filter (optional) input reads 'com.docker.compose.project=staging' with the helper text 'Only prune resources matching this Docker label.' Cron Expression is '0 4 * * 0' with the preview 'At 04:00 AM, only on Sunday'. The Enabled toggle reads ON. The Delete after successful run checkbox is unchecked. Cancel and Create buttons sit at the bottom right." />
|
||||
</Frame>
|
||||
|
||||
### Stack lifecycle scheduling
|
||||
|
||||
Stop Stack, Take Stack Down, Start Stack, and Backup Stack Files let you schedule lifecycle events for individual stacks on a per-node basis. A common pattern for development or staging stacks is:
|
||||
|
||||
- A **Stop Stack** or **Take Stack Down** task at the end of the working day (for example, `0 19 * * 1-5`, 7 PM on weekdays).
|
||||
- A **Start Stack** task at the start of the working day (for example, `0 8 * * 1-5`, 8 AM on weekdays).
|
||||
|
||||
**Stop vs Take Down.** Use Stop Stack when you want containers ready to resume quickly: Docker keeps the container filesystem in place and the matching Start Stack simply restarts the existing containers. Use Take Stack Down to fully release resources; Start Stack then recreates the containers from the compose file on the next run.
|
||||
|
||||
**Backup Stack Files** keeps a single slot per stack under `<DATA_DIR>/backups/<stack>/`. Each run overwrites the previous backup. For a versioned, point-in-time archive across all nodes, use a [Fleet Snapshot](/features/fleet-backups) task instead.
|
||||
|
||||
The four lifecycle actions execute against the local Sencho instance. To schedule lifecycle operations for a remote node, switch the active node to that node first and create the schedule from its UI; the schedule then lives on the remote node and runs there.
|
||||
|
||||
### Delete after successful run
|
||||
|
||||
When **Delete after successful run** is enabled, the task removes itself, including its run history, after its first successful execution. Use this for one-shot operations: a pre-deploy preparation step, a one-off backup before a config change, or a stop request that should fire exactly once.
|
||||
|
||||
If the run fails, the task stays in the schedule so you can inspect the error in the run history and retry with **Run now** or by re-enabling the schedule. Disabling the option at any time before the successful run prevents auto-deletion.
|
||||
|
||||
<Note>
|
||||
Manual runs via **Run Now** also trigger the delete-after-run logic. If you run the task manually and it succeeds, the task removes itself.
|
||||
Manual runs via **Run now** also trigger the delete-after-run logic. If you run a one-shot task manually and it succeeds, the task removes itself.
|
||||
</Note>
|
||||
|
||||
## Cron Expression Reference
|
||||
## Cron expression reference
|
||||
|
||||
Sencho uses standard 5-field cron expressions:
|
||||
|
||||
@@ -148,119 +157,106 @@ Sencho uses standard 5-field cron expressions:
|
||||
* * * * *
|
||||
```
|
||||
|
||||
### Common Examples
|
||||
Common examples:
|
||||
|
||||
| Expression | Description |
|
||||
|-----------|-------------|
|
||||
| Expression | Meaning |
|
||||
|---|---|
|
||||
| `0 3 * * *` | Every day at 3:00 AM |
|
||||
| `0 */6 * * *` | Every 6 hours |
|
||||
| `0 3 * * 0` | Every Sunday at 3:00 AM |
|
||||
| `30 2 1 * *` | 1st of every month at 2:30 AM |
|
||||
| `30 2 1 * *` | The 1st of every month at 2:30 AM |
|
||||
| `0 0 * * 1-5` | Midnight on weekdays |
|
||||
|
||||
## Filtering by Node
|
||||
## Filtering by node
|
||||
|
||||
When managing a multi-node fleet, you can filter the schedule list to show only tasks targeting a specific node. There are two ways to access this:
|
||||
When you manage a multi-node fleet, you can filter the schedule list to show only tasks targeting a specific node. There are two entry points:
|
||||
|
||||
- **From the Nodes table:** Click the **calendar icon** on any node row in **Settings → Nodes** to jump directly to the Schedules view filtered to that node.
|
||||
- **From the Schedules view:** A filter bar appears at the top showing which node you're viewing, with a **Clear filter** button to return to the full list.
|
||||
- **From the Nodes table:** click the calendar icon on any node row in **Settings · Nodes** to jump straight to Schedules filtered to that node.
|
||||
- **From the Schedules view:** a filter chip appears at the top of the page showing the active node filter, with a **Clear filter** button to return to the full list.
|
||||
|
||||
When creating a new task while a node filter is active, Sencho pre-selects that node in the create dialog.
|
||||
When you create a new task while a node filter is active, Sencho pre-selects that node in the create modal.
|
||||
|
||||
## Managing Tasks
|
||||
## Managing tasks
|
||||
|
||||
Each task row has four action buttons:
|
||||
Each row in the All tasks table has four per-row buttons:
|
||||
|
||||
- **Run Now** (play icon): Immediately execute the task without waiting for the next scheduled run. Manual runs are labeled "Manual" in the execution history.
|
||||
- **Execution History** (clock icon): Open the run history panel for this task.
|
||||
- **Edit** (pencil icon): Modify the task name, action, target, or schedule.
|
||||
- **Delete** (trash icon): Permanently remove the task and all its execution history after confirmation.
|
||||
- **Run now** (play icon). Immediately execute the task without waiting for the next firing. While the run is in flight, the button pulses; once it completes, the row's Status badge updates and the run appears in the history as a `Manual` source.
|
||||
- **Execution history** (clock icon). Open the run history sheet for the task. See [Execution history](#execution-history).
|
||||
- **Edit** (pencil icon). Open the task in the same modal used to create it. All fields are editable; the modal keeps the existing target unless you change the Action.
|
||||
- **Delete** (trash icon). Open a destructive confirmation dialog that permanently removes the task and its run history. There is no soft-delete.
|
||||
|
||||
Use the **Enabled** toggle switch in the task list to pause or resume a schedule without deleting it.
|
||||
Use the **Enabled** toggle in the row to pause a schedule without losing its configuration or history. Disabled tasks still appear in the table; they are excluded from the Timeline.
|
||||
|
||||
## Failure Notifications
|
||||
## Failure notifications
|
||||
|
||||
When a scheduled task fails, Sencho automatically dispatches an **error-level alert** through your configured notification channels (Discord, Slack, or custom webhooks). The alert includes the task name, action type, and error message so you can diagnose the issue immediately.
|
||||
When a scheduled task fails, Sencho dispatches an `error`-level notification in the `system` category through your configured channels (Discord, Slack, custom webhooks, and the in-app bell). The notification carries the task name, the action, and the error message so you can diagnose without opening the run history.
|
||||
|
||||
When a previously failing task succeeds again, Sencho sends an **info-level recovery notification** to confirm the issue is resolved. This recovery-only approach avoids notification noise from tasks that succeed on every run.
|
||||
When a task that previously failed succeeds again, Sencho dispatches an `info`-level recovery notification in the same category to confirm the issue is resolved. Tasks that succeed on every run do not produce per-run notifications: the recovery-only approach keeps the channel quiet during steady-state operation.
|
||||
|
||||
Vulnerability Scan tasks always send a completion notification, even on a clean run, because the message carries severity counts you may want to react to. See [Alerts & Notifications → Scheduled scan completion](/features/alerts-notifications#scheduled-scan-completion) for the full message format.
|
||||
Vulnerability Scan tasks are an exception: they always dispatch a completion notification, even on a clean run, because the message body carries severity counts you may want to react to. The severity is `info` on a clean run and `warning` when findings are present; the category is `scan_finding`, not `system`. The full message format and routing behavior is documented in [Alerts & Notifications · Vulnerability scanning](/features/alerts-notifications#vulnerability-scanning).
|
||||
|
||||
To configure notification channels, go to **Settings > Notifications**.
|
||||
Configure delivery channels in **Settings · Notifications**.
|
||||
|
||||
## Execution history
|
||||
|
||||
The Execution history button on any row opens a right-side sheet with the breadcrumb **Schedules › `<task name>` › Runs**. The sheet header reports the total run count, and a **Download CSV** secondary action exports the full history. The footer reads **Next run `<timestamp>`** when one is scheduled.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/scheduled-operations/failure-notification.png" alt="Failure notification alert shown in the notification bell popover" />
|
||||
<img src="/images/scheduled-operations/run-history.png" alt="The Execution history sheet for a task named 'Vul Scan'. Breadcrumb reads 'Schedules › VUL SCAN › RUNS'. The italic display heading 'Vul Scan' sits above a '28 runs' meta line and a 'Download CSV' button. A table lists nineteen runs with columns Time, Source (every row a Scheduled badge), Status (every row a green Success badge), Duration (mixing fast cache-hit runs of 0.0s with full scans up to 21.9s), and Details (mixing 'Scanned 15 image(s). Found 27 critical…' rows with 'All 15 image(s) already scanned recently (cache hit)' rows). The footer reads 'NEXT RUN 5/15/2026, 3:00:00 AM'." />
|
||||
</Frame>
|
||||
|
||||
## Execution History
|
||||
The table columns are:
|
||||
|
||||
Click the clock icon on any task to open the run history panel. The history is displayed as a table with the following columns:
|
||||
| Column | What it shows |
|
||||
|---|---|
|
||||
| **Time** | When the run started, in your local timezone. |
|
||||
| **Source** | A badge: **Scheduled** for cron firings, **Manual** for Run now triggers. |
|
||||
| **Status** | **Success** (green), **Failed** (red), or **Running** (in-flight). |
|
||||
| **Duration** | How long the run took. |
|
||||
| **Details** | A summary of the outcome for successful runs, or the captured error for failures. |
|
||||
|
||||
| Column | Description |
|
||||
|--------|-------------|
|
||||
| **Time** | When the run started |
|
||||
| **Source** | Whether the run was triggered by the **Scheduler** or **Manual** (via Run Now) |
|
||||
| **Status** | Success, Failed, or Running |
|
||||
| **Duration** | How long the run took (in seconds) |
|
||||
| **Details** | Output summary or error message |
|
||||
The sheet paginates at 20 rows per page; navigate with the **Previous** / **Next** controls in the footer. Run history is retained for 30 days; older runs are purged automatically.
|
||||
|
||||
Run history is paginated at 20 entries per page. You can export the full history as CSV using the download button in the panel header.
|
||||
The CSV export includes every run in the history with columns: Timestamp, Source, Status, Duration (s), Details.
|
||||
|
||||
Execution history is retained for 30 days.
|
||||
## How it works
|
||||
|
||||
<Frame>
|
||||
<img src="/images/scheduled-operations/run-history.png" alt="Execution history showing run source, status, duration, and details with export button" />
|
||||
</Frame>
|
||||
A background Scheduler Service evaluates due tasks every 60 seconds. When a task's next run time has passed:
|
||||
|
||||
## How It Works
|
||||
|
||||
The Scheduler Service runs in the background and checks for due tasks every 60 seconds. When a task's next run time has passed:
|
||||
|
||||
1. The scheduler verifies your license tier matches the action (Skipper for update, scan, snapshot; Admiral for everything else).
|
||||
2. It executes the configured action using the same internal services that power the UI buttons (restart, snapshot, prune).
|
||||
3. Results are logged to the execution history.
|
||||
4. On failure, an alert is dispatched via your configured notification channels.
|
||||
1. The scheduler checks that your license tier still allows the action; if not, the run is logged as failed and skipped.
|
||||
2. It executes the configured action using the same internal services that power the equivalent UI buttons (Restart, Snapshot, Prune, and so on).
|
||||
3. The result is logged to the execution history.
|
||||
4. On failure, an `error`-level notification goes out through the configured channels. On recovery from a previously failing state, an `info`-level recovery notification is dispatched.
|
||||
5. The next run time is recalculated from the cron expression.
|
||||
|
||||
If a task is still running from a previous execution, the scheduler skips it to prevent overlap.
|
||||
Most actions execute on whichever Sencho instance owns the schedule. The exception is **Auto-update Stack** and **Auto-update All Stacks**: when the target is a remote node, the gateway forwards execution to the remote node's `/api/auto-update/execute` endpoint, mirroring the per-node update flow used by the manual button. The remaining actions execute directly against the local instance's Compose directory, which is why lifecycle schedules for a remote node must be created from that node's own UI.
|
||||
|
||||
If the server restarts while a task is mid-execution, the orphaned run record is automatically marked as failed with a "Server restarted during execution" message on the next startup. No manual cleanup is needed.
|
||||
If a task is still running from a previous firing, the scheduler skips the new firing to prevent overlap.
|
||||
|
||||
If Sencho restarts while a task is mid-execution, the orphaned run record is marked as failed with the message `Server restarted during execution` on the next startup. No manual cleanup is needed.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### Task was automatically disabled
|
||||
|
||||
If a task's cron expression becomes invalid after creation (for example, due to a corrupt database edit), the scheduler disables the task and records the reason in the last error column. To fix this:
|
||||
|
||||
1. Open the task in edit mode.
|
||||
2. Re-enter a valid cron expression.
|
||||
3. Re-enable the task using the toggle switch.
|
||||
|
||||
### Run shows "Server restarted during execution"
|
||||
|
||||
This means Sencho was restarted (or crashed) while this task was mid-execution. The run was marked as failed automatically on startup. The task itself is still enabled and will run at its next scheduled time. If you want to re-run it immediately, use the **Run Now** button.
|
||||
|
||||
### Scheduled scan completed but no notification arrived
|
||||
|
||||
The scan notification shares the same delivery path as every other Sencho alert. Check, in order:
|
||||
|
||||
1. At least one channel is enabled in **Settings > Notifications** and the **Test** button succeeds for it.
|
||||
2. If the stack the scan is associated with has notification routes defined, make sure at least one matching route is enabled; the routing layer takes priority over the global channels.
|
||||
3. Open the notification bell. The in-app bell receives every notification regardless of channel configuration, and a failed external delivery is logged there as an error entry you can inspect.
|
||||
4. On a remote node, notification channels must be configured on the remote instance itself; channel settings are per-node.
|
||||
|
||||
### Scan task fails with "Trivy binary is not available"
|
||||
|
||||
Trivy must be installed on the node that runs the scan, not just on the primary instance. Follow [Installing Trivy](/operations/trivy-setup) on the target node, then trigger **Run Now** from the schedule to confirm the scan succeeds before waiting for the next cron tick.
|
||||
|
||||
### Start Stack fails with a compose file error
|
||||
|
||||
The Start Stack action runs `docker compose up -d` against the stack's compose directory. If the stack folder has been removed or its compose file is missing, the run fails. Re-create or restore the stack in the [Editor](/features/editor), then trigger **Run Now** to confirm the start succeeds before the next cron tick.
|
||||
|
||||
### Auto-backup always overwrites the previous backup
|
||||
|
||||
Backup Stack Files keeps a single backup slot per stack under `<DATA_DIR>/backups/<stackName>/`. Each run overwrites the previous backup. This is by design for simplicity. If you need a timestamped archive of compose files across all nodes, use a [Fleet Snapshot](/features/fleet-backups) schedule, which stores versioned snapshots.
|
||||
|
||||
### One-shot task disappeared after a successful run
|
||||
|
||||
A task with **Delete after successful run** enabled removes itself automatically after the first successful execution, including its run history. This is expected behavior. If you need a record of the run before deletion, export the execution history to CSV via the clock icon before the next successful run.
|
||||
<AccordionGroup>
|
||||
<Accordion title="A task was automatically disabled">
|
||||
If a task's cron expression stops being parseable (for example, due to a corrupt database edit), the scheduler disables the task and records the validation error in its last-error field. Open the task in Edit mode, re-enter a valid cron expression, and flip the Enabled toggle back on. The next firing recalculates from the new expression on the next tick.
|
||||
</Accordion>
|
||||
<Accordion title="A run shows 'Server restarted during execution'">
|
||||
Sencho was restarted (or the process crashed) while this run was in flight, so the scheduler marked it failed on startup to avoid a stuck `Running` row. The task itself is still enabled and will fire again at its next cron tick. If you want to re-execute immediately, click **Run now** on the row.
|
||||
</Accordion>
|
||||
<Accordion title="A scheduled scan completed but no notification arrived">
|
||||
Scan notifications share the same delivery path as every other Sencho notification. Check, in order: 1) at least one channel is enabled in **Settings · Notifications** and its **Test** button succeeds; 2) if the stack the scan is associated with has notification routes defined, at least one matching route is enabled (the routing layer takes priority over the global channels); 3) open the notification bell, where every notification is recorded regardless of channel configuration, and look for a delivery-error entry; 4) on a remote node, channels must be configured on that node itself because channel settings are per-node.
|
||||
</Accordion>
|
||||
<Accordion title="A Vulnerability Scan fails with 'Trivy binary is not available'">
|
||||
Trivy must be installed on the node that runs the scan, not just on the gateway. Follow [Installing Trivy](/operations/trivy-setup) on the target node, then click **Run now** on the schedule to confirm the scan succeeds before waiting for the next cron tick.
|
||||
</Accordion>
|
||||
<Accordion title="A Start Stack run fails with a compose file error">
|
||||
Start Stack runs `docker compose up -d` against the stack's compose directory. If the stack folder has been removed or its compose file is missing, the run fails. Re-create or restore the stack from the [Editor](/features/editor), then click **Run now** on the schedule to confirm it succeeds before the next cron tick.
|
||||
</Accordion>
|
||||
<Accordion title="A Backup Stack Files run overwrote the previous backup">
|
||||
Backup Stack Files keeps a single slot per stack under `<DATA_DIR>/backups/<stack>/`. Each run overwrites the previous backup. This is by design: for a versioned archive of compose files across the fleet, use a [Fleet Snapshot](/features/fleet-backups) task, which retains every snapshot under the snapshots directory.
|
||||
</Accordion>
|
||||
<Accordion title="A one-shot task disappeared from the list">
|
||||
A task with **Delete after successful run** enabled removes itself automatically after the first successful execution, along with its run history. This is expected behavior. If you need a record of the run before deletion, click **Download CSV** in the run history sheet before the next successful run.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
@@ -1,66 +1,160 @@
|
||||
---
|
||||
title: Sencho Mesh
|
||||
description: Connect containers across nodes by hostname over an authenticated WebSocket tunnel, no VPN, no firewall changes, no extra ports.
|
||||
description: Cross-node container networking. Reach any meshed service on any node by hostname, over the same authenticated channel Sencho already uses to manage the fleet.
|
||||
---
|
||||
|
||||
<Note>
|
||||
Sencho Mesh requires a Sencho **Admiral** license. Skipper and Community Edition do not include this feature.
|
||||
Sencho Mesh requires an [Admiral license](/features/licensing). Community and Skipper do not include this feature.
|
||||
</Note>
|
||||
|
||||
Sencho Mesh makes a multi-node fleet feel like one machine. Opt a stack into the mesh and its services become reachable from any other meshed stack on the fleet by a stable hostname. Traffic rides an authenticated WebSocket tunnel between Sencho instances, so there are no new ports to open and no separate VPN to manage.
|
||||
Sencho Mesh gives a multi-node fleet the network topology of a single machine. Once a stack opts in, every service it exposes becomes reachable from any other meshed stack on the fleet by a stable hostname. Cross-node traffic rides the same authenticated channel Sencho already uses to manage the fleet, so a node behind NAT or a residential firewall participates exactly like a public VPS.
|
||||
|
||||
Mesh works with any remote mode:
|
||||
The audience is operators running a small fleet of Docker Compose hosts (homelab, lab + colo, cross-region production) who want service-to-service connectivity across hosting boundaries without standing up Tailscale, WireGuard, or a service mesh sidecar per container.
|
||||
|
||||
- **Pilot Agent** nodes carry mesh traffic over the long-lived agent tunnel.
|
||||
- **Distributed API** nodes carry mesh traffic over a short-lived tunnel that central opens on demand using the node's API token. Central tears the tunnel down after five minutes of idle, so a node that sees no mesh traffic costs nothing extra to keep configured.
|
||||
## Mental model
|
||||
|
||||
<Card title="Pilot Agent" icon="link" href="/features/pilot-agent">
|
||||
Pilot Agent is the easiest remote mode for nodes that cannot expose an inbound port. See Pilot Agent for how to enroll a node into your fleet.
|
||||
</Card>
|
||||
Three moving parts cooperate per node.
|
||||
|
||||
## How it works
|
||||
1. **The `sencho_mesh` Docker bridge.** Each Sencho instance creates an internal Docker network on first boot (default `172.30.0.0/24`) and pins itself to the first usable host address on it. This is the lane every meshed container uses to talk to the local Sencho.
|
||||
2. **The alias registry.** When a stack opts in, Sencho publishes a hostname for every service port: `<service>.<stack>.<node>.sencho`. A Postgres `db` service in a stack called `api` on a node called `opsix` is `db.api.opsix.sencho`. The registry is fleet-wide; every Sencho knows the full set.
|
||||
3. **The cross-node transport.** Sencho terminates the alias on the destination node's `sencho_mesh` bridge and carries TCP bytes over the channel between the two Sencho instances. On Pilot Agent nodes that is the existing agent tunnel. On Distributed API Proxy nodes it is a separate authenticated WebSocket bridge that central dials on demand using the node's API token.
|
||||
|
||||
Each Sencho instance creates an internal Docker bridge network called `sencho_mesh` (default subnet `172.30.0.0/24`) on first boot and pins itself to a static IP on that network. When you opt a stack into the mesh, Sencho:
|
||||
The user-facing effect is `psql -h db.api.opsix.sencho` from a container on any other meshed node, with no port forwarding on the host that runs Postgres.
|
||||
|
||||
1. Generates a Compose override file that injects `extra_hosts` for every cross-node alias the fleet currently exposes, pointing each one at the local Sencho's static IP.
|
||||
2. Attaches every service in the stack to the `sencho_mesh` network so the alias IP is reachable from inside the user's containers.
|
||||
3. Redeploys the stack with the override applied so its containers pick up the new entries and the network attachment.
|
||||
## Key capabilities
|
||||
|
||||
Aliases follow a predictable scheme:
|
||||
**Cross-node service discovery by hostname.** Aliases follow a predictable scheme (`<service>.<stack>.<node>.sencho`), so application config can name remote services without hardcoding container IPs, host IPs, or per-node DNS entries.
|
||||
|
||||
```
|
||||
<service>.<stack>.<node>.sencho
|
||||
```
|
||||
**Per-stack opt-in.** Mesh participation is per stack per node, not per node. A node can mesh some stacks and leave others isolated. Opt-in is sticky: stopping a stack pauses its aliases without losing the opt-in record, and they republish when the stack starts again.
|
||||
|
||||
A Postgres `db` service in a stack named `api` on a node named `opsix` is reachable as `db.api.opsix.sencho` from any other meshed stack.
|
||||
**Bidirectional traffic over a single channel.** Mesh traffic is multiplexed onto the same channel Sencho already uses for fleet operations, so a node behind NAT can both receive and originate connections without exposing any new inbound port. The only inbound port that ever matters is the one Sencho itself is already listening on for fleet operations.
|
||||
|
||||
**Live fleet-wide diagnostics.** Every alias has a one-click probe that runs across the real code path. Each node exposes a diagnostics panel showing forwarder liveness, pilot tunnel state, active TCP streams, and the resolver cache. A fleet-wide activity log records routing decisions and tunnel state changes as they happen.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
- An Admiral license on the central Sencho.
|
||||
- At least one [enrolled remote node](/features/multi-node) (Pilot Agent or Distributed API Proxy).
|
||||
- Functioning `sencho_mesh` data plane on every node that should participate. The Routing tab shows a red banner if a node's data plane did not come up; the troubleshooting section below covers each cause.
|
||||
- Per-stack opt-in. Enabling mesh on a node does not automatically place every stack into the mesh; each stack is opted in individually.
|
||||
|
||||
App-layer authentication is **not** in scope. Postgres still needs a password, Redis still needs an ACL, your internal HTTP API still needs whatever auth it normally uses. The mesh moves bytes; it does not authenticate the protocols inside those bytes.
|
||||
|
||||
## Enable the mesh
|
||||
|
||||
Mesh lives under **Fleet → Routing**.
|
||||
|
||||
<Frame caption="Fleet → Routing, Table view. Each node card shows mesh stack count, published aliases, per-alias Test probe, and an Add stack to mesh action.">
|
||||
<img src="/images/sencho-mesh/routing-tab-overview.png" alt="Sencho Mesh Routing tab Table view with node cards" />
|
||||
</Frame>
|
||||
|
||||
1. Open **Fleet → Routing**.
|
||||
2. Toggle **mesh** on for each node you want to participate.
|
||||
3. Click **Add stack to mesh** on any node and tick the stacks whose services should be reachable cross-node.
|
||||
2. Flip the mesh toggle (`ON` / `OFF`) on each node that should participate.
|
||||
3. Click **Add stack to mesh** on a node and confirm one or more stacks.
|
||||
|
||||
Each opt-in or opt-out triggers an automatic redeploy of the affected stack so its `/etc/hosts` and network attachments refresh. The opt-in sheet shows a confirmation prompt before it starts the redeploy.
|
||||
<Frame caption="The opt-in sheet. Already-meshed stacks show an `in mesh` pill plus a Topology shortcut; the rest get an Add to mesh button.">
|
||||
<img src="/images/sencho-mesh/opt-in-sheet.png" alt="Mesh opt-in sheet listing stacks on a node" />
|
||||
</Frame>
|
||||
|
||||
## What's exposed and what isn't
|
||||
Each opt-in and opt-out triggers a redeploy of the affected stack so its `/etc/hosts` entries and `sencho_mesh` network attachment refresh. The confirmation modal makes this explicit (`Add and redeploy` / `Remove and redeploy`) and the post-action toast says `<stack> added to mesh, redeploying`.
|
||||
|
||||
Four guarantees:
|
||||
## Lifecycle and behavior
|
||||
|
||||
1. **Only opted-in services are reachable.** A stack reaches another stack's services through the mesh only if both stacks have explicitly opted in. The Pilot agent on the target node refuses any request for a non-opted service.
|
||||
2. **Aliases are not internet-reachable.** Sencho listens on an internal Docker network; nothing about the mesh exposes new ports beyond the host's existing firewall posture.
|
||||
3. **Traffic is encrypted in transit.** Cross-node bytes ride an authenticated WSS tunnel between Sencho instances. The Node Token generated during fleet enrollment is the credential that authorizes the mesh tunnel; restricted API token scopes (read-only, deploy-only) cannot carry mesh traffic.
|
||||
4. **Tier-gated and audit-logged.** Only Admiral users can configure the mesh. Enable, disable, opt-in, and opt-out events write durable rows to the audit log with the actor's identity.
|
||||
**Opt-in.** Sencho writes a Compose override file (`<DATA_DIR>/mesh/overrides/<nodeId>/<stack>.override.yml`) that adds the `sencho_mesh` network attachment to every service in the stack and injects `extra_hosts` lines for every cross-node alias the fleet currently exposes. The stack redeploys with that override applied. Containers come back up with the alias names resolvable in their `/etc/hosts`.
|
||||
|
||||
What the mesh does **not** do:
|
||||
**Opt-out.** The override file is removed and the stack redeploys without it. Containers come back without the mesh network attachment and without the `extra_hosts` lines.
|
||||
|
||||
- No application-layer authentication. Your Postgres still needs a password.
|
||||
- No per-port firewall. Any opted-in stack can reach any other opted-in service.
|
||||
- No rate limiting or quotas.
|
||||
- No persistent traffic metrics. Live diagnostics only.
|
||||
**Stack stopped.** The opt-in record is sticky. A stack that opts in and then stops continues to show on the node card with a `suspended` pill and the caption `Stack stopped, alias resumes when services start.` Aliases republish on the next alias-refresh tick (within roughly one minute) when the stack starts again. No need to opt out and back in.
|
||||
|
||||
## Customizing the mesh subnet
|
||||
**Peer reconnect.** On every alias-refresh tick (60 s), and immediately on a tunnel-up event, central recomputes the global alias set and pushes refreshed overrides to any node whose set drifted. Stack redeploys are skipped if the override content is unchanged.
|
||||
|
||||
Each Sencho creates `sencho_mesh` as a `/24` bridge network at `172.30.0.0/24` by default. If that range collides with an existing network on a host, override it with the `SENCHO_MESH_SUBNET` environment variable on that node:
|
||||
**Proxy-mode bridge.** When central needs to forward TCP to a Distributed API Proxy node, it dials `/api/mesh/proxy-tunnel` on the remote using the node's API token, opens a persistent bidirectional WebSocket, and multiplexes TCP frames over it. The bridge stays open by default; central reconciles missing bridges every 60 s and reactively redials on any non-terminal teardown. Auth failures are terminal and skip redial. Recent dial failures are cached for 60 s so a misconfigured remote does not see a redial storm.
|
||||
|
||||
To make a proxy-mode bridge tear down after idle and reopen on demand, set `SENCHO_MESH_PROXY_TUNNEL_IDLE_MS` on the central node to a non-zero millisecond value. The default is `0` (no idle close).
|
||||
|
||||
## Diagnostics
|
||||
|
||||
Every node card has a **Diagnostics** button that opens a live view.
|
||||
|
||||
<Frame caption="Per-node diagnostics. Forwarder liveness, pilot tunnel state, active TCP streams, and the resolver cache mapping aliases to backend host:port pairs.">
|
||||
<img src="/images/sencho-mesh/diagnostics-sheet.png" alt="Sencho Mesh diagnostics sheet for a node" />
|
||||
</Frame>
|
||||
|
||||
The sheet shows:
|
||||
|
||||
- **Forwarder** state and number of listening ports. The forwarder is the in-process TCP listener Sencho binds for each opted-in service port.
|
||||
- **Pilot tunnel** state (`connected` / `disconnected`) and last-seen timestamp. For local diagnostics this reports the local Sencho's own forwarder; for a remote it reports central's view of the tunnel to that remote.
|
||||
- **Active streams** with byte counters in and out, and how long each stream has been open.
|
||||
- **Resolver cache** showing the aliases currently registered on this node and the backend `host:port` they resolve to.
|
||||
|
||||
Mesh runs in-process on each node; there is no separate mesh container to inspect with `docker ps`.
|
||||
|
||||
## Mesh activity log
|
||||
|
||||
The Routing-tab masthead has a **Mesh activity** button that opens a fleet-wide event log.
|
||||
|
||||
<Frame caption="Mesh activity. Every route resolution, tunnel state change, opt-in or opt-out, probe, and forwarder event is recorded with a source, type, and message. Filterable by alias, type, or message.">
|
||||
<img src="/images/sencho-mesh/mesh-activity-sheet.png" alt="Sencho Mesh activity log sheet" />
|
||||
</Frame>
|
||||
|
||||
The log is the first place to look when something flips state unexpectedly. It is an in-memory ring buffer of the most recent ~1000 events and resets on Sencho restart; for long-term retention, ship the audit log (opt-in and opt-out events also land there) to an external system.
|
||||
|
||||
## Topology view
|
||||
|
||||
The Routing tab has a **Table** / **Graph** toggle. Graph mode draws the fleet as a node-and-edge diagram so the live state is visible at a glance.
|
||||
|
||||
A second toggle picks what the edges encode.
|
||||
|
||||
**Tunnels** colours one edge per remote node by tunnel state.
|
||||
|
||||
<Frame caption="Graph view, Tunnels mode. The brand-coloured edge labelled `pilot · ok` is an active Pilot Agent tunnel. The dashed `proxy` edges are Distributed API Proxy peers that central can dial on demand.">
|
||||
<img src="/images/sencho-mesh/graph-tunnels.png" alt="Sencho Mesh topology graph, Tunnels mode" />
|
||||
</Frame>
|
||||
|
||||
Edge labels:
|
||||
|
||||
- `pilot · ok`: Pilot Agent tunnel is connected.
|
||||
- `pilot · idle`: Pilot Agent tunnel is configured but not currently up.
|
||||
- `proxy`: Distributed API Proxy peer. Central dials the mesh bridge on demand.
|
||||
- `unreachable`: credentials or remote version cannot carry mesh traffic. The node card spells out the specific reason.
|
||||
|
||||
**Aliases** keeps the same node layout and labels each edge with the number of aliases the remote node publishes.
|
||||
|
||||
<Frame caption="Graph view, Aliases mode. Each edge label shows what the remote node publishes for the rest of the fleet to consume.">
|
||||
<img src="/images/sencho-mesh/graph-aliases.png" alt="Sencho Mesh topology graph, Aliases mode" />
|
||||
</Frame>
|
||||
|
||||
Click any node card in graph mode to open its opt-in sheet.
|
||||
|
||||
**Per-stack topology.** Inside the opt-in sheet, each opted-in stack row has a **Topology** button that opens a focused diagram for that stack.
|
||||
|
||||
<Frame caption="Per-stack topology. Stack at the centre, every alias it publishes to the right, and a column of meshed consumer nodes with their tunnel state.">
|
||||
<img src="/images/sencho-mesh/stack-topology-sheet.png" alt="Per-stack mesh topology sheet" />
|
||||
</Frame>
|
||||
|
||||
The "consumer nodes" column lists meshed peers that could reach this stack's aliases via DNS. Whether a container on a consumer actually dials an alias depends on that consumer's own opt-in stacks.
|
||||
|
||||
The Routing tab polls `/mesh/status` and `/mesh/aliases` every 30 seconds while the browser tab is in the foreground, so tunnel state changes and alias additions appear without a manual reload. Polling pauses automatically when the tab is hidden.
|
||||
|
||||
The graph is designed for fleet sizes typical of self-hosted Compose setups (up to roughly 50 nodes). Larger fleets render but become dense; the Table view is more readable for inventory at scale.
|
||||
|
||||
## Test upstream
|
||||
|
||||
Every alias row has a one-click test that runs a real probe along the same code path traffic uses. The result appears as a toast:
|
||||
|
||||
- **Success.** `<alias> ok (<rtt>ms)`.
|
||||
- **Failure.** `<alias> <stage>: <code>`, where `<stage>` is one of:
|
||||
- `no_route`: the alias does not resolve on central. The destination stack is not opted in, or its opt-in is sticky and the stack is currently stopped.
|
||||
- `pilot_tunnel`: no bridge to the destination node could be opened. The Pilot Agent tunnel is down, or the proxy-mode dial failed.
|
||||
- `agent_resolve`: the bridge is up but the remote refuses the alias. The destination stack is not opted in on its home node.
|
||||
- `agent_dial`: the remote accepted the request but could not connect to the target container.
|
||||
- `target_port`: the remote dialed the target container but no service answered on the declared port.
|
||||
|
||||
Use Test before assuming the issue is your application. It tells you whether the mesh path itself is working.
|
||||
|
||||
## Customising the mesh subnet
|
||||
|
||||
The default `172.30.0.0/24` will collide if a host already has a Docker bridge in that range. Override it per node with `SENCHO_MESH_SUBNET`:
|
||||
|
||||
```yaml
|
||||
services:
|
||||
@@ -69,74 +163,46 @@ services:
|
||||
- SENCHO_MESH_SUBNET=10.42.0.0/24
|
||||
```
|
||||
|
||||
Sencho's static IP on the network is `<network address> + 2` (so `10.42.0.2` for the example above). Operators can configure each node independently; the override generator pushes alias hostnames that resolve to whichever IP the deploying node uses locally.
|
||||
Sencho pins itself to `<network address> + 2` (so `10.42.0.2` for the example above; the bridge's `+1` gateway sits between). Each node is configured independently; the alias registry pushes the correct local IP to each node's override file.
|
||||
|
||||
## Bidirectional traffic
|
||||
## Security and trust boundaries
|
||||
|
||||
Mesh traffic flows in both directions over a single persistent WebSocket
|
||||
between central and each proxy-mode peer. Central dials the bridge to every
|
||||
mesh-enabled peer on startup and re-dials any dropped bridge on its
|
||||
reconcile tick (every 60 seconds by default). When a container on the peer
|
||||
needs to reach a service on central, the request multiplexes over that same
|
||||
bridge: no extra inbound listener on central, no public URL required on
|
||||
the central side. This is the path that makes "call a service on any node
|
||||
by hostname" work in NAT'd homelab topologies out of the box.
|
||||
**Authentication.** Cross-node traffic rides an authenticated WebSocket between Sencho instances. Pilot Agent nodes authenticate with the long-lived JWT issued at enrollment (stored at `/app/data/pilot.jwt` on the agent). Distributed API Proxy nodes authenticate with the [Node Token](/features/multi-node#add-a-remote-node-distributed-api-proxy) generated on the remote. Restricted API token scopes (read-only, deploy-only) cannot carry mesh traffic; only a full-admin Node Token can.
|
||||
|
||||
## Test upstream
|
||||
**Inbound exposure.** Mesh adds no new inbound port to central. A Pilot Agent node's outbound tunnel carries mesh traffic in both directions. A Distributed API Proxy node only needs the inbound port Sencho already listens on for fleet operations (so it can be dialed by central). Aliases themselves are not internet-reachable; they only resolve inside the `sencho_mesh` Docker bridge on each participating node.
|
||||
|
||||
Every alias row has a one-click **Test** button that runs a real probe across the same code path traffic uses. Result is shown inline:
|
||||
**Encryption in transit.** The WebSocket between Sencho instances is whatever TLS posture you already configured for fleet management. If you front Sencho with a TLS-terminating reverse proxy, mesh inherits that. If you use a private CA, see Pilot Agent's [`SENCHO_PILOT_CA_FILE`](/features/pilot-agent#environment-variables) for the supported trust-store override.
|
||||
|
||||
- **Green tick** with round-trip time when the path is healthy.
|
||||
- **Red badge** with the failing stage (`pilot_tunnel`, `agent_resolve`, `agent_dial`, or `target_port`) when something is wrong.
|
||||
**Audit trail.** Opt-in and opt-out events write durable rows to the audit log with the actor's identity. Lower-level events (tunnel state changes, alias publishes, route resolutions, probe results) land in the mesh activity log described above, which is in-memory only.
|
||||
|
||||
Use the Test button before assuming an issue is your application's fault. It tells you whether the mesh path itself is the problem.
|
||||
**App-layer auth is your problem.** The mesh transports bytes. Database passwords, API keys, ACLs, and mTLS between application services are all unchanged by Mesh. Treat a meshed network as you would any flat L3 segment.
|
||||
|
||||
## Diagnostics
|
||||
## Limitations and non-goals
|
||||
|
||||
Every node card has a **Diagnostics** button that opens a live view of:
|
||||
These are the explicit boundaries of the v1 mesh.
|
||||
|
||||
- Forwarder liveness and pilot tunnel state on this node.
|
||||
- Active TCP streams with byte counters and open age.
|
||||
- The resolver cache showing which aliases are registered.
|
||||
- **One alias per TCP port across the fleet.** If two stacks expose the same port (for example two Postgres instances on 5432), only the first can be added to the mesh. The opt-in sheet returns a clear inline error if a second tries.
|
||||
- **Sencho's API port is reserved.** A meshed service exposing port 1852 is rejected at opt-in to prevent collision with Sencho's own listener.
|
||||
- **Remote-to-remote routes through central.** Central can carry mesh traffic between any two remote nodes by relaying frames through itself. Direct peer-to-peer tunnels between two remotes are not in v1.
|
||||
- **Shared stream pool with the Pilot tunnel.** Each Pilot tunnel multiplexes up to **1024** concurrent streams covering HTTP, WebSocket, and mesh TCP. A heavy mesh workload counts against the same ceiling as ordinary fleet API traffic. See Pilot Agent's [Resource limits](/features/pilot-agent#resource-limits) for the full picture.
|
||||
- **No TLS termination, no L7 features.** Mesh is L4. There is no built-in HTTPS, no host-based routing, no header rewriting, no blue/green cutover. Run those at your existing reverse proxy.
|
||||
- **No application authentication.** Mesh does not add a credential layer to the transported protocol. Your services still authenticate their callers themselves.
|
||||
- **`network_mode: host` services cannot join.** A service that runs on the host network namespace cannot attach to `sencho_mesh` and therefore cannot publish an alias. The rest of the stack can still mesh; the host-network service stays out.
|
||||
- **Mesh activity log is in-memory.** The fleet-wide log resets when Sencho restarts. For long-term retention rely on the audit log (opt-in / opt-out) or export from your reverse proxy.
|
||||
|
||||
This is the first place to look when a connection isn't behaving as expected.
|
||||
## Example workflows
|
||||
|
||||
## Mesh activity
|
||||
**Home node plus cloud VPS.** A homelab runs Postgres for the household; a small VPS in a public datacentre runs a public-facing web app. Enable mesh on both. Opt the home node's Postgres stack and the cloud node's app stack into the mesh. The app's config sets `DATABASE_URL=postgres://app@db.postgres.home.sencho/app`. No port forwarding on the home router; no WireGuard tunnel to maintain.
|
||||
|
||||
The masthead has a **Mesh activity** button that opens the fleet-wide event log. Every route resolution, tunnel state change, opt-in, opt-out, and probe is recorded there. Filter by alias, source, type, or message. Useful for understanding what just happened when something flips state.
|
||||
**Two cloud nodes, blue and green.** Two VPSes each run a copy of a stateless app; both pull from a shared Redis stack on a third "infra" node. All three nodes are meshed. The two app nodes opt their app stacks in; the infra node opts its Redis stack in. Either app reaches Redis as `cache.redis.infra.sencho`. Promoting blue → green is a redeploy on the app node; nothing about the mesh changes.
|
||||
|
||||
## Topology view
|
||||
|
||||
The Routing tab has a **Table** / **Graph** toggle in its header. Graph mode draws the fleet as a node-and-edge diagram so it is clear at a glance which nodes are meshed, which tunnels are live, and where aliases are published.
|
||||
|
||||
A second toggle picks what the edges encode:
|
||||
|
||||
- **Tunnels** shows one edge per remote node, coloured by tunnel state. A solid brand edge labelled `pilot · ok` or `proxy` means traffic is ready to flow. A dashed muted edge labelled `pilot · idle` means the Pilot agent is offline. A dashed red edge labelled `unreachable` means the credentials or remote build cannot carry mesh traffic; the node card shows the specific reason.
|
||||
- **Aliases** keeps the same node layout and labels each edge with the number of aliases the remote node publishes. A remote that publishes nothing reads `no aliases`.
|
||||
|
||||
Click any node card to open the opt-in sheet for that node. On each opted-in stack row the sheet shows a **Topology** button that opens a focused diagram for that one stack: the stack at the centre, every alias it publishes branching out, and a column of meshed consumer nodes with their tunnel state. Use it to confirm what a stack exposes and which peers can reach it.
|
||||
|
||||
In the stack diagram, *consumer nodes* are meshed peers that could reach this stack's aliases via DNS. Whether a container on a consumer actually dials an alias depends on the consumer's own opt-in stacks.
|
||||
|
||||
The graph reads the same `/mesh/status` and `/mesh/aliases` data the Table view does, so any opt-in or opt-out refreshes both views. The Routing tab also refreshes both feeds every 30 seconds while the browser tab is in the foreground, so tunnel state changes and alias additions appear without a manual reload. Polling pauses automatically when the tab is hidden.
|
||||
|
||||
The graph is designed for fleet sizes typical of self-hosted Compose setups (up to roughly 50 nodes). Larger fleets render but become visually dense; the Table view is the more readable surface for inventory at scale.
|
||||
|
||||
## V1 limitations
|
||||
|
||||
A few things are deliberately out of scope for the first release:
|
||||
|
||||
- **One alias per TCP port across the fleet.** If two stacks expose the same port (for example two Postgres instances on 5432), only the first can be added to the mesh. The opt-in sheet shows a clear inline error if the second tries.
|
||||
- **Sencho's API port is reserved.** A meshed service exposing port 1852 is rejected at opt-in to prevent collision with the Sencho UI / API listener.
|
||||
- **Remote-to-remote routing rides through central.** Mesh works for traffic between central and any remote in either direction, and between two remotes via central relay. Direct peer-to-peer tunnels between remotes are not supported.
|
||||
- **Stream pool shared with the Pilot tunnel.** Each Pilot tunnel multiplexes up to 1024 concurrent streams covering HTTP, WebSocket, and mesh TCP traffic. A heavy mesh workload counts against the same ceiling as ordinary fleet API traffic. See the [Pilot Agent](/features/pilot-agent) resource-limit notes for the full picture.
|
||||
- **No TLS termination, no blue/green cutover.** Layer 7 features land in a follow-up.
|
||||
**Triage a "my app can't connect."** A developer reports their app stack can't reach `db.api.opsix.sencho`. Open Fleet → Routing and click the Test button on the alias row. A green toast confirms the mesh path is healthy and the issue is inside the application. A red toast with `pilot_tunnel: ...` pushes you to the Diagnostics sheet on `opsix` and the Mesh activity log to chase the tunnel state.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Mesh data plane is down">
|
||||
The Routing tab shows a red banner when this Sencho's `sencho_mesh` setup did not complete. The banner names the specific reason; the same reason appears in the mesh activity log and on `/api/health` as `mesh.dataPlane.reason`. The fix depends on which reason fired:
|
||||
The Routing tab shows a red banner when the local Sencho's `sencho_mesh` setup did not complete. The banner names the specific reason; the same reason appears in the mesh activity log and on `/api/health` as `mesh.dataPlane.reason`. The fix depends on which reason fired:
|
||||
|
||||
- `subnet_overlap`: the requested CIDR overlaps another Docker bridge network on this host. Run `docker network ls -q | xargs -L1 docker network inspect --format '{{.Name}} {{range .IPAM.Config}}{{.Subnet}} {{end}}'` to list every existing subnet, then set `SENCHO_MESH_SUBNET` to a free `/24` (for example `10.42.0.0/24`) and recreate the Sencho container.
|
||||
- `subnet_mismatch`: `sencho_mesh` already exists with a different subnet. Either remove the network (`docker network rm sencho_mesh` after detaching any containers) or set `SENCHO_MESH_SUBNET` to match the existing subnet.
|
||||
@@ -148,47 +214,107 @@ A few things are deliberately out of scope for the first release:
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Opt-in rejects with 'host-network service'">
|
||||
Stacks whose services declare `network_mode: host` cannot join `sencho_mesh` and therefore cannot participate in the mesh. Switch the affected service to bridge networking and redeploy, or accept that the stack stays out of the mesh.
|
||||
Stacks whose services declare `network_mode: host` cannot join `sencho_mesh` and therefore cannot publish a mesh alias. Switch the affected service to bridge networking and redeploy, or accept that the stack stays out of the mesh.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="A route shows `tunnel down`">
|
||||
The Pilot tunnel to the target node is gone. Check **Fleet → Overview** for the node's status. Mesh recovers automatically when the tunnel reconnects.
|
||||
<Accordion title="Opt-in returns 409 'Port already claimed by another mesh stack'">
|
||||
The mesh enforces one alias per TCP port across the fleet. Another stack on the fleet already publishes a service on the same port. Move one of the services to a different port and redeploy, or leave the second stack out of the mesh.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="A route shows `unreachable`">
|
||||
The tunnel is up but the destination port did not answer. Check that the target stack is running and that its service is listening on the declared port. Click **Test** to see the exact failing stage.
|
||||
<Accordion title="A node card shows `pilot offline`">
|
||||
The Pilot Agent tunnel to that node is not connected. Open **Fleet → Overview** for the node's status and follow the [Pilot Agent troubleshooting](/features/pilot-agent#troubleshooting) entries. Mesh recovers automatically when the tunnel reconnects; no opt-out is needed.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="A route shows `not authorized`">
|
||||
The destination stack is not opted into the mesh on its home node. Open Routing on that node and add the stack.
|
||||
<Accordion title="A node card shows `unreachable`">
|
||||
The node card hosts the reason text directly. While a node is unreachable the mesh toggle and `Add stack to mesh` action on its card are disabled so a redeploy is not triggered against a broken target. Common causes for proxy-mode peers:
|
||||
|
||||
- `api token rejected by remote`: the credential central uses to dial this node is not accepted. Open the remote Sencho, generate a fresh Node Token via **Settings → Nodes → Generate Token**, and paste it back into the node's credentials in central's **Settings → Nodes**.
|
||||
- `remote does not support proxy mesh`: the remote Sencho is on a version that does not implement proxy-mode mesh. Update the remote to the same build as central and the badge clears on the next refresh.
|
||||
- `TLS handshake failed`: the remote serves a certificate Node's default trust store does not accept. Use a certificate issued by a trusted authority on the remote.
|
||||
- `api_url not set` or `api token missing`: the node was added without credentials. Edit the node in **Settings → Nodes** and supply the URL and token.
|
||||
- `remote unreachable`: central could not establish a TCP connection to the remote Sencho. Check that the remote is running and that the URL in **Settings → Nodes** is reachable from central.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="The Test button returns a stage I don't recognise">
|
||||
The probe reports the exact stage it failed at:
|
||||
|
||||
- `no_route`: the alias is not registered. Confirm the destination stack is opted in on its home node, and that the stack is running.
|
||||
- `pilot_tunnel`: no bridge to the destination node could be opened. Open Diagnostics on the destination node card to inspect the tunnel state.
|
||||
- `agent_resolve`: the bridge is up but the remote refuses the alias. The destination stack is not opted in on its home node.
|
||||
- `agent_dial`: central reached the remote but the remote could not connect to the target container.
|
||||
- `target_port`: the remote dialed the target container but no service answered on the declared port. Confirm the service is running and listening.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Adding a stack hangs on redeploy">
|
||||
Mesh redeploys the stack on opt-in to refresh hostnames and network attachments. A stuck redeploy usually means the stack itself failed to come back up. Check the stack's deploy logs.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="A Distributed API node shows `unreachable` on the Routing tab">
|
||||
Central could not open a mesh tunnel to this node. The badge tooltip shows the specific reason. While a node is in this state the **mesh toggle and Add stack to mesh action on the node card are disabled** so a redeploy is not triggered against an unreachable target. Common causes:
|
||||
|
||||
- `api token rejected by remote` — the credential central uses to dial this node is not accepted. Open the remote Sencho, generate a fresh Node Token via **Settings → Nodes → Generate Token**, and paste it back into the node's credentials in central's **Settings → Nodes**.
|
||||
- `remote does not support proxy mesh` — the remote Sencho is on a version that predates proxy-mode mesh. Update the remote and the badge clears on the next refresh.
|
||||
- `TLS handshake failed` — the remote serves a certificate Node's default trust store does not accept. Use a certificate issued by a trusted authority on the remote.
|
||||
- `api_url not set` or `api token missing` — the node was added without credentials. Edit the node in **Settings → Nodes** and supply the URL and token.
|
||||
The opt-in flow waits for the redeploy to complete before clearing the spinner. A stuck redeploy usually means the stack itself failed to come back up after the override was injected. Check the stack's deploy logs.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Stack topology sheet says 'No published mesh services'">
|
||||
The stack is opted into the mesh but exposes no service ports that became aliases. The stack joins `sencho_mesh` (other meshed containers can talk to it directly by container name) but no fleet-wide hostname is published. To publish an alias, declare a port on a service in the stack's compose file and redeploy.
|
||||
The stack is opted in but exposes no service ports. The stack joins `sencho_mesh` (other meshed containers can reach it directly by container name on that network) but no fleet-wide hostname is published. To publish an alias, declare a port on a service in the stack's compose file and redeploy.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="A stack on the Routing tab shows a `suspended` pill">
|
||||
The stack is opted into the mesh, but its services are not currently running, so there are no aliases to publish. The opt-in is sticky: when the stack starts again, its aliases reappear automatically on the next refresh (within roughly one minute) without needing a manual opt-out and re-opt-in. To clear the suspended state, start the stack from its **Overview** page. To remove the opt-in entirely, open the node card and use the opt-out action.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Mesh activity is empty after restart">
|
||||
The mesh activity log is an in-memory ring buffer. It resets when Sencho restarts. The audit log (under the top-nav **Audit** view) retains opt-in and opt-out events across restarts; for tunnel-state and probe-level events you need to watch the activity log live or export the logs from your reverse proxy.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="A `proxy` peer keeps failing immediately after a remote upgrade">
|
||||
After upgrading the remote Sencho, central may cache a failed dial for up to 60 seconds. Click **Refresh** on the Routing tab or wait one minute. If the failure persists, follow the `unreachable` entry above; check that the Node Token on the remote is still valid (Settings → Nodes → Generate Token).
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Graph reflects a stale node state">
|
||||
The Routing tab polls `/mesh/status` and `/mesh/aliases` every 30 seconds while the browser tab is focused. To force an immediate refresh, leave and return to the Routing tab, or toggle any stack's mesh state to trigger an action-driven refresh. Polling pauses when the tab is hidden, so a long-dormant tab catches up on the first poll after it regains focus.
|
||||
The Routing tab polls `/mesh/status` and `/mesh/aliases` every 30 seconds while the browser tab is focused. To force an immediate refresh, toggle any stack's mesh state to trigger an action-driven refresh, or leave and return to the tab. Polling pauses when the tab is hidden, so a long-dormant tab catches up on the first poll after it regains focus.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Graph mode is hard to read with a large fleet">
|
||||
The diagram suits typical fleet sizes of up to roughly 50 nodes. Larger fleets render but the layout becomes dense. Use the Table view for inventory at scale and reach for Graph mode for spot checks of tunnel state and alias publication.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Common questions
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Do I need to open a port on my router?">
|
||||
Not for Pilot Agent nodes. The agent dials outbound to central and carries mesh traffic in both directions over that tunnel. For Distributed API Proxy nodes, central needs to reach the remote Sencho's existing API port; if the remote is behind NAT, either expose that one port or run the remote in Pilot Agent mode instead.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Is the alias DNS internet-reachable?">
|
||||
No. Aliases resolve only inside the `sencho_mesh` Docker bridge on each participating node. Nothing about the mesh adds public DNS records or exposes services on the public internet. You can still front a meshed service with a public reverse proxy as normal; that proxy then resolves the alias internally.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Does opting in change my compose file?">
|
||||
No. Sencho writes a separate override file (`<DATA_DIR>/mesh/overrides/<nodeId>/<stack>.override.yml`) and applies it at deploy time alongside the stack's own compose file. The override is removed on opt-out. Your stack's compose file is unchanged.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="What happens if two stacks expose the same port?">
|
||||
The mesh enforces one alias per TCP port across the fleet. The second stack to try to opt in returns a 409 with the message `Port already claimed by another mesh stack`. Move one service to a different port to mesh both.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Is Mesh the same as Federation?">
|
||||
No. [Federation](/features/fleet-federation) decides *where* a blueprint deploys (cordon, pin, blueprint placement). Mesh handles *how* containers on different nodes talk to each other once they are running. They are independent and compose freely.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Can I run Mesh without Pilot Agent?">
|
||||
Yes. Mesh works with any enrolled remote node. Distributed API Proxy peers get a dedicated mesh bridge that central dials over the remote's existing API port. Pilot Agent peers get mesh traffic on the same tunnel they already use for fleet operations. The choice is independent.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## Where Mesh fits
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Pilot Agent" icon="link" href="/features/pilot-agent">
|
||||
Outbound-only remote mode. Carries mesh traffic on the existing agent tunnel.
|
||||
</Card>
|
||||
<Card title="Multi-Node Management" icon="server" href="/features/multi-node">
|
||||
Enrol remote nodes. Generate the Node Token that authorises proxy-mode mesh dials.
|
||||
</Card>
|
||||
<Card title="Fleet Federation" icon="network-wired" href="/features/fleet-federation">
|
||||
Steer blueprint placement across the same fleet Mesh networks together.
|
||||
</Card>
|
||||
<Card title="Licensing" icon="key" href="/features/licensing">
|
||||
Mesh ships in the Admiral tier. See what else Admiral unlocks.
|
||||
</Card>
|
||||
</CardGroup>
|
||||
|
||||
@@ -1,12 +1,12 @@
|
||||
---
|
||||
title: Stack Sidebar
|
||||
description: Manage, group, and pin your stacks from the primary sidebar.
|
||||
description: Manage, group, pin, and bulk-act on your stacks from the primary sidebar.
|
||||
---
|
||||
|
||||
The stack sidebar is your cockpit for every stack on the active node. Stacks are grouped by label, your most-used stacks can be pinned to the top, and a live activity footer keeps you aware of what just happened.
|
||||
The stack sidebar is your cockpit for every stack on the active node. Stacks are grouped by label, your most-used stacks can be pinned to the top, a one-click bulk mode lets you act on several at once, and a live activity footer keeps you aware of what just happened.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/sidebar/sidebar-overview.png" alt="Stack sidebar overview" />
|
||||
<img src="/images/sidebar/sidebar-overview.png" alt="Stack sidebar overview with branding header, node switcher, Create Stack button, Bulk mode and Scan icons, search box, ALL/UP/DOWN/UPDATES filter chips, three label groups (MEDIA, UTILITIES, NETWORK) with stack rows and label dots, and a LIVE activity footer at the bottom" />
|
||||
</Frame>
|
||||
|
||||
## Layout
|
||||
@@ -15,39 +15,86 @@ From top to bottom:
|
||||
|
||||
1. **Branding header** shows your Sencho build version.
|
||||
2. **Node switcher** selects which Sencho instance you are managing.
|
||||
3. **Create Stack** opens the new-stack dialog. The folder icon scans your compose directory.
|
||||
4. **Search** filters stacks by name. Press <kbd>⌘K</kbd> (or <kbd>Ctrl+K</kbd>) to focus it.
|
||||
5. **Stack list** shows your stacks grouped by label.
|
||||
6. **Activity footer** shows the most recent stack event.
|
||||
3. **Action row** holds **Create Stack**, the **Bulk mode** toggle (square icon), and **Scan stacks folder** (folder-search icon).
|
||||
4. **Search** filters the stack list. Start typing into **Search stacks...** to filter in place.
|
||||
5. **Filter chips** narrow the list by status (described below).
|
||||
6. **Stack list** shows your stacks grouped by label.
|
||||
7. **Activity footer** shows the most recent stack event on the node.
|
||||
|
||||
## Filter chips
|
||||
|
||||
Four chips sit below the search box. Each shows a live count to the right of its label, and counts above 99 render as **99+** to keep the row from overflowing.
|
||||
|
||||
- **All**: every stack on the node.
|
||||
- **Up**: stacks where at least one container is running.
|
||||
- **Down**: stacks where no container is running.
|
||||
- **Updates**: stacks with at least one image update available. The chip turns orange when the count is non-zero so you can spot pending updates at a glance.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/sidebar/sidebar-filter-chips.png" alt="Filter chip row with ALL pressed (count 14), UP (count 14), DOWN (count 0), and UPDATES highlighted in orange with count 1" />
|
||||
</Frame>
|
||||
|
||||
Click the **−** (minus) icon on the right to collapse the chip row when you need more vertical room; it becomes a **+** (plus) icon you click to bring the chips back. The collapsed state is remembered in your browser.
|
||||
|
||||
## Groups
|
||||
|
||||
Stacks are grouped by label. Stacks with multiple labels appear in each matching group so you always see the full fleet membership per label. Stacks without labels appear in an **Unlabeled** group at the bottom. Click a group header to collapse or expand it; the collapsed state is remembered per node.
|
||||
Stacks are grouped by label. Groups are sorted by size (most populated first), then alphabetically. Stacks with multiple labels appear in every matching group so you always see the full fleet membership per label. Stacks without labels appear in an **Unlabeled** group at the bottom. Click a group header to collapse or expand it; the collapsed state is remembered per node.
|
||||
|
||||
## Pinning
|
||||
|
||||
Right-click a stack and choose **Pin to top**. Pinned stacks sit in a dedicated **PINNED** group at the top of the list. Up to 10 stacks can be pinned per node; pinning an 11th evicts the oldest.
|
||||
Right-click a stack and choose **Pin to top**. Pinned stacks sit in a dedicated **★ PINNED** group at the top of the list. Up to 10 stacks can be pinned per node; pinning an 11th evicts the oldest. Right-click a pinned stack and choose **Unpin** to remove it from the group.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/sidebar/sidebar-pinned.png" alt="Pinned stacks appear in a PINNED group at top" />
|
||||
<img src="/images/sidebar/sidebar-pinned.png" alt="Sidebar with a starred PINNED group at the top containing plex, sonarr, and cloudflared, followed by the regular MEDIA group below" />
|
||||
</Frame>
|
||||
|
||||
## Stack rows
|
||||
|
||||
Each row gives you everything you need to read the stack at a glance, in a fixed column order:
|
||||
|
||||
- **Status pill** on the left. Two uppercase letters in mono type, or a spinner while a lifecycle action is in flight. `UP` means at least one container is running; `DN` means the stack is stopped.
|
||||
- **Stack name** in mono type, truncated with an ellipsis when the row gets tight.
|
||||
- **Label dots** to the right of the name. Up to three colored dots representing the stack's labels render here. If a stack carries more than three labels, a **+N** counter appears for the extras.
|
||||
- **Update indicator**. When a stack has an image update pending, an extra colored dot appears alongside the label dots. If only a Git source update is pending (no image update), a small Git branch icon shows instead. The image-update dot takes priority when both apply.
|
||||
- **Hover kebab** on the right edge. Hover the row to reveal a vertical three-dot menu that opens the same actions as right-clicking the row.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/sidebar/sidebar-row-anatomy.png" alt="Two stack rows shown stacked: dozzle with two label dots (utilities in grey and media in magenta) and a hover kebab, and cloudflared with an orange indicator dot and a hover kebab" />
|
||||
</Frame>
|
||||
|
||||
## Bulk mode
|
||||
|
||||
Click the **Bulk mode** icon next to **Create Stack** (or press <kbd>B</kbd>) to enter selection mode. A checkbox appears at the left of every row, and a sticky toolbar slides in just above the list:
|
||||
|
||||
- **Start**, **Stop**, **Restart** apply the action to every selected stack.
|
||||
- **Update** pulls the latest images for every selected stack and requires a **Skipper** or **Admiral** license.
|
||||
|
||||
Click rows to toggle their selection. The toolbar header shows a running count. Click the **×** in the corner of the toolbar to clear the selection, and click the icon again (or press <kbd>B</kbd>) to leave bulk mode.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/sidebar/sidebar-bulk-mode.png" alt="Sidebar in bulk mode with the Bulk mode icon highlighted, a sticky toolbar reading 3 selected with Start, Stop, Restart, and Update buttons, and checkboxes visible on every stack row" />
|
||||
</Frame>
|
||||
|
||||
## Context menu
|
||||
|
||||
Right-click any stack (or open the kebab that appears on hover) for its context menu, grouped by purpose:
|
||||
Right-click any stack (or open the kebab that appears on hover) for its context menu. Items are grouped by purpose:
|
||||
|
||||
- **Inspect**: view alerts, open the auto-heal sheet, check for image updates, open the running app.
|
||||
- **Organize**: assign labels, pin or unpin the stack.
|
||||
- **Lifecycle**: deploy, stop, restart, or update the stack.
|
||||
- **Destructive**: delete the stack.
|
||||
- **Inspect**: **Alerts**, **Auto-Heal**, the **Auto-update** toggle, **Check updates**, and **Open App** (the last only appears when the stack is running and exposes a port).
|
||||
- **Organize**: **Labels** (a submenu listing every label on the node, with a check next to each label the stack already carries) and **Pin to top** / **Unpin**.
|
||||
- **Lifecycle**: **Deploy** (when stopped), **Stop**, **Restart**, **Update** (when running), and **Schedule task**.
|
||||
- **Destructive**: **Delete**.
|
||||
|
||||
<Note>
|
||||
**Auto-Heal**, the **Auto-update** toggle, and **Schedule task** require a **Skipper** or **Admiral** license.
|
||||
</Note>
|
||||
|
||||
<Frame>
|
||||
<img src="/images/sidebar/sidebar-context-menu.png" alt="Grouped stack context menu" />
|
||||
<img src="/images/sidebar/sidebar-context-menu.png" alt="Right-click context menu on a running plex stack with four sections. INSPECT lists Alerts, Auto-Heal, Auto-update: Enabled, Check updates, and Open App. ORGANIZE lists Labels with a submenu arrow and Pin to top. LIFECYCLE lists Stop, Restart, Update, and Schedule task. DESTRUCTIVE lists Delete in red." />
|
||||
</Frame>
|
||||
|
||||
## Keyboard shortcuts
|
||||
|
||||
Every action in the context menu has a keyboard shortcut. Shortcuts fire on the currently selected stack and are blocked when a text input is focused or the command palette is open.
|
||||
Shortcuts fire on the currently selected stack. They are blocked while a text input is focused or the global command palette is open.
|
||||
|
||||
### Lifecycle
|
||||
|
||||
@@ -65,15 +112,35 @@ On macOS, use <kbd>Cmd</kbd> in place of <kbd>Ctrl</kbd>.
|
||||
|
||||
| Key | Action |
|
||||
|-----|--------|
|
||||
| <kbd>A</kbd> | Open alerts sheet |
|
||||
| <kbd>H</kbd> | Open auto-heal sheet (Skipper and above) |
|
||||
| <kbd>A</kbd> | Open the alerts sheet |
|
||||
| <kbd>H</kbd> | Open the auto-heal sheet (Skipper or Admiral) |
|
||||
| <kbd>U</kbd> | Check for image updates |
|
||||
| <kbd>P</kbd> | Pin or unpin the stack |
|
||||
| <kbd>B</kbd> | Toggle bulk mode |
|
||||
|
||||
The **↗** and **L ›** glyphs inside the context menu are visual hints, not global bindings. Open **Open App** or **Labels** through the menu (or the hover kebab) to use them.
|
||||
|
||||
## Activity footer
|
||||
|
||||
The footer shows the most recent stack lifecycle event within the last hour. Click it to jump to the global activity log. When nothing has happened recently, the footer reads **IDLE**.
|
||||
The footer surfaces the most recent stack lifecycle event on the node. Each ticker shows the stack name in brand color, a short event description, and a relative time. The kicker text below alternates between **LIVE · VIEW ACTIVITY →** when an event is showing and **IDLE · NO RECENT ACTIVITY** when nothing has happened recently. A pulsing dot to the left indicates the WebSocket subscription is healthy. Click anywhere on the footer to open the global activity view.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/sidebar/sidebar-activity.png" alt="Activity footer ticker" />
|
||||
<img src="/images/sidebar/sidebar-activity.png" alt="Activity footer with a green pulsing dot, the stack name 'cloudflared' in brand color, an event description, and the kicker LIVE · VIEW ACTIVITY" />
|
||||
</Frame>
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="I can't see my newly created stack in the sidebar">
|
||||
Confirm the active filter chip is **All** so a status filter is not hiding the row. Clear the **Search stacks...** box in case a stale query is filtering the list. Then click the **Scan stacks folder** icon next to **Create Stack** to re-index the compose directory on disk.
|
||||
</Accordion>
|
||||
<Accordion title="Pinning is not sticking past the 10th stack">
|
||||
Each node has a 10-stack pin limit. When you pin an 11th stack, the oldest pin is automatically evicted so the new one fits. Unpin a stack you no longer need before adding another, or accept that the oldest will roll off.
|
||||
</Accordion>
|
||||
<Accordion title="Keyboard shortcuts are not firing">
|
||||
Shortcuts only fire when a stack row is selected in the sidebar (one row is highlighted). They are intentionally blocked while a text input is focused or the global command palette (<kbd>Ctrl</kbd>+<kbd>K</kbd>) is open, so they do not collide with typing. Click a stack row, then try the shortcut again.
|
||||
</Accordion>
|
||||
<Accordion title="The filter chips disappeared from the sidebar">
|
||||
Click the **+** icon to the right of the search box to bring them back. The chip row collapses to a thin **−** / **+** toggle, and the state is remembered in your browser, so an earlier collapse persists across reloads until you expand it again.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
@@ -3,7 +3,9 @@ title: SSO & LDAP Authentication
|
||||
description: Authenticate with your existing identity provider, including LDAP, Google, GitHub, Okta, and any spec-compliant OIDC provider.
|
||||
---
|
||||
|
||||
Sencho lets your team sign in using existing identity providers instead of managing separate credentials. SSO works **alongside** password authentication; it does not replace it. SSO is available in every Sencho edition; higher tiers add preset providers and enterprise directory support.
|
||||
Sencho lets your team sign in with the identity provider you already use instead of maintaining a second set of credentials. SSO works **alongside** password authentication; it does not replace it.
|
||||
|
||||
SSO is available on every Sencho tier. Higher tiers add preset providers for Google, GitHub, and Okta, plus enterprise directory support via LDAP and Active Directory.
|
||||
|
||||
## Supported providers
|
||||
|
||||
@@ -15,48 +17,42 @@ Sencho lets your team sign in using existing identity providers instead of manag
|
||||
| **Okta** | OpenID Connect | Skipper | Preset for any Okta org or Okta-compatible IdP, with branded login button |
|
||||
| **LDAP / Active Directory** | LDAP bind + search | Admiral | Works with OpenLDAP, Active Directory, FreeIPA, and any LDAPv3 server |
|
||||
|
||||
<Tip>
|
||||
**Self-hosters on the Community tier** can still integrate Google, GitHub, Okta, or any other identity provider by using the Custom OIDC option pointed at the provider's discovery endpoint. The paid-tier presets ship with provider-aware defaults (issuer URL, claim mapping) and a branded login button; the underlying protocol is the same. You still create the OAuth app in the provider's console and paste the Client ID and Secret either way.
|
||||
</Tip>
|
||||
|
||||
## How it works
|
||||
|
||||
### LDAP flow
|
||||
|
||||
1. User enters their directory username and password on the Sencho login page
|
||||
2. Sencho binds to LDAP with a service account, searches for the user, then verifies their password
|
||||
3. If this is their first login, a Sencho account is automatically created
|
||||
4. Sencho issues a JWT and the user is logged in, identical to a password login
|
||||
1. User enters their directory username and password on the Sencho login page.
|
||||
2. Sencho binds to LDAP with a service account, locates the user, then verifies their password.
|
||||
3. On the user's first login, a Sencho account is created automatically.
|
||||
4. Sencho issues a session JWT and the user is logged in, identical to a password login.
|
||||
|
||||
### OIDC / OAuth flow (Google, GitHub, Okta, Custom OIDC)
|
||||
|
||||
1. User clicks the provider button on the login page (e.g., "Sign in with Google")
|
||||
2. Browser redirects to the identity provider for authentication
|
||||
3. After granting consent, the provider redirects back to Sencho with an authorization code
|
||||
4. Sencho exchanges the code for tokens, verifies the ID token, and reads user information
|
||||
5. If this is their first login, a Sencho account is automatically created
|
||||
6. Sencho issues a JWT and redirects to the dashboard
|
||||
|
||||
All OIDC flows use **PKCE** (Proof Key for Code Exchange) and a **state parameter** for CSRF protection.
|
||||
1. User clicks the provider button on the login page (e.g., **Sign in with Google**).
|
||||
2. The browser is redirected to the identity provider for authentication.
|
||||
3. After granting consent, the provider redirects back to Sencho with an authorization code.
|
||||
4. Sencho exchanges the code for tokens, verifies the ID token, and reads user information.
|
||||
5. On the user's first login, a Sencho account is created automatically.
|
||||
6. Sencho issues a session JWT and the user lands on the dashboard.
|
||||
|
||||
## Auto-provisioning
|
||||
|
||||
When a user logs in via SSO for the first time, Sencho automatically creates a local account:
|
||||
When a user signs in via SSO for the first time, Sencho creates a local account:
|
||||
|
||||
- **Username** is derived from their identity provider profile (display name, email prefix, or login handle)
|
||||
- **Role** is assigned based on [role mapping](#role-mapping); defaults to Viewer if no mapping matches
|
||||
- **Password** is set to an unusable placeholder. SSO users cannot log in with a password
|
||||
- **Seat limits** from your license apply. If admin seats are full, the user is downgraded to Viewer. If all seats are full, login is denied with a clear error message.
|
||||
- **Username** is derived from the identity provider profile (display name, email prefix, or login handle).
|
||||
- **Role** is assigned from [role mapping](#role-mapping); defaults to Viewer if no mapping matches.
|
||||
- **Password** is set to an unusable placeholder. SSO users cannot sign in with the password form.
|
||||
- **Seat limits** from your license apply. If admin seats are full, the user is downgraded to Viewer. If every seat is full, sign-in is denied with a clear error message.
|
||||
|
||||
On subsequent logins, the existing account is reused. The user's **email** and **role** are synced from the identity provider on every login. If a user is added to your admin group, they will be promoted to Admin on their next login. If removed, they will be demoted to their default role. Seat limits are respected: if admin seats are full, the promotion is deferred until a seat opens up.
|
||||
On every subsequent sign-in, the existing account is reused and the user's **email** and **role** are synced from the identity provider. Adding someone to your admin group promotes them to Admin on their next sign-in; removing them demotes them to the default role. Promotions defer cleanly when admin seats are full and apply as soon as a seat opens up.
|
||||
|
||||
## Role mapping
|
||||
|
||||
### LDAP group mapping
|
||||
|
||||
Set the **Admin Group DN** to a group in your directory. Members of that group get the Admin role; everyone else gets the default role (Viewer).
|
||||
Set the **Admin Group DN** to a group in your directory. Members of that group receive the Admin role; everyone else receives the default role (Viewer).
|
||||
|
||||
Example: If your admin group is `cn=sencho-admins,ou=groups,dc=example,dc=com`, set that as the Admin Group DN. Users who are a `member` of that group will be provisioned as Admin.
|
||||
Example: if your admin group is `cn=sencho-admins,ou=groups,dc=example,dc=com`, set that as the **Admin Group DN**. Users whose `memberOf` attribute lists that DN are provisioned as Admin.
|
||||
|
||||
### OIDC claim mapping
|
||||
|
||||
@@ -64,37 +60,43 @@ For OIDC providers, configure two fields:
|
||||
|
||||
| Field | Description | Example |
|
||||
|-------|-------------|---------|
|
||||
| **Admin Claim** | The JWT claim name that contains role information | `groups` |
|
||||
| **Admin Claim** | The token claim name that contains role information | `groups` |
|
||||
| **Admin Claim Value** | The value within that claim that grants Admin | `sencho-admins` |
|
||||
|
||||
If the user's ID token contains a `groups` claim with the value `sencho-admins`, they get Admin. Otherwise, they get the default role.
|
||||
If the user's ID token contains a `groups` claim with the value `sencho-admins`, they receive Admin. Otherwise, they receive the default role.
|
||||
|
||||
<Note>
|
||||
Not all providers include a `groups` claim by default. You may need to configure custom claims in your identity provider's admin console.
|
||||
Not every provider includes a `groups` claim by default. You may need to configure custom claims or scopes in your identity provider's admin console to surface group membership in the ID token.
|
||||
</Note>
|
||||
|
||||
## Configuration
|
||||
|
||||
SSO can be configured two ways:
|
||||
|
||||
1. **Settings UI** - Go to **Settings > SSO** in the Sencho dashboard. Enable providers, enter credentials, and test connections from the UI. Changes take effect immediately without restarting.
|
||||
2. **Environment variables** - Set `SSO_*` variables in your Docker Compose file. These seed the database on first boot. After that, the database configuration is authoritative.
|
||||
1. **Settings UI**: go to **Settings → SSO** in the Sencho dashboard. Enable providers, paste credentials, and test connections from the UI. Changes take effect immediately, no restart required.
|
||||
2. **Environment variables**: set `SSO_*` variables in your Docker Compose file. They seed the database on first boot; afterwards the Settings UI is authoritative.
|
||||
|
||||
### Via Settings UI
|
||||
|
||||
Admins can manage SSO providers in **Settings > SSO**. Each provider is displayed as a collapsible card with:
|
||||
|
||||
- An **enable/disable** toggle and an **Active** badge when enabled
|
||||
- Provider-specific configuration fields (expand the card to configure)
|
||||
- A **Save** button to persist changes
|
||||
- A **Test Connection** button to verify connectivity before saving
|
||||
- A **Remove** button to delete an existing provider configuration
|
||||
Admins manage SSO providers in **Settings → SSO**. The page lists every provider as a collapsible card with a label, an **enable / disable** toggle pill on the right, and an **Active** badge on the header when the provider is on.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/sso/sso-settings.png" alt="SSO settings panel showing all five identity providers" />
|
||||
<img src="/images/sso/sso-settings.png" alt="SSO settings panel listing the five identity providers as collapsible cards with enable / disable toggles" />
|
||||
</Frame>
|
||||
|
||||
Expand a provider card to configure it. The LDAP configuration form includes:
|
||||
Click a card to expand it. The footer of every expanded form has the same actions:
|
||||
|
||||
- **Save**: persists changes for that provider.
|
||||
- **Test Connection**: runs a live check (LDAP bind for LDAP, OIDC discovery for the OIDC providers) and shows a green check or red X next to the button with the result.
|
||||
- **Remove**: clears the saved configuration (only present once a config has been saved at least once).
|
||||
|
||||
A static helper sits below all five cards. It reminds you that SSO users are auto-provisioned on first login and shows the OAuth callback URL template for OIDC providers:
|
||||
|
||||
```
|
||||
https://<your-sencho-url>/api/auth/sso/oidc/<provider>/callback
|
||||
```
|
||||
|
||||
### LDAP fields
|
||||
|
||||
| Field | Description |
|
||||
|-------|-------------|
|
||||
@@ -102,52 +104,60 @@ Expand a provider card to configure it. The LDAP configuration form includes:
|
||||
| **Bind DN** | Service account DN used to search the directory |
|
||||
| **Bind Password** | Service account password |
|
||||
| **Search Base** | Base DN for user searches (e.g., `ou=users,dc=example,dc=com`) |
|
||||
| **Search Filter** | LDAP filter template using `{{username}}` as placeholder |
|
||||
| **Search Filter** | LDAP filter template using `{{username}}` as a placeholder |
|
||||
| **Admin Group DN** | DN of the group whose members receive the Admin role |
|
||||
| **Default Role** | Role assigned to users not in the admin group (Viewer or Admin) |
|
||||
| **Verify TLS certificate** | Toggle to enable or disable TLS certificate verification |
|
||||
|
||||
For Active Directory, set **Search Filter** to `(sAMAccountName={{username}})`. The form's helper text shows the same example inline.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/sso/sso-settings-ldap.png" alt="LDAP configuration form with server URL, bind DN, search base, and role mapping" />
|
||||
<img src="/images/sso/sso-settings-ldap.png" alt="LDAP / Active Directory configuration form with server URL, bind DN, search base, role mapping, and Verify TLS toggle" />
|
||||
</Frame>
|
||||
|
||||
OIDC providers (Google, GitHub, Okta) share a common configuration form:
|
||||
### OIDC fields (Google, GitHub, Okta)
|
||||
|
||||
| Field | Description |
|
||||
|-------|-------------|
|
||||
| **Issuer URL** | (Okta and Custom OIDC) Your provider's OIDC issuer URL |
|
||||
| **Issuer URL** | (Okta only) Your Okta org issuer URL, e.g., `https://dev-123456.okta.com` |
|
||||
| **Client ID** | OAuth client ID from your identity provider |
|
||||
| **Client Secret** | OAuth client secret |
|
||||
| **Admin Claim** | JWT claim name inspected for role mapping (e.g., `groups`) |
|
||||
| **Admin Claim Value** | Value within the claim that grants Admin (e.g., `sencho-admins`) |
|
||||
| **Scopes** | Space-separated OAuth scopes (default: `openid email profile`). Customize if your provider requires additional scopes for group claims. |
|
||||
| **Admin Claim** | Token claim name inspected for role mapping (default: `groups`) |
|
||||
| **Admin Claim Value** | Value within the claim that grants Admin (default: `sencho-admins`) |
|
||||
| **Scopes** | Space-separated OAuth scopes (default: `openid email profile`). Customize if your provider needs additional scopes to emit group claims. |
|
||||
| **Default Role** | Role assigned when no claim mapping matches (Viewer or Admin) |
|
||||
|
||||
Google and GitHub already know their own issuer URL, so the form omits that field for those providers.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/sso/sso-settings-oidc.png" alt="Google OIDC configuration form with client ID, client secret, and role claim mapping" />
|
||||
<img src="/images/sso/sso-settings-oidc.png" alt="Google OIDC configuration form with client ID, client secret, admin claim mapping, scopes, and default role" />
|
||||
</Frame>
|
||||
|
||||
### Custom OIDC configuration
|
||||
### Custom OIDC fields
|
||||
|
||||
The **Custom OIDC** provider has additional fields beyond the standard OIDC configuration:
|
||||
The **Custom OIDC** form adds the fields needed to point Sencho at a self-hosted or third-party identity provider:
|
||||
|
||||
| Field | Description |
|
||||
|-------|-------------|
|
||||
| **Display Name** | Label shown on the login button (e.g., "Corporate SSO") |
|
||||
| **Display Name** | Label shown on the login button (e.g., `Corporate SSO`) |
|
||||
| **Issuer URL** | Base URL of the OIDC discovery endpoint (without `/.well-known/openid-configuration`) |
|
||||
| **User ID Claim** | Claim name for the unique user identifier (default: `sub`) |
|
||||
| **Username Claim** | Claim name for the display name (default: `preferred_username`) |
|
||||
| **Email Claim** | Claim name for the email address (default: `email`) |
|
||||
| **User ID Claim** | Claim used as the unique user identifier (default: `sub`) |
|
||||
| **Username Claim** | Claim used for the display name (default: `preferred_username`) |
|
||||
| **Email Claim** | Claim used for the email address (default: `email`) |
|
||||
|
||||
The claim mapping fields let you tell Sencho which token claims correspond to user identity fields. Most spec-compliant providers use the standard claim names, so you can leave these blank unless your provider uses non-standard names.
|
||||
Leave the three claim fields blank to use the standard OIDC defaults. Most spec-compliant providers will work out of the box; override only when your provider emits non-standard claim names.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/sso/sso-settings-custom-oidc.png" alt="Custom OIDC configuration form with discovery URL, claim mapping, and role settings" />
|
||||
<img src="/images/sso/sso-settings-custom-oidc.png" alt="Custom OIDC configuration form with display name, issuer URL, claim mapping fields, scopes, and default role" />
|
||||
</Frame>
|
||||
|
||||
<Note>
|
||||
The **User ID Claim**, **Username Claim**, and **Email Claim** fields are also accepted on Google, GitHub, and Okta as environment variables (see [Custom OIDC env vars](#custom-oidc)). The Settings UI hides them on the presets because the defaults match those providers; reach for them only if you have a custom claim layout to map.
|
||||
</Note>
|
||||
|
||||
### Via environment variables
|
||||
|
||||
Environment variables are useful for initial deployment or infrastructure-as-code workflows. They seed the SSO configuration on first startup. After that, changes made in the Settings UI take precedence.
|
||||
Environment variables are useful for initial deployment and infrastructure-as-code workflows. They seed the SSO configuration on first startup. After that, changes made in the Settings UI take precedence.
|
||||
|
||||
## SSO environment variables reference
|
||||
|
||||
@@ -156,7 +166,7 @@ Environment variables are useful for initial deployment or infrastructure-as-cod
|
||||
| Variable | Default | Description |
|
||||
|----------|---------|-------------|
|
||||
| `SSO_LDAP_ENABLED` | `false` | Enable LDAP authentication |
|
||||
| `SSO_LDAP_DISPLAY_NAME` | `LDAP` | Label shown on the login button (e.g., "Corporate AD") |
|
||||
| `SSO_LDAP_DISPLAY_NAME` | `LDAP` | Label shown on the login button (e.g., `Corporate AD`) |
|
||||
| `SSO_LDAP_URL` | - | LDAP server URL (e.g., `ldap://ldap.example.com:389` or `ldaps://ldap.example.com:636`) |
|
||||
| `SSO_LDAP_BIND_DN` | - | Service account DN for searching users |
|
||||
| `SSO_LDAP_BIND_PASSWORD` | - | Service account password (encrypted at rest in the database) |
|
||||
@@ -164,7 +174,7 @@ Environment variables are useful for initial deployment or infrastructure-as-cod
|
||||
| `SSO_LDAP_SEARCH_FILTER` | `(uid={{username}})` | LDAP filter template. Use `(sAMAccountName={{username}})` for Active Directory |
|
||||
| `SSO_LDAP_ADMIN_GROUP_DN` | - | DN of the group whose members receive the Admin role |
|
||||
| `SSO_LDAP_DEFAULT_ROLE` | `viewer` | Role assigned to LDAP users not in the admin group |
|
||||
| `SSO_LDAP_TLS_REJECT_UNAUTHORIZED` | `true` | Whether to verify the LDAP server's TLS certificate |
|
||||
| `SSO_LDAP_TLS_REJECT_UNAUTHORIZED` | `true` | Set to the literal string `false` to skip TLS certificate verification (useful for self-signed certs in development) |
|
||||
|
||||
### Google OIDC
|
||||
|
||||
@@ -205,11 +215,13 @@ Environment variables are useful for initial deployment or infrastructure-as-cod
|
||||
| `SSO_OIDC_CUSTOM_USERNAME_CLAIM` | `preferred_username` | Token claim for the display name |
|
||||
| `SSO_OIDC_CUSTOM_EMAIL_CLAIM` | `email` | Token claim for the email address |
|
||||
|
||||
The `*_ID_CLAIM`, `*_USERNAME_CLAIM`, and `*_EMAIL_CLAIM` variables are also accepted for Google, GitHub, and Okta (substitute the provider prefix, e.g., `SSO_OIDC_OKTA_USERNAME_CLAIM`). They override the per-provider defaults when your token layout differs.
|
||||
|
||||
### General
|
||||
|
||||
| Variable | Default | Description |
|
||||
|----------|---------|-------------|
|
||||
| `SSO_OIDC_ADMIN_CLAIM` | `groups` | JWT claim name inspected for Admin role mapping |
|
||||
| `SSO_OIDC_ADMIN_CLAIM` | `groups` | Token claim name inspected for Admin role mapping |
|
||||
| `SSO_OIDC_ADMIN_CLAIM_VALUE` | `sencho-admins` | Value in the admin claim that maps to the Admin role |
|
||||
| `SSO_DEFAULT_ROLE` | `viewer` | Default role for all SSO users when no mapping matches |
|
||||
| `SSO_CALLBACK_URL` | auto-detect | External base URL for OAuth callback URLs (see below) |
|
||||
@@ -220,7 +232,7 @@ Environment variables are useful for initial deployment or infrastructure-as-cod
|
||||
If Sencho is behind a reverse proxy (nginx, Traefik, Caddy), you **must** set `SSO_CALLBACK_URL` to your external URL. Otherwise, OAuth callbacks will fail.
|
||||
</Warning>
|
||||
|
||||
Set `SSO_CALLBACK_URL` to the URL users access Sencho from, for example, `https://sencho.example.com`. Sencho uses this to construct the OAuth redirect URI that your identity provider calls back to.
|
||||
Set `SSO_CALLBACK_URL` to the URL users use to reach Sencho, for example, `https://sencho.example.com`. Sencho uses this to construct the OAuth redirect URI that your identity provider calls back to.
|
||||
|
||||
If not set, Sencho auto-detects the URL from the request's `Host` header and protocol, which works for direct access but fails behind proxies that rewrite the host.
|
||||
|
||||
@@ -228,101 +240,82 @@ If not set, Sencho auto-detects the URL from the request's `Host` header and pro
|
||||
|
||||
### Keycloak
|
||||
|
||||
1. Create a new client in your Keycloak realm (Client type: **OpenID Connect**)
|
||||
2. Set **Valid redirect URIs** to `https://sencho.example.com/api/auth/sso/oidc/oidc_custom/callback`
|
||||
3. Enable **Client authentication** (confidential access type) and copy the client secret from the Credentials tab
|
||||
4. The Issuer URL is your realm URL: `https://keycloak.example.com/realms/myrealm`
|
||||
5. Keycloak uses standard claim names by default, so you can leave claim mapping blank
|
||||
1. Create a new client in your Keycloak realm (Client type: **OpenID Connect**).
|
||||
2. Set **Valid redirect URIs** to `https://sencho.example.com/api/auth/sso/oidc/oidc_custom/callback`.
|
||||
3. Enable **Client authentication** (confidential access type) and copy the client secret from the **Credentials** tab.
|
||||
4. The Issuer URL is your realm URL: `https://keycloak.example.com/realms/myrealm`.
|
||||
5. Keycloak uses standard claim names by default, so claim mapping can be left blank.
|
||||
|
||||
### Authentik
|
||||
|
||||
1. Create a new OAuth2/OpenID Provider in Authentik
|
||||
2. Set the redirect URI to `https://sencho.example.com/api/auth/sso/oidc/oidc_custom/callback`
|
||||
3. Copy the Client ID and Client Secret
|
||||
4. The Issuer URL is `https://authentik.example.com/application/o/<slug>/`
|
||||
5. Default claims work. For group-based admin mapping, configure a `groups` scope in Authentik
|
||||
1. Create a new **OAuth2 / OpenID Provider** in Authentik.
|
||||
2. Set the redirect URI to `https://sencho.example.com/api/auth/sso/oidc/oidc_custom/callback`.
|
||||
3. Copy the Client ID and Client Secret.
|
||||
4. The Issuer URL is `https://authentik.example.com/application/o/<slug>/`.
|
||||
5. Standard claims work. For group-based admin mapping, configure a `groups` scope in Authentik so it emits the claim into the ID token.
|
||||
|
||||
### Authelia
|
||||
|
||||
1. Add an OpenID Connect client to your Authelia configuration
|
||||
2. Set `redirect_uris` to include `https://sencho.example.com/api/auth/sso/oidc/oidc_custom/callback`
|
||||
3. The Issuer URL is your Authelia domain: `https://auth.example.com`
|
||||
4. Authelia uses standard OIDC claims
|
||||
1. Add an OpenID Connect client to your Authelia configuration under `identity_providers.oidc.clients`.
|
||||
2. Set `redirect_uris` to include `https://sencho.example.com/api/auth/sso/oidc/oidc_custom/callback`.
|
||||
3. The Issuer URL is your Authelia domain: `https://auth.example.com`.
|
||||
4. Authelia uses standard OIDC claims.
|
||||
|
||||
### Zitadel
|
||||
|
||||
1. Create a new Web application in your Zitadel project
|
||||
2. Add `https://sencho.example.com/api/auth/sso/oidc/oidc_custom/callback` as a redirect URI
|
||||
3. The Issuer URL is your Zitadel instance URL: `https://zitadel.example.com`
|
||||
4. Copy the Client ID and Client Secret from the application settings
|
||||
1. Create a new **Web** application in your Zitadel project.
|
||||
2. Add `https://sencho.example.com/api/auth/sso/oidc/oidc_custom/callback` as a redirect URI.
|
||||
3. The Issuer URL is your Zitadel instance URL: `https://zitadel.example.com`.
|
||||
4. Copy the Client ID and Client Secret from the application settings.
|
||||
|
||||
### KanIDM
|
||||
|
||||
1. Create a new OAuth2 client in KanIDM
|
||||
2. Set the redirect URI and copy the client credentials
|
||||
3. The Issuer URL is your KanIDM domain: `https://kanidm.example.com/oauth2/openid/<client_id>`
|
||||
4. KanIDM may use `name` instead of `preferred_username` for the username claim. Set **Username Claim** to `name` if usernames are not being mapped correctly.
|
||||
1. Create a new OAuth2 client in KanIDM (`kanidm system oauth2 create ...`).
|
||||
2. Add the redirect URL with `kanidm system oauth2 add-redirect-url ...` and fetch the basic secret with `kanidm system oauth2 show-basic-secret ...`.
|
||||
3. The Issuer URL is `https://kanidm.example.com/oauth2/openid/<client_id>`.
|
||||
4. KanIDM may emit `name` instead of `preferred_username` for the username claim. If usernames look wrong after the first login, set **Username Claim** to `name`.
|
||||
|
||||
### Pocket ID
|
||||
|
||||
1. Create a new OIDC client in Pocket ID
|
||||
2. Set the callback URL to `https://sencho.example.com/api/auth/sso/oidc/oidc_custom/callback`
|
||||
3. The Issuer URL is your Pocket ID instance URL
|
||||
4. Copy the Client ID and Secret from the client configuration
|
||||
1. Create a new OIDC client in Pocket ID.
|
||||
2. Set the callback URL to `https://sencho.example.com/api/auth/sso/oidc/oidc_custom/callback`.
|
||||
3. The Issuer URL is your Pocket ID instance URL.
|
||||
4. Copy the Client ID and Client Secret from the client view (Pocket ID shows the secret only at creation; reset it from the same screen if you lose it).
|
||||
|
||||
## Security
|
||||
|
||||
- **PKCE** - All OIDC flows use `code_challenge_method=S256` to prevent authorization code interception
|
||||
- **State parameter** - A cryptographic random value protects against CSRF attacks on the OAuth callback
|
||||
- **Encrypted secrets** - LDAP bind passwords and OIDC client secrets are encrypted at rest
|
||||
- **No local password** - SSO users are created with an unusable password hash. They cannot bypass SSO by using the password login form
|
||||
- **Admin-only configuration** - Only administrators can enable or configure SSO providers
|
||||
- **PKCE**: All OIDC flows use `code_challenge_method=S256` to prevent authorization code interception.
|
||||
- **State parameter**: A cryptographic random value protects against CSRF on the OAuth callback.
|
||||
- **Encrypted secrets**: LDAP bind passwords and OIDC client secrets are encrypted at rest in the Sencho database.
|
||||
- **No local password**: SSO users are created with an unusable password hash and cannot bypass SSO by using the password sign-in form.
|
||||
- **Admin-only configuration**: Only administrators can enable or configure SSO providers.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### Discovery URL errors
|
||||
<AccordionGroup>
|
||||
<Accordion title="Test Connection returns 'Discovery failed' or a network timeout">
|
||||
The Sencho container could not reach the provider's discovery URL. Verify the **Issuer URL** is reachable from inside the container (not just from your browser), confirm it does not include `/.well-known/openid-configuration` (just the base issuer URL), and check that container DNS can resolve the hostname. For providers with HTTPS, make sure the certificate chain is valid; self-signed certs may need additional container configuration.
|
||||
</Accordion>
|
||||
|
||||
If the **Test Connection** button returns an error like "Discovery failed" or a network timeout:
|
||||
<Accordion title="Sign-in fails with an issuer validation error">
|
||||
The `issuer` value in the provider's discovery document does not match what Sencho expects. This commonly happens when the **Issuer URL** has a trailing-slash mismatch (e.g., `https://auth.example.com` vs `https://auth.example.com/`), or when the provider is accessed via a different hostname than it advertises in its discovery document. Fix: set the **Issuer URL** to exactly match the `issuer` field returned by your provider's `/.well-known/openid-configuration` endpoint.
|
||||
</Accordion>
|
||||
|
||||
- Verify the Issuer URL is reachable from the Sencho container (not just your browser)
|
||||
- Confirm the URL does not include `/.well-known/openid-configuration`, just the base issuer URL
|
||||
- For providers behind a corporate firewall, ensure the Sencho container's DNS can resolve the hostname
|
||||
- Check that HTTPS certificates are valid. Self-signed certificates may require additional container configuration
|
||||
<Accordion title="Users land with the wrong username or no email">
|
||||
Enable **Developer Mode** (Settings → Developer) to log the raw claims Sencho receives from the provider in the server logs. Check your provider's documentation for which claims it includes in the ID token and `userinfo` response, and verify that the configured **Scopes** include everything your provider needs to emit `email` and group claims. For Custom OIDC, override **User ID Claim**, **Username Claim**, or **Email Claim** to match the names your provider actually emits.
|
||||
</Accordion>
|
||||
|
||||
### Issuer mismatch
|
||||
<Accordion title="The provider returns 'invalid redirect URI' during sign-in">
|
||||
The callback URL registered with your identity provider must exactly match `https://sencho.example.com/api/auth/sso/oidc/<provider>/callback`, where `<provider>` is `oidc_google`, `oidc_github`, `oidc_okta`, or `oidc_custom`. If Sencho sits behind a reverse proxy, set `SSO_CALLBACK_URL` to your external URL. Some providers are strict about trailing slashes and HTTP vs HTTPS.
|
||||
</Accordion>
|
||||
|
||||
If login fails with an issuer validation error, the `issuer` value in the provider's discovery document does not match what Sencho expects. This commonly happens when:
|
||||
<Accordion title="SSO buttons do not appear on the login page">
|
||||
Verify the provider is **enabled** (toggle on the Active state) in **Settings → SSO** and that the configuration saved successfully. The login page fetches the list of enabled providers when it loads; hard-refresh the tab if changes were just made.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
- The Issuer URL has a trailing slash mismatch (e.g., `https://auth.example.com` vs `https://auth.example.com/`)
|
||||
- The provider is accessed via a different hostname than it advertises in its discovery document
|
||||
|
||||
Fix: set the Issuer URL to exactly match the `issuer` field returned by your provider's `/.well-known/openid-configuration` endpoint.
|
||||
|
||||
### Claim mapping not working
|
||||
|
||||
If users are created with incorrect usernames or missing emails:
|
||||
|
||||
- Enable **Developer Mode** (Settings > Developer Mode toggle) to see the raw claims Sencho receives from the provider in the server logs
|
||||
- Check your provider's documentation for which claims it includes in the ID token and userinfo response
|
||||
- Verify that the scopes you configured include the necessary permissions (some providers require explicit `profile` or `email` scopes)
|
||||
- Set the appropriate claim names in the Custom OIDC claim mapping fields
|
||||
|
||||
### Redirect URI mismatch
|
||||
|
||||
If the provider returns an "invalid redirect URI" error during login:
|
||||
|
||||
- The callback URL configured in your identity provider must exactly match: `https://sencho.example.com/api/auth/sso/oidc/oidc_custom/callback`
|
||||
- If Sencho is behind a reverse proxy, set `SSO_CALLBACK_URL` to your external URL
|
||||
- Some providers are strict about trailing slashes and HTTP vs HTTPS
|
||||
|
||||
### SSO buttons not appearing on login page
|
||||
|
||||
- Verify the provider is **enabled** (toggle on) in Settings > SSO
|
||||
- Check that the provider configuration was saved successfully
|
||||
- The login page fetches enabled providers on load. Hard refresh the page if changes were just made.
|
||||
|
||||
For common SSO issues (LDAP connection errors, OAuth callback mismatches), see the [Troubleshooting](/operations/troubleshooting#ldap-connection-refused) page.
|
||||
The [operations troubleshooting page](/operations/troubleshooting#ldap-connection-refused) covers a few more cases that come up during initial setup, including LDAP connection refused, TLS certificate errors, and OAuth callback URL mismatches.
|
||||
|
||||
## Combining SSO with two-factor authentication
|
||||
|
||||
SSO and [two-factor authentication](/features/two-factor-authentication) can coexist. By default, users who have 2FA enrolled skip the TOTP challenge when they sign in through SSO, since the identity provider has already authenticated them. Users who want stricter sign-in can opt in to requiring 2FA on SSO from their Account & Security screen.
|
||||
SSO and [two-factor authentication](/features/two-factor-authentication) work together. By default, SSO sign-ins skip the TOTP challenge, since the identity provider has already authenticated the user. Operators who want a stricter posture can flip **Require 2FA on SSO sign-in** on their **Settings → Account & Security** card to require both factors on every SSO sign-in. The toggle only appears once at least one SSO provider is enabled.
|
||||
|
||||
@@ -1,100 +1,146 @@
|
||||
---
|
||||
title: Stack Labels
|
||||
description: Organize your stacks with colored labels for quick filtering, grouping, and bulk actions.
|
||||
title: "Stack Labels"
|
||||
description: "Per-node tags that group your stacks by purpose, surface them under collapsible headers in the sidebar, and unlock cross-stack bulk actions across the fleet."
|
||||
---
|
||||
|
||||
Stack Labels let you tag your Docker Compose stacks with custom colored labels like `production`, `staging`, or `media-server`. Once labeled, you can filter your stack list, identify stacks at a glance, and run bulk actions on entire groups.
|
||||
A **Stack Label** is a per-node tag (name plus color) you can stick on any stack. Once a stack carries a label, the sidebar groups it under that label's header instead of dumping every stack into a flat list, and Fleet View can filter the overview by tag. Operators on a Skipper or Admiral license also get a pair of fleet-wide actions powered by labels: stop every stack labeled `prod` across every node, or replace the label set on a batch of stacks in one shot.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/stack-labels/sidebar-with-labels.png" alt="Sidebar showing stacks with colored label dots and a label filter pill" />
|
||||
<img src="/images/stack-labels/sidebar-grouping.png" alt="Sidebar showing stacks grouped under three uppercase label headers (MEDIA 6, UTILITIES 5, NETWORK 3) and an UNLABELED 1 group at the bottom. Each stack row carries a small colored dot on the trailing edge that matches the group's color." />
|
||||
</Frame>
|
||||
|
||||
## Creating labels
|
||||
## What problem this solves
|
||||
|
||||
### From Settings
|
||||
A flat sidebar of fifteen stacks all rendering at the same level forces you to scan every name to find the one you want. With Stack Labels:
|
||||
|
||||
1. Open **Settings > Labels**
|
||||
2. Click **+ New Label**
|
||||
3. Enter a name and pick a color from the 10 available swatches
|
||||
4. Click **Create**
|
||||
- **Stacks group by purpose, not by alphabet.** Each label becomes a collapsible section header, sorted by stack count, so the busy buckets (your media stack, your network stack) sit at the top and rarely-touched ones can be folded away.
|
||||
- **A glance is enough.** Each row also carries up to three colored dots on its trailing edge, so a stack tagged with two purposes (for example `prod` and `media`) shows both colors without you having to open the assignment menu.
|
||||
- **Bulk operations stop being copy-paste.** Stop every `prod` stack across the fleet from one card. Re-tag eight stacks at once when a service moves between concerns. No scripting, no per-stack menu hunt.
|
||||
|
||||
## Anatomy of a label
|
||||
|
||||
| Field | Rule |
|
||||
|---|---|
|
||||
| **Name** | 1 to 30 characters. Letters, digits, spaces, and hyphens only (`^[a-zA-Z0-9 -]+$`). Case-sensitive and unique per node. |
|
||||
| **Color** | One of ten swatches: teal, blue, purple, rose, amber, green, orange, pink, cyan, slate. The color drives the dot on each row, the bullet on the group header, and the swatch in the Fleet View **Tags** filter. |
|
||||
| **Scope** | Per-node. The same name can exist on two different nodes with two different colors; the fleet-stop card matches by name across nodes. |
|
||||
| **Limit** | 50 labels per node. The **New label** primary button switches to **Limit reached** when you hit the cap. |
|
||||
|
||||
## Where labels appear
|
||||
|
||||
### Sidebar grouping
|
||||
|
||||
The stack list is split into one collapsible section per label. Group headers render in uppercase mono with a count chip on the right (`MEDIA 6`). Order is fixed: a `★ PINNED` group first if any stacks are pinned, then label groups sorted by stack count descending and by label name ascending, then `UNLABELED` last for stacks that carry no label. A stack tagged with two labels appears in both groups, the same row twice. Search and the **All / Up / Down / Updates** filter chips above the list operate on rows inside whichever groups are expanded.
|
||||
|
||||
Every row also carries up to three colored trailing dots that mirror the assigned labels. Beyond three, an additional `+N` counter appears next to the dots so the row never grows unbounded.
|
||||
|
||||
### Fleet View tags filter
|
||||
|
||||
The [Fleet View](/features/fleet-view) overview toolbar carries a **Filters** popover with a **Tags** multi-select. The dropdown lists every label that exists on any node in the fleet, with each entry rendered as a colored dot plus the label name. Selecting one or more tags filters the node cards to nodes that contain at least one stack with that label.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/stack-labels/settings-labels.png" alt="Settings Labels section showing existing labels with assignment counts" />
|
||||
<img src="/images/stack-labels/fleet-tags-filter.png" alt="Fleet Overview Filters popover open. The popover shows four sections (Status, Type, Severity, Tags). The Tags multi-select is expanded into a dropdown with five options: Media, Network, Prod, staging, Utilities, each prefixed by a small colored dot." />
|
||||
</Frame>
|
||||
|
||||
The label list shows each label's color, name, and how many stacks are currently assigned to it. Hover over a label row to reveal **Edit** and **Delete** buttons.
|
||||
The Tags filter aggregates label rows across nodes by name, so a label called `prod` that only exists on one of four nodes still shows up in the dropdown but the filter resolves to that single node.
|
||||
|
||||
## Working with labels
|
||||
|
||||
### Manage labels from Settings
|
||||
|
||||
**Settings · Advanced · Labels** is the canonical place to create, rename, recolor, and delete labels. The masthead shows a `LABELS N/50` counter so you can see how close the active node is to the cap, and a per-row stack count tells you how many stacks currently carry each label.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/stack-labels/create-label-dialog.png" alt="Create Label dialog with name input and color picker" />
|
||||
<img src="/images/stack-labels/settings-labels.png" alt="Settings page open on the Advanced · Labels section. The right pane shows a 'Per-node labels for stacks and containers.' description, a 'New label' primary button, and three label rows: Media (purple dot, 6 stacks), Network (orange dot, 3 stacks), Utilities (slate dot, 5 stacks). The masthead shows the LABELS 3/50 stat." />
|
||||
</Frame>
|
||||
|
||||
### From the stack context menu
|
||||
|
||||
You can also create labels without leaving the sidebar:
|
||||
|
||||
1. Right-click any stack in the sidebar
|
||||
2. Hover over **Labels** to open the submenu
|
||||
3. Click **Manage labels...** at the bottom to open the Labels section in Settings
|
||||
|
||||
### From the three-dot menu
|
||||
|
||||
1. Click the three-dot menu on any stack row in the sidebar
|
||||
2. Click **Labels** to open the label assignment popover
|
||||
3. Click **Create new label** at the bottom of the popover to create and assign in one step
|
||||
|
||||
## Assigning labels to stacks
|
||||
|
||||
There are two ways to assign labels:
|
||||
|
||||
**Right-click context menu:** Right-click a stack in the sidebar, hover over **Labels**, and click a label to toggle it on or off. A checkmark indicates the label is currently assigned.
|
||||
Hover any row to reveal a **Pencil** edit icon and a destructive **Trash** icon on the trailing edge. The edit dialog shares its chrome with the create dialog: the kicker reads `LABELS · NEW` for a new label or `LABELS · EDIT` when you opened it from the pencil, the body has a single `Label name` input plus the ten color swatches, and the footer has **Cancel** and **Create** (or **Save**) buttons.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/stack-labels/context-menu-labels.png" alt="Right-click context menu with Labels submenu showing assigned label with checkmark" />
|
||||
<img src="/images/stack-labels/create-label-dialog.png" alt="Create label modal. Kicker reads 'LABELS · NEW', title 'Create label', subtitle 'Manage label properties'. The body has a 'Label name' text input above a Color section with ten circular swatches arranged in a wrap (nine in the first row, slate alone in the second). Footer buttons are Cancel and a disabled Create." />
|
||||
</Frame>
|
||||
|
||||
**Three-dot menu popover:** Click the three-dot menu on a stack row, then click **Labels**. The popover shows all labels with checkmarks for assigned ones. Click any label to toggle it.
|
||||
Deleting a label opens a destructive confirmation with the kicker `LABELS · DELETE · IRREVERSIBLE` and the body line `Removes the label from every stack across the fleet.` There is no undo: the label row is dropped, every assignment row pointing to it is dropped, and the affected stacks fall back to whatever other labels they still carry. Stacks left with no remaining labels move into the `UNLABELED` group on the next sidebar refresh.
|
||||
|
||||
A stack can have multiple labels. Changes save immediately.
|
||||
### Create and assign inline from the stack menu
|
||||
|
||||
## Filtering by label
|
||||
Right-clicking a stack in the sidebar (or using the three-dot kebab menu on its row) opens the same context menu under the **organize** group. Click **Labels** to open a submenu listing every label that exists on the active node, with a checkmark next to each one currently assigned to this stack. Clicking a label toggles the assignment immediately. The two trailing items handle creation and full management:
|
||||
|
||||
### Sidebar
|
||||
|
||||
When at least one label exists, a pill bar appears between the search box and the stack list. Click a label pill to filter: only stacks with that label are shown. Click multiple pills to see stacks matching **any** of the selected labels. Click an active pill again to deselect it.
|
||||
- **New label** drops an inline form into the same submenu (text input with placeholder `Label name`, the ten color swatches, **Create** / **Cancel** buttons). Submitting creates the label on this node and assigns it to the stack in a single round trip. The entry hides itself once the node hits 50 labels.
|
||||
- **Manage labels...** sends you to **Settings · Advanced · Labels** for bulk renames, recolors, and deletions.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/stack-labels/sidebar-filtered.png" alt="Sidebar filtered by the prod label showing only matching stacks" />
|
||||
<img src="/images/stack-labels/context-menu-labels.png" alt="Right-click context menu on a sidebar stack row. The Labels submenu is open and shows three label rows (Media with a checkmark on the right, Network, Utilities), a separator, a + New label entry, and a Manage labels... entry." />
|
||||
</Frame>
|
||||
|
||||
### Fleet View
|
||||
<Frame>
|
||||
<img src="/images/stack-labels/inline-create-form.png" alt="The Labels submenu after the user clicked New label. The submenu now shows a text input with the placeholder 'Label name', a row of ten colored circles (teal selected by default), and a row of two buttons labelled Create (disabled while the input is empty) and Cancel." />
|
||||
</Frame>
|
||||
|
||||
The **Tags** filter in the [Fleet View](/features/fleet-view) toolbar lets you filter nodes by stack labels. Selecting a label filters to nodes that contain stacks with that label.
|
||||
A stack can carry multiple labels and will then appear under each label's group in the sidebar. There is no per-stack label cap; the only cap is the per-node total of 50.
|
||||
|
||||
## Bulk actions
|
||||
## Fleet · Fleet Actions
|
||||
|
||||
<Note>
|
||||
Bulk actions on a label require a Sencho **Skipper** or **Admiral** license. Creating, assigning, and filtering by labels is available on every tier.
|
||||
Bulk actions on a label require a Sencho **Skipper** or **Admiral** license. Creating, assigning, and filtering by labels is available on every tier; only the cross-node bulk surface is paid.
|
||||
</Note>
|
||||
|
||||
Right-click a label pill in the sidebar to access bulk actions:
|
||||
|
||||
| Action | Effect |
|
||||
|--------|--------|
|
||||
| **Deploy all** | Runs deploy on every stack with that label |
|
||||
| **Stop all** | Stops every stack with that label |
|
||||
| **Restart all** | Restarts every stack with that label |
|
||||
Two cards in the **Fleet · Fleet Actions** tab use labels to drive cross-node operations. Both endpoints are gated on a paid tier and an admin role: Community installs see the tab in the navigation but the body renders a single empty-state card explaining the gate, and a non-admin operator on a paid tier sees the cards rendered but receives a 403 from the backend on submit.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/stack-labels/bulk-actions-menu.png" alt="Right-click context menu on a label pill showing Deploy all, Stop all, and Restart all options" />
|
||||
<img src="/images/stack-labels/fleet-actions.png" alt="Fleet Actions tab with two cards side by side. Left card 'Stop fleet by label' (rose accent rail) has a Label name combobox containing 'Media' and a Stop matching stacks button beneath. Right card 'Bulk label assign' (purple accent rail) has a node selector reading Local (local), a stacks checklist showing plex and radarr ticked, a Labels row with a highlighted Media pill plus inactive Network and Utilities, and an Apply to 3 stacks button." />
|
||||
</Frame>
|
||||
|
||||
A confirmation dialog shows which stacks will be affected before executing. If any stacks fail during a bulk action, the error toast lists the specific stack names that failed. Only one bulk action can run at a time per node.
|
||||
### Stop fleet by label
|
||||
|
||||
## Managing labels
|
||||
Type a label name; Sencho fans the request out to every online node and stops every stack on that node that carries a label with the same name. The card autocompletes the input from the union of label names on every reachable node, so you do not need to remember whose label rows exist where. The result list shows a per-node breakdown with success and failure counts, and the warning callout under the input restates the fleet-wide semantics: `Different nodes can have their own label rows. Stops are dispatched per node and report per-stack results below.` A confirmation modal titled `Stop all stacks labeled "<name>"?` with the **Stop fleet** primary action runs the action.
|
||||
|
||||
Open **Settings > Labels** to view, edit, and delete your labels.
|
||||
Offline nodes are skipped during autocomplete loading and are reported as failures during the actual run, so a partial-fleet stop is observable rather than silent.
|
||||
|
||||
- **Edit**: Click the pencil icon on a label row to change its name or color
|
||||
- **Delete**: Click the trash icon to remove the label. A confirmation dialog explains that the label will be removed from all stacks. This cannot be undone.
|
||||
### Bulk label assign
|
||||
|
||||
Label names must be unique per node. You can create up to 50 labels per node. There are 10 color options: teal, blue, purple, rose, amber, green, orange, pink, cyan, and slate.
|
||||
Pick a node from the dropdown; the card loads that node's stacks and labels in parallel. Tick the stacks you want to relabel and click the label pills you want to apply. The footer button reads `Apply to N stack(s)` and the confirmation modal title reads `Apply N label(s) to M stack(s)?` so the scope is unambiguous before you commit. A line under the controls makes the replacement semantics explicit:
|
||||
|
||||
> Selected labels replace each chosen stack's existing label set on this node. Selecting no labels clears assignments.
|
||||
|
||||
The **Bulk label assign** card is per-node only by design. To re-tag stacks on a different node, switch the picker; the stack and label list refreshes and your previous selection clears.
|
||||
|
||||
## Limits and rules
|
||||
|
||||
- **50 labels per node.** Settings hides the **New label** button at the cap; the inline `+ New label` entry in the stack menu hides itself too.
|
||||
- **Names are unique per node**, case-sensitive. The same name on two nodes is two separate label rows. Cross-node fleet stop matches on name; bulk assign always operates on one node's labels at a time.
|
||||
- **Allowed name characters**: letters, digits, spaces, and hyphens. Empty names and names beyond 30 characters are rejected at the API.
|
||||
- **Bulk-action concurrency**: only one label-driven bulk action can run on a single node at a time. A second concurrent attempt against the same node returns HTTP 429 and the operator sees an error toast; the in-flight action keeps running.
|
||||
- **Tier visibility**: organizing with labels is Community. Sidebar grouping, trailing dots on stack rows, the **Settings · Advanced · Labels** panel, the inline create form in the stack menu, and the Fleet View **Tags** filter all light up on every tier. Only the **Fleet · Fleet Actions** body content (the Stop-fleet-by-label and Bulk-label-assign cards) is paid; Community installs see the tab in the navigation but the body renders a single upgrade-context empty state.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Stacks are not grouped under labels in the sidebar">
|
||||
Grouping requires at least one stack to carry at least one label. With zero labels assigned the sidebar collapses into a single `UNLABELED` group, which is rendered as a flat list. Create a label from **Settings · Advanced · Labels** or right-click any stack and use **Labels · New label**, assign it to a stack, and the grouped layout takes over.
|
||||
</Accordion>
|
||||
<Accordion title="The trailing colored dots are missing on stack rows">
|
||||
A stack row only renders trailing dots when at least one label is assigned to that stack. Right-click the row, open the **Labels** submenu, and tick at least one label; the dots appear on the next sidebar refresh. If a stack already has labels assigned but the dots still do not appear, check that the active node is the one that owns the assignments. Labels are per-node, so switching the node switcher to a different instance shows that instance's assignments only.
|
||||
</Accordion>
|
||||
<Accordion title="`+ New label` is missing from the stack context menu">
|
||||
The active node already has 50 labels (the per-node cap). Both the inline `New label` entry in the stack submenu and the **New label** button in **Settings · Advanced · Labels** hide themselves at the cap. Delete an unused label or rename an existing one to free a slot.
|
||||
</Accordion>
|
||||
<Accordion title="The Tags filter does not list a label I just created">
|
||||
The Tags filter aggregates labels across every node in the fleet by name. If the new label only exists on one node and that node was offline at the moment the page loaded, the dropdown may not include it. Refresh **Fleet · Overview** with the **Refresh** button in the toolbar to repull node state.
|
||||
</Accordion>
|
||||
<Accordion title="`Stop fleet by label` reports `No nodes have a label by that name`">
|
||||
Labels are per-node, so the fleet-stop matches by name across nodes. If the label you typed only exists on the active node and you typed the wrong case (`prod` versus `Prod`), no node will match. The combobox autocompletes from the union of label names on reachable nodes; pick from the suggestion list rather than typing freehand to avoid case mistakes.
|
||||
</Accordion>
|
||||
<Accordion title="`Bulk label assign` cleared every label on my stacks unexpectedly">
|
||||
The card replaces, it does not merge. Selecting no labels and clicking **Apply to N stack(s)** is the documented way to clear assignments, and the confirmation copy on the **Bulk label assign** modal restates this: `No labels selected, this will clear existing assignments on the selected stacks.` Re-pick the labels you want and run the action again to restore them.
|
||||
</Accordion>
|
||||
<Accordion title="Two stacks with the same name on different nodes only got relabeled on one">
|
||||
`Bulk label assign` is per-node by design. The node picker at the top is the source of truth and the stacks list only shows stacks on that node. Run the card a second time with the other node selected, or use **Stop fleet by label** instead if the goal is fleet-wide.
|
||||
</Accordion>
|
||||
<Accordion title="The Fleet Actions tab body shows an upgrade card on Community">
|
||||
Cross-node bulk actions stay on Skipper and Admiral. The tab itself is visible to every tier, but the body renders a single empty-state card on Community pointing at the upgrade. All other Stack Labels surfaces (sidebar grouping, trailing dots, the Settings panel, the inline create form, the Fleet View Tags filter) work the same on Community.
|
||||
</Accordion>
|
||||
<Accordion title="Deleting a label removed it from every stack">
|
||||
Working as designed. The destructive confirmation reads `Removes the label from every stack across the fleet.` The label row and every assignment row that pointed to it are dropped in a single transaction. There is no undo; recreate the label by name and color and reassign the affected stacks if you need to recover.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
@@ -1,148 +1,260 @@
|
||||
---
|
||||
title: Two-Factor Authentication
|
||||
description: Protect your Sencho account with a time-based one-time password (TOTP) and single-use backup codes.
|
||||
description: Protect your Sencho account with a time-based one-time password (TOTP), single-use recovery codes, and optional SSO enforcement.
|
||||
---
|
||||
|
||||
Two-factor authentication (2FA) adds a second step to sign-in. After your password is accepted, Sencho asks for a six-digit code from your authenticator app. Someone who steals your password still cannot sign in without that code.
|
||||
<Note>
|
||||
Two-factor authentication is available on every Sencho tier including Community. It is enrolled per operator account, so each user (admin, viewer, deployer, node admin, auditor) can turn it on independently.
|
||||
</Note>
|
||||
|
||||
2FA is available on all Sencho tiers, including Community.
|
||||
<Note>
|
||||
The Account panel is hub-only. It does not appear under Settings on a remote node, since each operator's password and 2FA live on the gateway that proxies the fleet. Switch back to the local node via the node switcher in the masthead to enrol or disable.
|
||||
</Note>
|
||||
|
||||
Two-factor authentication adds a second step to sign-in. After your password is accepted, Sencho asks for a six-digit code from your authenticator app. Someone who steals your password still cannot sign in without that code. Sencho also issues ten single-use recovery codes during enrolment, so a lost phone is not a lockout.
|
||||
|
||||
## How it works
|
||||
|
||||
1. You enrol once by scanning a QR code with an authenticator app (or typing a secret by hand)
|
||||
2. The app generates a fresh six-digit code every 30 seconds
|
||||
3. On every sign-in, Sencho asks for the current code after your password passes
|
||||
4. You also save ten single-use backup codes for the days when your phone is not available
|
||||
1. You enrol once by scanning a QR code with an authenticator app (or pasting the secret by hand).
|
||||
2. The app generates a fresh six-digit code every 30 seconds.
|
||||
3. On every sign-in, Sencho asks for the current code after your password passes. The form auto-submits the moment you enter the sixth digit.
|
||||
4. You also save ten single-use recovery codes for the days when your phone is not available. Each code is consumed the first time it is used.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/two-factor-auth/challenge.png" alt="Sign-in challenge screen asking for the six-digit code" />
|
||||
<img src="/images/two-factor-auth/challenge.png" alt="Sign-in challenge screen with the kicker SENCHO VERIFY, an italic display Verify hero, the caption Enter the 6-digit code from your authenticator, six empty digit cells, a Use backup code link in mono brackets below, and a Console Verify / Cancel Sign out footer." />
|
||||
</Frame>
|
||||
|
||||
## Enrol
|
||||
|
||||
Open **Settings → Account & Security** and click **Set up 2FA**. Sencho walks you through three steps.
|
||||
|
||||
### Step 1 – Scan the QR code
|
||||
|
||||
Scan the QR code with an authenticator app such as **1Password**, **Bitwarden**, **Google Authenticator**, **Authy**, or **Microsoft Authenticator**.
|
||||
Open **Settings · Account** and locate the **Two-factor authentication** section under the Password section. When 2FA is off, a warn callout reads `Two-factor is off` with the subtitle `Add a time-based code from your authenticator app, every sign-in.` and a `Set up 2FA` button on the right.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/two-factor-auth/enroll-qr.png" alt="Enrolment dialog showing the QR code" />
|
||||
<img src="/images/two-factor-auth/account-card-disabled.png" alt="Settings Account panel, Two-factor authentication section in the off state. The kicker on the right reads OFF, and the warn callout shows a shield icon, the title Two-factor is off, the subtitle Add a time-based code from your authenticator app, every sign-in, and a Set up 2FA button on the far right." />
|
||||
</Frame>
|
||||
|
||||
If your app cannot scan or you would rather type the secret into a password manager, click **Can't scan? Show secret key** and copy the base32 string. Paste it into the app's manual-entry field.
|
||||
Clicking `Set up 2FA` opens a three-step modal with the kicker `SECURITY · MFA`. A step rail at the top shows `01 PAIR` → `02 CONFIRM` → `03 ARCHIVE` and highlights the active step.
|
||||
|
||||
### Step 2 – Confirm
|
||||
### Step 1 · Pair your authenticator
|
||||
|
||||
Enter the six-digit code your app is currently showing. This proves the app is paired correctly before Sencho turns 2FA on.
|
||||
The first step shows a QR code and the message `Scan the code with 1Password, Bitwarden, Google Authenticator, or any TOTP app.` Any RFC 6238 TOTP authenticator works; the four names listed are common ones, not an exhaustive list.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/two-factor-auth/enroll-confirm.png" alt="Enrolment confirmation step asking for the six-digit code" />
|
||||
<img src="/images/two-factor-auth/enroll-qr.png" alt="Enrol step 1 modal titled Pair your authenticator. The modal shows the SECURITY MFA kicker, the active 01 PAIR step in the rail, an instruction line about scanning with 1Password Bitwarden Google Authenticator or any TOTP app, a square QR code, a SECRET MANUAL ENTRY label below it with the base32 secret in a monospace field plus a copy icon button, and Cancel and Continue buttons at the bottom." />
|
||||
</Frame>
|
||||
|
||||
### Step 3 – Save your backup codes
|
||||
If your app cannot scan or you would rather paste the secret into a password manager, the base32 string sits directly below the QR under the `SECRET · MANUAL ENTRY` label. The copy icon next to it puts the raw secret on your clipboard.
|
||||
|
||||
Sencho shows ten single-use backup codes. Each code can sign you in once without the authenticator app. **Save them somewhere safe before closing this dialog**, they are not shown again.
|
||||
Click **Continue** to advance to the confirmation step.
|
||||
|
||||
### Step 2 · Confirm the pairing
|
||||
|
||||
The second step asks for the current six-digit code so Sencho can verify the app is paired correctly before turning 2FA on. The instruction line reads `Enter the 6-digit code shown in your authenticator to confirm the pairing.`
|
||||
|
||||
<Frame>
|
||||
<img src="/images/two-factor-auth/enroll-backup-codes.png" alt="Ten backup codes shown after enrolment" />
|
||||
<img src="/images/two-factor-auth/enroll-confirm.png" alt="Enrol step 2 modal titled Confirm the pairing. The step rail shows step 1 PAIR as a small filled dot (complete), step 02 CONFIRM active with a brand underline, and step 03 ARCHIVE muted. The body shows the instruction line and six empty digit cells (one focused on the left). A Back button sits at the bottom right and there is no submit button." />
|
||||
</Frame>
|
||||
|
||||
Copy them to a password manager or download them as a text file. Common places users store them:
|
||||
Sencho submits the moment you enter the sixth digit. There is no submit button on this step; if the code is wrong, the cells turn briefly red and clear themselves so you can try again with the next 30-second code. The `Back` button at the bottom right returns to the QR step if you need to re-pair.
|
||||
|
||||
- Their primary password manager as a secure note on the Sencho entry
|
||||
- An encrypted document on a separate device
|
||||
- Printed and stored with other emergency recovery material
|
||||
### Step 3 · Save your recovery codes
|
||||
|
||||
After a successful confirmation Sencho issues ten single-use recovery codes, shown grouped as `ABCDE-FGHIJ` for easier reading. **Save them somewhere safe before closing this dialog.** They are not shown again.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/two-factor-auth/enroll-backup-codes.png" alt="Enrol step 3 modal titled Save your recovery codes. The step rail shows 1 PAIR and 2 CONFIRM as complete dots and 03 ARCHIVE active with a brand underline. The body has the instruction Each code unlocks your account once if your authenticator is unavailable, a RECOVERY CODES card with the 10 ISSUED counter on the right, two columns of five codes each in the ABCDE-FGHIJ format with row numbers 01 to 10, a Copy all button and a Download button side by side, and a Done button at the bottom right." />
|
||||
</Frame>
|
||||
|
||||
The dialog offers two actions:
|
||||
|
||||
- **Copy all** puts the ten codes (newline-separated) on the clipboard.
|
||||
- **Download** saves them as `sencho-backup-codes.txt`, a plain-text file you can move to a password manager or print.
|
||||
|
||||
Common places operators store recovery codes:
|
||||
|
||||
- A secure note attached to the Sencho entry in their primary password manager.
|
||||
- An encrypted document on a separate device.
|
||||
- Printed and kept with other emergency recovery material.
|
||||
|
||||
Clicking **Done** finishes enrolment. The Account panel now shows the **enabled** state.
|
||||
|
||||
## Sign in with 2FA
|
||||
|
||||
After entering your password, Sencho shows the 2FA challenge screen.
|
||||
|
||||
1. Open your authenticator app, find the Sencho entry, read the six-digit code
|
||||
2. Type it into the **Verification code** field
|
||||
3. Sencho submits automatically once you enter the sixth digit. No click required.
|
||||
|
||||
<Note>
|
||||
If you prefer the explicit route, the **Verify and sign in** button still works. The auto-submit only applies to the six-digit TOTP field; backup codes always require a click to confirm.
|
||||
</Note>
|
||||
|
||||
If your phone is unavailable, click **Use a backup code instead**, enter one of the codes you saved during enrolment, and click **Verify and sign in**. That code is now used up.
|
||||
|
||||
### Backup code entry tips
|
||||
|
||||
Backup codes are shown grouped as `ABCDE-FGHIJ` to make them easier to read. When signing in, you can enter them any of these ways:
|
||||
|
||||
- Paste the code exactly as shown: `ABCDE-FGHIJ`
|
||||
- Paste without the dash: `ABCDEFGHIJ`
|
||||
- Paste with extra whitespace or lowercase letters; Sencho normalises the input before sending it to the server
|
||||
|
||||
Only letters and digits are significant; dashes, spaces, and case are ignored.
|
||||
|
||||
## Regenerate backup codes
|
||||
|
||||
If you think your backup codes have been exposed, or you have used most of them, regenerate them from **Settings → Account & Security → Regenerate backup codes**. Sencho asks for a current code, then issues a fresh set of ten and invalidates the old set immediately.
|
||||
|
||||
The Account & Security card nudges you about low code counts so you notice before you are locked out:
|
||||
|
||||
- **3 or more codes remaining:** a muted count under the status message.
|
||||
- **1 or 2 codes remaining:** the count turns warning-coloured with an alert icon, inviting you to regenerate a fresh set.
|
||||
- **0 codes remaining:** a dedicated warning card appears with a **Regenerate now** button. At this point, losing your authenticator app means recovery needs an administrator, so regenerate before that happens.
|
||||
After your password is accepted, Sencho replaces the password form with the verify screen.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/two-factor-auth/account-card-enabled.png" alt="Account card showing 2FA enabled with regenerate and disable actions" />
|
||||
<img src="/images/two-factor-auth/challenge.png" alt="Sign-in challenge in TOTP mode showing the SENCHO VERIFY kicker, an italic display Verify hero, the caption Enter the 6-digit code from your authenticator, six digit cells (first focused), a mono-bracketed Use backup code link, and a footer reading Console Verify on the left and Cancel Sign out on the right." />
|
||||
</Frame>
|
||||
|
||||
1. Open your authenticator, find the Sencho entry, read the six-digit code.
|
||||
2. Type it into the digit cells. Sencho submits automatically once you enter the sixth digit. No click required.
|
||||
|
||||
The mono-bracketed `[ Use backup code ]` link below the cells switches the form to backup-code mode.
|
||||
|
||||
### Backup-code mode
|
||||
|
||||
<Frame>
|
||||
<img src="/images/two-factor-auth/challenge-backup.png" alt="Sign-in challenge in backup-code mode showing the SENCHO VERIFY kicker, the Verify hero, the caption Enter one of your saved backup codes to continue, a BACKUP CODE 10 CHARS label, a single wide input pre-filled with the placeholder ABCDE-FGHIJ, a Verify button below, and a Use authenticator link to toggle back." />
|
||||
</Frame>
|
||||
|
||||
The input is one wide field labelled `BACKUP CODE · 10 CHARS`. Type or paste a code in any of these forms:
|
||||
|
||||
- Exactly as shown on the dialog: `ABCDE-FGHIJ`.
|
||||
- Without the dash: `ABCDEFGHIJ`.
|
||||
- With extra whitespace or lowercase letters; Sencho normalises before sending to the server.
|
||||
|
||||
Only letters and digits are significant; dashes, spaces, and case are ignored on both client and server. The form here does **not** auto-submit (a 10-character paste is too easy to truncate by mistake); click **Verify** to send the code. That code is now used up.
|
||||
|
||||
The mono-bracketed `[ Use authenticator ]` link toggles back to the six-digit TOTP form.
|
||||
|
||||
### Rate limiting
|
||||
|
||||
Five failed verifications in a row lock the account for 15 minutes. The challenge screen replaces the input with a `Retry in MM:SS` countdown and a `Rate limited` label while the lockout is in effect. The failure counter resets only on a successful sign-in: if you retry with another wrong code after the window expires, the counter is still at five and the lockout re-engages immediately. See [Lockout recovery](#lockout-recovery) below.
|
||||
|
||||
## Account panel anatomy
|
||||
|
||||
Once 2FA is enabled, the **Two-factor authentication** section on the Account panel renders three rows and a destructive ghost link. The kicker on the right of the section header reads `ENABLED` (it reads `OFF` before enrolment).
|
||||
|
||||
<Frame>
|
||||
<img src="/images/two-factor-auth/account-card-enabled.png" alt="Two-factor authentication section in the enabled state. The header reads TWO-FACTOR AUTHENTICATION on the left and ENABLED on the right. The Authenticator app row shows a shield-check icon next to the mono badge ENROLLED. The Backup codes row shows 10 remaining next to an outline Regenerate button. A destructive ghost Disable 2FA link sits at the bottom right." />
|
||||
</Frame>
|
||||
|
||||
| Row | What it shows |
|
||||
|------|---------------|
|
||||
| **Authenticator app** | A shield-check icon and the mono badge `ENROLLED`. Helper text: `Sign-in requires a time-based code from your authenticator. Keep your backup codes safe.` |
|
||||
| **Backup codes** | The current remaining count (`N remaining`) and an outline **Regenerate** button. Helper text varies by count (see [Recovery codes](#recovery-codes)). |
|
||||
| **Require 2FA on SSO sign-in** | A TogglePill, only rendered when at least one SSO provider is configured. Default is off. Helper text: `By default, SSO logins skip the second factor. Enforce it here to require both.` See [SSO sign-in](#sso-sign-in) below. |
|
||||
| **Disable 2FA** | A destructive ghost link at the bottom right. See [Disable 2FA](#disable-2fa). |
|
||||
|
||||
The masthead at the top of the Settings page also carries a `2FA on` chip while 2FA is enabled. When the remaining count drops to two or below, a second masthead chip appears: `BACKUP · 2 left` (warn) or `BACKUP · 0 left` (error).
|
||||
|
||||
## Recovery codes
|
||||
|
||||
Each recovery code can be used once. Regenerate when you have used most of them, or if you think the saved set has been exposed.
|
||||
|
||||
### How the row reacts to the remaining count
|
||||
|
||||
The **Backup codes** row helper text and tone change as the count drops:
|
||||
|
||||
| Codes remaining | Helper text | Tone | Extra UI |
|
||||
|-----------------|-------------|------|----------|
|
||||
| 3 or more | `Single-use codes for when your authenticator is unavailable.` | default | none |
|
||||
| 1–2 | `Running low. Regenerate a fresh set.` | warn | masthead `BACKUP · N left` chip |
|
||||
| 0 | `No backup codes remain. Regenerate a fresh set before you lose access to your authenticator.` | error | masthead `BACKUP · 0 left` chip plus the standalone callout described below |
|
||||
|
||||
<Frame>
|
||||
<img src="/images/two-factor-auth/account-card-low-codes.png" alt="Two-factor authentication section with 2 backup codes remaining. The Backup codes row helper text Running low Regenerate a fresh set is rendered in a warning amber tone. The value column shows 2 remaining next to an outline Regenerate button." />
|
||||
</Frame>
|
||||
|
||||
When the count hits zero, an additional error callout renders below the section action row:
|
||||
|
||||
<Frame>
|
||||
<img src="/images/two-factor-auth/account-card-no-codes.png" alt="Two-factor authentication section with 0 backup codes remaining. The Backup codes row shows the error-tone helper text No backup codes remain Regenerate a fresh set before you lose access to your authenticator, the value 0 remaining, and a Regenerate button. Below the section, a standalone error callout with an alert-triangle icon shows the kicker NO BACKUP CODES LEFT, the subtitle Without codes recovery needs an administrator if you lose your authenticator, and a Regenerate button on the right." />
|
||||
</Frame>
|
||||
|
||||
### Regenerate
|
||||
|
||||
Click the **Regenerate** button on the Backup codes row. The dialog opens with the kicker `SECURITY · BACKUP CODES` and the title `Confirm identity`. The body reads `Enter a code from your authenticator to generate a new set. The previous codes stop working immediately.`
|
||||
|
||||
<Frame>
|
||||
<img src="/images/two-factor-auth/regenerate-confirm.png" alt="Regenerate dialog confirm step. The modal header shows the kicker SECURITY BACKUP CODES and the italic display title Confirm identity. The body has the instruction Enter a code from your authenticator to generate a new set The previous codes stop working immediately and six empty digit cells, with the first cell focused. A Cancel ghost button sits at the bottom right and there is no submit button." />
|
||||
</Frame>
|
||||
|
||||
Sencho submits the moment you enter the sixth digit. On success the dialog advances to the `New recovery codes` step:
|
||||
|
||||
<Frame>
|
||||
<img src="/images/two-factor-auth/regenerate-show.png" alt="Regenerate dialog show step. The header reads SECURITY BACKUP CODES kicker and New recovery codes title. A warn rail at the top reads PREVIOUS CODES HAVE BEEN INVALIDATED. Below it is the line Each code can be used once Store them safely They will not be shown again and a RECOVERY CODES card with 10 ISSUED on the right showing ten new codes in two columns. Copy all and Download buttons sit side by side, with a Done button at the bottom right." />
|
||||
</Frame>
|
||||
|
||||
The previous ten codes stop working the instant the new set is issued, including any that you have not yet used. Save the new set the same way you saved the originals: copy to a password manager, download as `sencho-backup-codes.txt`, or both.
|
||||
|
||||
This dialog accepts only a TOTP code, not a backup code. Regenerating with a backup code would invalidate that very code mid-flight, so the action requires a fresh six-digit TOTP from the authenticator.
|
||||
|
||||
## SSO sign-in
|
||||
|
||||
If Sencho is configured with LDAP or an OIDC provider (Custom OIDC, Google, GitHub, Okta), the SSO flow signs you in without the second factor by default. Your identity provider has already authenticated you, and stacking a TOTP on top is friction most teams do not need.
|
||||
|
||||
You can opt your own account into stricter behaviour with the **Require 2FA on SSO sign-in** toggle. It only appears on the Account panel when at least one SSO provider is enabled.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/two-factor-auth/sso-enforce.png" alt="Require 2FA on SSO sign-in row. The row title is bold on the left, the helper text reads By default SSO logins skip the second factor Enforce it here to require both, and a TogglePill on the right is in the OFF state." />
|
||||
</Frame>
|
||||
|
||||
This toggle is a per-user opt-in. For a fleet-wide policy, admins can require MFA enrolment at the SSO provider level via the **Require MFA** toggle on the provider config (see [SSO Authentication](/features/sso) and the [RBAC page](/features/rbac#sso-auto-provisioning)). The two toggles are independent: the per-user toggle decides whether a TOTP is asked for on every SSO sign-in; the per-provider toggle decides whether new SSO-provisioned users are forced to enrol TOTP before they can use the rest of the console.
|
||||
|
||||
## Disable 2FA
|
||||
|
||||
From the same Account & Security section, click **Disable 2FA**. Sencho asks you to confirm with a current authenticator code (or a backup code) so that no one else can disable protection without your phone. Once disabled, only your password protects the account.
|
||||
|
||||
## Signing in with SSO
|
||||
|
||||
If Sencho is configured with LDAP or an OIDC provider (Google, GitHub, Okta), the SSO flow signs you in without the second factor by default. Your identity provider has already authenticated you.
|
||||
|
||||
If you want Sencho to ask for a TOTP code even after a successful SSO sign-in, flip the **Require 2FA even when signing in via SSO** switch on the Account & Security card. The switch only appears when at least one SSO provider is enabled.
|
||||
The destructive ghost **Disable 2FA** link at the bottom of the Two-factor authentication section opens a red-chromed confirmation modal:
|
||||
|
||||
<Frame>
|
||||
<img src="/images/two-factor-auth/sso-enforce.png" alt="Toggle that requires 2FA even when signing in via SSO" />
|
||||
<img src="/images/two-factor-auth/disable-confirm.png" alt="Disable 2FA confirmation modal. The destructive header shows the kicker SECURITY MFA DISABLE in red and the italic display title Turn off two-factor with a red top rail. The body reads Disabling 2FA removes this login layer Your backup codes become invalid Confirm with a current code to proceed, with six digit cells below and a mono-bracketed Use backup code toggle. The footer shows a Cancel button on the left and a red Disable button on the right." />
|
||||
</Frame>
|
||||
|
||||
| Surface | Value |
|
||||
|---------|-------|
|
||||
| Kicker | `SECURITY · MFA · DISABLE` |
|
||||
| Title | `Turn off two-factor` |
|
||||
| Body | `Disabling 2FA removes this login layer. Your backup codes become invalid. Confirm with a current code to proceed.` |
|
||||
| Toggle | `[ Use backup code ]` / `[ Use authenticator ]` |
|
||||
| Confirm button | `Disable` (red destructive variant) |
|
||||
|
||||
Confirmation needs either a current TOTP or a backup code, so nobody who has only your password can disable the second factor. Once 2FA is disabled, your backup codes are invalidated immediately and only your password protects the account.
|
||||
|
||||
## Admin reset (lost authenticator)
|
||||
|
||||
If you lose both your authenticator app and your remaining backup codes, you cannot sign in on your own. Any administrator can reset your 2FA from **Settings · Users**: the shield-off icon on your row opens a confirmation modal targeted at your account.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/two-factor-auth/admin-reset.png" alt="Admin reset confirmation modal opened from the Users panel. The header shows the kicker USERS RESET 2FA and the italic display title Reset 2FA for viewer. The body reads Removes the user's authenticator enrolment and backup codes They will sign in with just their password on their next login and can re-enrol from their account settings Use this when a user has lost access to their authenticator. The footer has Cancel and Reset 2FA buttons." />
|
||||
</Frame>
|
||||
|
||||
The modal copy is verbatim:
|
||||
|
||||
> Removes the user's authenticator enrolment and backup codes. They will sign in with just their password on their next login and can re-enrol from their account settings. Use this when a user has lost access to their authenticator.
|
||||
|
||||
The reset bumps the target user's token version, so any active sessions of theirs return `401` on the next API request. After the reset, the user signs in with their password and re-enrols TOTP from **Settings · Account**. The full procedure for administrators, including a CLI fallback if every admin has lost access, lives at [Managing Two-Factor Authentication](/operations/two-factor-admin).
|
||||
|
||||
## Lockout recovery
|
||||
|
||||
The 15-minute lockout that fires after five failed verifications is per-account, not per-IP. The counter clears on a successful sign-in. Practical consequence: if you keep retrying with wrong codes after the 15-minute window expires, the counter is still at five and the lockout re-engages on the first wrong code. To break the cycle, either wait until you have a known-good code (clock synced authenticator) and sign in cleanly, or have an administrator reset your 2FA from the Users panel to clear both the enrollment and the failure counter at once.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### "Code invalid" even though the code is current
|
||||
<AccordionGroup>
|
||||
<Accordion title="Code invalid even though the code is current">
|
||||
The most common cause is clock drift between your phone and the Sencho host. TOTP codes are computed from the current Unix time; even a 30-second skew breaks verification.
|
||||
|
||||
The most common cause is clock drift between your phone and the Sencho server. Authenticator apps compute codes from the current Unix time, and even a 30 second skew breaks verification.
|
||||
- On iOS: **Settings · General · Date & Time** → enable **Set Automatically**.
|
||||
- On Android: **Settings · System · Date & time** → enable **Automatic date & time**.
|
||||
- In Google Authenticator: open the menu → **Settings · Time correction for codes · Sync now**.
|
||||
- In Authy: **Settings · My account · Sync device time**.
|
||||
|
||||
- On iOS, open **Settings → General → Date & Time** and enable **Set Automatically**
|
||||
- On Android, open **Settings → System → Date & time** and enable **Automatic date & time**
|
||||
- In Google Authenticator, open the menu → **Settings → Time correction for codes → Sync now**
|
||||
- In Authy, open **Settings → My account → Sync device time**
|
||||
If Sencho runs in a container, the container's clock follows the host. Make sure the host clock is synced with NTP.
|
||||
</Accordion>
|
||||
<Accordion title="Wrong account selected in the authenticator">
|
||||
Authenticator apps let you store many accounts. The label Sencho uses is the literal issuer `Sencho` plus your username. If you have more than one Sencho instance enrolled in the same authenticator, every entry shows the same issuer; rename them in the app to distinguish (most apps let you edit the label without invalidating the secret).
|
||||
</Accordion>
|
||||
<Accordion title="The QR code will not scan">
|
||||
Use the **Secret · manual entry** row directly below the QR. Click the copy icon to put the base32 string on your clipboard, then paste it into your authenticator's manual-entry screen. The secret is identical; the QR is just a convenience.
|
||||
</Accordion>
|
||||
<Accordion title="Lost phone and no backup codes left">
|
||||
Ask an administrator to reset your 2FA from **Settings · Users**. The shield-off icon on your row opens the confirmation modal documented above. If you are the only admin and have also lost access, the administrator can run the CLI reset on the host running Sencho; see [Managing Two-Factor Authentication](/operations/two-factor-admin) for the command.
|
||||
</Accordion>
|
||||
<Accordion title="Lost your backup codes but still have the authenticator">
|
||||
Sign in normally with a TOTP, then **Settings · Account · Two-factor authentication · Regenerate**. Sencho asks for a current authenticator code and then issues a fresh set of ten codes. The previous set is invalidated the moment the new set is shown.
|
||||
</Accordion>
|
||||
<Accordion title="Ran out of backup codes">
|
||||
Each code is single-use, so the count drops by one every time you sign in with a backup code instead of the authenticator. As soon as you sign back in with an authenticator code, regenerate a new set from **Settings · Account · Two-factor authentication · Regenerate**. Without codes, recovery from a lost authenticator requires an administrator.
|
||||
</Accordion>
|
||||
<Accordion title="SSO sign-in is unexpectedly asking for a code">
|
||||
Two independent toggles can cause this. Check both:
|
||||
|
||||
If your server runs in a container, make sure its host clock is synced with NTP.
|
||||
|
||||
### Wrong account selected in the authenticator
|
||||
|
||||
Authenticator apps let you store many accounts. If you have more than one Sencho instance, or a Sencho entry and a different service that looks similar, pick the correct one. The issuer label Sencho uses is simply `Sencho`.
|
||||
|
||||
### The QR code will not scan
|
||||
|
||||
Click **Can't scan? Show secret key** during enrolment, copy the base32 string, and paste it into the authenticator app's manual-entry screen. The secret is identical, the QR code is just a convenience.
|
||||
|
||||
### Lost phone and no backup codes left
|
||||
|
||||
Contact an administrator. An admin can reset your 2FA from the Users section in **Settings → Users**. See the [admin guide](/operations/two-factor-admin) for the steps. If you are the only admin and have lost access, the administrator can also reset 2FA from the command line on the host running Sencho.
|
||||
|
||||
### Lost your backup codes
|
||||
|
||||
If you still have your authenticator app, sign in as normal and regenerate the codes from **Settings → Account & Security → Regenerate backup codes**. The previous set stops working immediately. If you no longer have the authenticator app either, follow the "Lost phone and no backup codes left" entry above and ask an administrator to reset 2FA.
|
||||
|
||||
### Ran out of backup codes
|
||||
|
||||
Each backup code can be used once. As soon as you sign back in with an authenticator code, regenerate a new set from **Settings → Account & Security**. Without codes, losing your phone means recovery requires an administrator.
|
||||
|
||||
### SSO sign-in is unexpectedly asking for a code
|
||||
|
||||
The **Require 2FA even when signing in via SSO** toggle is on for your account. Open **Settings → Account & Security** and flip it off if SSO alone is enough for your threat model.
|
||||
|
||||
### Enrolment did not activate after confirmation
|
||||
|
||||
The browser may have dropped the JSON response. Retry the confirmation step and watch for an error toast. If confirmation continues to fail, cancel the dialog, ensure your clock is synced, and start enrolment again.
|
||||
- **Your own toggle.** **Settings · Account · Two-factor authentication · Require 2FA on SSO sign-in** is on for your account. Flip it off if SSO alone is enough for your threat model.
|
||||
- **The provider's toggle.** Each SSO provider config has a `Require MFA` switch. When that is on, every SSO-provisioned user must enrol TOTP after their first successful sign-in. Only an administrator can change this on the provider config under **Settings · SSO**.
|
||||
</Accordion>
|
||||
<Accordion title="Repeatedly locked out after waiting 15 minutes">
|
||||
The 15-minute window expires on the clock, but the failure counter only resets on a successful sign-in. If you retry with another wrong code after the window expires, the counter is still at five and the lockout re-engages immediately. Have an administrator reset your 2FA from the Users panel to clear both the enrolment and the failure counter, then sign in with your password and re-enrol TOTP from your account settings.
|
||||
</Accordion>
|
||||
<Accordion title="The shield icon is missing on a user I expected to reset">
|
||||
The shield-off icon on a Users-panel row only renders when that user has a finished TOTP enrolment. A user who started enrolment but never completed step 2 has no enrolment to reset; ask them to finish enrolment from their account settings.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
@@ -1,12 +1,12 @@
|
||||
---
|
||||
title: "Vulnerability Scanning"
|
||||
description: "Scan container images for known CVEs, surface severity badges in the Resources Hub, and alert on policy violations."
|
||||
description: "Scan container images and stack compose files for CVEs, secrets, and misconfigurations. Surface severity badges in the Resources Hub, compare scans over time, and gate deploys on policy violations."
|
||||
---
|
||||
|
||||
Sencho integrates with [Trivy](https://trivy.dev) to scan container images for known vulnerabilities (CVEs), surface severity badges next to your images, and alert when a scan result exceeds a configured threshold. Manual scanning, secret and misconfiguration detection, scan comparison, and CVE suppressions are all available on every tier. Skipper and Admiral add automation, policy enforcement, and compliance exports.
|
||||
Sencho integrates with [Trivy](https://trivy.dev) to scan container images and Compose files for vulnerabilities (CVEs), hardcoded secrets, and misconfigurations. Findings surface as severity badges in the Resources Hub and as drillable reports in the scan drawer. Manual scanning, secret and misconfig detection, scan history, comparison, and CVE suppressions are available on every tier. Skipper and Admiral add scheduled fleet scans, policy enforcement, SBOM, and SARIF exports.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/vulnerability-scanning/resources-badges.png" alt="Resources Hub showing vulnerability severity badges next to image tags" />
|
||||
<img src="/images/vulnerability-scanning/resources-badges.png" alt="Resources Hub Images table with severity badges (CRITICAL, HIGH, MEDIUM) on managed image rows alongside the Scan history button" />
|
||||
</Frame>
|
||||
|
||||
## Prerequisites
|
||||
@@ -14,7 +14,7 @@ Sencho integrates with [Trivy](https://trivy.dev) to scan container images for k
|
||||
The Trivy CLI must be available on the machine running Sencho. Trivy is not bundled with the Sencho Docker image; see [Installing Trivy](/operations/trivy-setup) for mount and installation options. Sencho checks for Trivy on startup and hides scanning UI when the binary is not available.
|
||||
|
||||
<Note>
|
||||
Trivy is installed independently on each Sencho instance. When you select a remote node, **Settings > Security** shows only the scanner status for that node; install, update, or uninstall Trivy from there to manage the remote's binary. Scan policies and CVE suppressions are managed on the control node and apply fleet-wide at read time.
|
||||
Trivy is installed independently on each Sencho instance. When you select a remote node, **Settings → Security** shows only the scanner status for that node; install, update, or uninstall Trivy from there to manage the remote's binary. Scan policies, CVE suppressions, and misconfig acknowledgements are managed on the control instance and replicate fleet-wide.
|
||||
</Note>
|
||||
|
||||
## Tier availability
|
||||
@@ -22,79 +22,84 @@ The Trivy CLI must be available on the machine running Sencho. Trivy is not bund
|
||||
| Feature | Community | Skipper | Admiral |
|
||||
|---------|:---------:|:-------:|:-------:|
|
||||
| Install / update / uninstall Trivy from Settings | ✓ | ✓ | ✓ |
|
||||
| On-demand image scanning (vulnerabilities) | ✓ | ✓ | ✓ |
|
||||
| Severity badges in the Resources Hub | ✓ | ✓ | ✓ |
|
||||
| Scan results drawer with vulnerability table | ✓ | ✓ | ✓ |
|
||||
| Post-deploy automated scanning | ✓ | ✓ | ✓ |
|
||||
| Secret detection in image filesystems | ✓ | ✓ | ✓ |
|
||||
| On-demand image vulnerability scanning | ✓ | ✓ | ✓ |
|
||||
| Full scan (vulnerabilities + secrets) | ✓ | ✓ | ✓ |
|
||||
| Compose file misconfiguration scanning | ✓ | ✓ | ✓ |
|
||||
| Scan history and comparison | ✓ | ✓ | ✓ |
|
||||
| Severity badges in the Resources Hub | ✓ | ✓ | ✓ |
|
||||
| Scan results drawer with grouped tabs | ✓ | ✓ | ✓ |
|
||||
| Post-deploy automated scanning | ✓ | ✓ | ✓ |
|
||||
| Scan history sheet | ✓ | ✓ | ✓ |
|
||||
| Scan comparison | ✓ | ✓ | ✓ |
|
||||
| CVE suppressions | ✓ | ✓ | ✓ |
|
||||
| Misconfig acknowledgements | ✓ | ✓ | ✓ |
|
||||
| Scheduled fleet scans (all images on a node) | | ✓ | ✓ |
|
||||
| Scan policies with `block_on_deploy` enforcement | | ✓ | ✓ |
|
||||
| SBOM generation (SPDX, CycloneDX) | | ✓ | ✓ |
|
||||
| SARIF export (code scanning integration) | | ✓ | ✓ |
|
||||
| Auto-update of the managed Trivy binary | | | ✓ |
|
||||
| Scheduled fleet scans (all images on a node) | | ✓ | ✓ |
|
||||
| Scan policies with `block_on_deploy` enforcement | | ✓ | ✓ |
|
||||
| SBOM generation (SPDX, CycloneDX) | | ✓ | ✓ |
|
||||
| SARIF export (code scanning integration) | | ✓ | ✓ |
|
||||
| Auto-update of the managed Trivy binary | | | ✓ |
|
||||
|
||||
## On-demand scanning
|
||||
|
||||
Navigate to the **Resources** tab and open the **Images** panel. When Trivy is available, every image row shows a shield icon alongside the delete action.
|
||||
Open the **Resources** tab and switch to the **Images** panel. When Trivy is available, every image row shows a shield icon alongside the inspect and delete actions. Click it and the menu offers two scan modes:
|
||||
|
||||
1. Click the shield icon on any image row. A menu appears with two options:
|
||||
- **Scan (vulnerabilities)**: the default, fastest path. Trivy inspects package metadata only.
|
||||
- **Full scan (vulnerabilities + secrets)**: additionally walks the image filesystem for hardcoded credentials, tokens, and keys. This takes noticeably longer.
|
||||
2. The row shows a loading spinner while Trivy runs. Most vulnerability scans finish in 10 to 60 seconds depending on image size and whether the Trivy database is already cached. Full scans add the time needed to read the filesystem.
|
||||
3. When the scan completes, a severity badge appears next to the image status (e.g. `CRITICAL`, `HIGH`, `MEDIUM`, `LOW`, or `CLEAN`).
|
||||
4. Click the badge to open the scan results drawer.
|
||||
- **Scan (vulnerabilities)**: the fast path. Trivy reads package metadata only.
|
||||
- **Full scan (vulnerabilities + secrets)**: also walks the image filesystem looking for hardcoded credentials, tokens, and keys. Slower because of the filesystem traversal.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/vulnerability-scanning/scan-details-sheet.png" alt="Vulnerability scan results drawer showing CVE table with severity, package, and fix columns alongside a policy violation banner" />
|
||||
</Frame>
|
||||
A spinner replaces the shield while the scan runs. Most vulnerability scans finish in 10 to 60 seconds depending on image size and whether the Trivy database is cached. Full scans add the time needed to read the filesystem. When the scan completes, a severity badge appears in the row's Status column. Click the badge to open the results drawer.
|
||||
|
||||
### Reading severity badges
|
||||
|
||||
Hovering over a severity badge reveals the breakdown of vulnerabilities by severity and the timestamp of the last scan. The badge color reflects the highest severity found:
|
||||
Hovering a badge reveals the breakdown by severity and the scan timestamp. The badge color reflects the highest severity found:
|
||||
|
||||
| Badge | Meaning |
|
||||
|-------|---------|
|
||||
| **CRITICAL** (red) | At least one vulnerability with a CVSS score of 9.0 or higher |
|
||||
| **HIGH** (amber) | At least one high-severity vulnerability |
|
||||
| **HIGH** (orange) | At least one high-severity vulnerability |
|
||||
| **MEDIUM** (blue) | Only medium and lower vulnerabilities |
|
||||
| **LOW** (muted) | Only low-severity vulnerabilities |
|
||||
| **CLEAN** (green) | Scan completed with zero findings |
|
||||
| **Clean** (green) | Scan completed with zero findings |
|
||||
|
||||
Scan results are cached by image digest. If the same digest is scanned again within 24 hours, Sencho returns the cached result instantly instead of re-running Trivy.
|
||||
Scan results are cached by image digest. If the same digest is scanned again within 24 hours, Sencho returns the cached result instantly instead of re-running Trivy. To force a fresh scan, open the drawer and click **Re-scan**.
|
||||
|
||||
## The scan results drawer
|
||||
|
||||
The drawer shows a full breakdown of the most recent scan for an image and groups findings across three tabs: **Vulnerabilities**, **Secrets**, and **Misconfigs**. The summary header shows counts across all three so you can see the full risk picture at a glance.
|
||||
The drawer opens as a right-side sheet with the breadcrumb `Security › Scans › <image-ref>`. The header surfaces the primary scan actions, the summary header reports severity counts and fixable findings, and the body groups findings across three tabs: **Vulnerabilities**, **Secrets**, and **Misconfigs**.
|
||||
|
||||
- **Summary**: counts per severity (critical, high, medium, low), total vulnerabilities, how many have a fix available, the Trivy version used, and when the scan ran.
|
||||
- **Vulnerabilities tab**: severity filter pills narrow the table, paginated list of every CVE found, including:
|
||||
- **CVE ID** (CVE-prefixed identifiers link to [cve.org](https://www.cve.org); GHSA identifiers link to the GitHub Advisory Database)
|
||||
- **Package** name and installed version
|
||||
- **Severity** badge
|
||||
- **Fixed version** with a green indicator if a fix is available
|
||||
<Frame>
|
||||
<img src="/images/vulnerability-scanning/scan-details-sheet.png" alt="Vulnerability scan drawer for lscr.io/linuxserver/swag:latest showing the Re-scan, Compare, CSV, and SARIF header actions, the severity summary (2 CRITICAL, 51 HIGH, 65 MEDIUM, 1 LOW), an SBOM download button, the severity filter pills, and a paged CVE table tinted by severity" />
|
||||
</Frame>
|
||||
|
||||
Critical and high rows in the table are tinted with a left accent rail so the rows that need attention catch the eye even before you read the severity column.
|
||||
|
||||
If the scan was evaluated against a [scan policy](#scan-policies) and the highest severity meets or exceeds the policy threshold, a destructive **Policy violation** banner appears at the top of the drawer naming the policy and the threshold it crossed.
|
||||
- **Secrets tab**: hardcoded credentials or keys detected in the image filesystem, with severity, rule, title, and the file/line location. Secret values are redacted: only the first eight characters of the match are stored.
|
||||
- **Misconfigs tab**: misconfiguration findings with severity, check ID, title, target file, and a suggested resolution. For image scans this tab is empty; for stack config scans (see below) it is the primary view.
|
||||
|
||||
### Actions
|
||||
|
||||
From the drawer header you can:
|
||||
### Header actions
|
||||
|
||||
- **Re-scan**: kick off a fresh scan, ignoring the digest cache.
|
||||
- **Download SBOM**: export a Software Bill of Materials in SPDX JSON or CycloneDX format (Skipper and Admiral).
|
||||
- **Export CSV**: export the full vulnerability list for offline review.
|
||||
- **Export SARIF**: download the full scan (vulnerabilities, secrets, and misconfigs) as SARIF 2.1.0 for upload to GitHub code scanning or other SARIF-aware tooling (Skipper and Admiral).
|
||||
- **Compare**: pick a baseline scan from the dropdown to diff against this one.
|
||||
- **CSV**: export the full vulnerability list for offline review.
|
||||
- **SARIF**: download the full scan (vulnerabilities, secrets, and misconfigs) as SARIF 2.1.0 for upload to GitHub code scanning or any SARIF-aware tool. Skipper or Admiral required.
|
||||
|
||||
The summary header below the actions reports the per-severity counts, the total, how many findings have a fix available, when the scan ran, and what triggered it. An **SBOM** button below the summary downloads a Software Bill of Materials in SPDX JSON or CycloneDX format (Skipper or Admiral).
|
||||
|
||||
### Vulnerabilities tab
|
||||
|
||||
Severity filter pills narrow the table, and the paginated list shows every CVE found with these columns:
|
||||
|
||||
- **CVE ID**: links to [cve.org](https://www.cve.org) for `CVE-…` identifiers and to the [GitHub Advisory Database](https://github.com/advisories) for `GHSA-…` identifiers.
|
||||
- **Package**: the affected package name and installed version.
|
||||
- **Severity**: badge matching the row tint.
|
||||
- **Fixed**: the version that contains the fix, with a green checkmark when an upgrade is available.
|
||||
|
||||
Critical and high rows carry a left accent rail so the entries that need attention catch the eye even before you read the severity column. If the scan was evaluated against a [scan policy](#scan-policies) and the highest severity meets or exceeds the policy threshold, a destructive **Policy violation** banner appears at the top of the drawer naming the policy and the threshold it crossed.
|
||||
|
||||
### Secrets tab
|
||||
|
||||
Hardcoded credentials or keys detected in the image filesystem, with severity, rule id, title, and the file path including the line number range when available. Secret values are redacted: only the first eight characters of the match are stored, so exporting or comparing a scan cannot leak the underlying credential.
|
||||
|
||||
### Misconfigs tab
|
||||
|
||||
Misconfiguration findings with severity, check id, title, target file, and a suggested fix. Image scans typically have an empty Misconfigs tab; stack-config scans (see [Compose misconfiguration scanning](#compose-misconfiguration-scanning)) populate it.
|
||||
|
||||
## Post-deploy automated scanning
|
||||
|
||||
When Trivy is available, Sencho automatically scans every deployed image in the background after a successful deploy. This applies to all deploy paths:
|
||||
When Trivy is available, Sencho scans every deployed image in the background after a successful deploy. This applies to all deploy paths:
|
||||
|
||||
- Stack deploy and redeploy
|
||||
- Stack update
|
||||
@@ -102,34 +107,34 @@ When Trivy is available, Sencho automatically scans every deployed image in the
|
||||
- Git source apply
|
||||
- Git source create
|
||||
|
||||
The deploy itself is never blocked by scanning; scans run asynchronously and surface their results via the severity badges in the Resources Hub. If high or critical vulnerabilities are found, an alert is dispatched through your configured [notification channels](/features/alerts-notifications).
|
||||
The deploy itself is never blocked by scanning. Scans run asynchronously and surface their results via the severity badges in the Resources Hub. If high or critical findings appear, an alert is dispatched through your configured [notification channels](/features/alerts-notifications).
|
||||
|
||||
### Opting out per deployment
|
||||
|
||||
The App Store deploy sheet includes an **Scan images for vulnerabilities after deploy** toggle (enabled by default). Uncheck it to skip the post-deploy scan for that single deployment. This does not disable scan policies globally; it simply opts this deployment out of scanning.
|
||||
The App Store deploy sheet has a **Security** section with a **Scan images for vulnerabilities after deploy** checkbox, enabled by default. Uncheck it to skip the post-deploy scan for that single deployment. This does not disable scan policies; it simply opts that deployment out of post-deploy scanning.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/vulnerability-scanning/app-store-toggle.png" alt="App Store deploy sheet showing the auto-scan checkbox enabled by default" />
|
||||
<img src="/images/vulnerability-scanning/app-store-toggle.png" alt="App Store deploy sheet Security section with the auto-scan checkbox enabled" />
|
||||
</Frame>
|
||||
|
||||
## Scheduled fleet scans
|
||||
|
||||
<Note>
|
||||
Scheduled scans require a **Skipper** or **Admiral** license.
|
||||
Scheduled fleet scans require a **Skipper** or **Admiral** license.
|
||||
</Note>
|
||||
|
||||
You can run recurring scans of every image on a node through the standard [Scheduled Operations](/features/scheduled-operations) system. Create a new scheduled task with action **Scan** and a cron expression. The scheduler iterates every image on the target node with a short delay between scans and records the result in the task's run history.
|
||||
You can run recurring scans of every image on a node through the standard [Scheduled Operations](/features/scheduled-operations) system. Create a scheduled task with action **Scan** and a cron expression. The scheduler iterates every image on the target node with a short delay between scans and records the result in the task's run history.
|
||||
|
||||
Use scheduled scans to keep CVE badges fresh even for images that are rarely redeployed: nightly (`0 3 * * *`) is a good default for most fleets.
|
||||
Use scheduled scans to keep badge counts fresh even for images that are rarely redeployed. A nightly cron like `0 3 * * *` is a sensible default for most fleets.
|
||||
|
||||
### Completion notifications
|
||||
|
||||
When a scheduled scan finishes, Sencho dispatches an alert through your configured [notification channels](/features/alerts-notifications). The alert includes the task name and a summary of scanned, cached, and failed images:
|
||||
When a scheduled scan finishes, Sencho dispatches an alert through your configured [notification channels](/features/alerts-notifications). The alert names the task and summarises scanned, cached, and failed image counts:
|
||||
|
||||
- **Info** when every image scanned successfully.
|
||||
- **Warning** when one or more images in the run failed to scan.
|
||||
- **Warning** when at least one image failed to scan.
|
||||
|
||||
Failures are typically transient (registry timeouts, missing credentials) and do not stop the rest of the run from completing. Check the task's run history for the detailed output.
|
||||
Failures are usually transient (registry timeouts, missing credentials) and never stop the rest of the run from completing. The task's run history holds the detailed output.
|
||||
|
||||
## Scan policies
|
||||
|
||||
@@ -137,12 +142,12 @@ Failures are typically transient (registry timeouts, missing credentials) and do
|
||||
Scan policies require a **Skipper** or **Admiral** license.
|
||||
</Note>
|
||||
|
||||
Policies let you define severity thresholds that govern whether a stack can deploy at all. A policy with **Block on deploy** enabled runs a pre-flight scan on every image in the stack before `docker compose up` executes; if any image meets or exceeds the threshold, the deploy is rejected with a dialog listing the offending images. Policies with **Block on deploy** disabled still evaluate every post-deploy and scheduled scan, and dispatch warning alerts when the threshold is exceeded.
|
||||
Policies define severity thresholds that govern whether a stack can deploy. A policy with **Block on deploy** enabled runs a pre-flight scan on every image in the stack before `docker compose up` executes; if any image meets or exceeds the threshold, the deploy is rejected with a dialog listing the offending images. Policies with **Block on deploy** disabled still evaluate every post-deploy and scheduled scan and dispatch warning alerts when the threshold is exceeded.
|
||||
|
||||
See [Deploy Enforcement](/features/deploy-enforcement) for the full pre-flight flow, admin bypass path, and audit-log behavior.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/vulnerability-scanning/security-settings.png" alt="Security section of Settings showing the scan policies list with add policy button" />
|
||||
<img src="/images/vulnerability-scanning/security-settings.png" alt="Settings Security page showing the masthead with Node, Policies, and Trivy status, the Vulnerability Scanner card with Auto-update Trivy toggle, an Add Policy button, and the No scan policies configured empty state" />
|
||||
</Frame>
|
||||
|
||||
### Creating a policy
|
||||
@@ -151,7 +156,7 @@ Go to **Settings → Security** and click **Add Policy**.
|
||||
|
||||
| Field | Description |
|
||||
|-------|-------------|
|
||||
| **Name** | A descriptive label (e.g. "Production critical block"). |
|
||||
| **Name** | A descriptive label (e.g. `prod-critical-block`). |
|
||||
| **Stack pattern** | Optional glob against stack names (e.g. `prod-*`). Leave empty to match every stack. |
|
||||
| **Max severity** | The threshold. If a scan finds any vulnerability at or above this severity, the policy fires. |
|
||||
| **Block on deploy** | When enabled, deploys are rejected before `docker compose up` runs if any image violates the threshold. When disabled, the policy still evaluates post-deploy and scheduled scans and dispatches warning alerts on violations. |
|
||||
@@ -165,7 +170,7 @@ When multiple policies match a deploy, Sencho picks the most specific one:
|
||||
2. Policies with a stack pattern win over wildcard policies.
|
||||
3. Disabled policies are never applied.
|
||||
|
||||
Only one policy is evaluated per deploy; use a single tight pattern rather than overlapping policies for clarity.
|
||||
Only one policy is evaluated per deploy. Use a single tight pattern rather than overlapping policies for clarity.
|
||||
|
||||
<Note>
|
||||
Policies created on a control instance replicate to every remote in the fleet automatically. See [Fleet Sync](/features/fleet-sync) for the replication, push retry, and replica demote behavior.
|
||||
@@ -206,32 +211,11 @@ curl -X POST https://your-sencho-instance:1852/api/security/policies \
|
||||
}'
|
||||
```
|
||||
|
||||
## SBOM generation
|
||||
|
||||
<Note>
|
||||
SBOM generation requires a **Skipper** or **Admiral** license.
|
||||
</Note>
|
||||
|
||||
A Software Bill of Materials (SBOM) is a machine-readable inventory of every package present in a container image. SBOMs are required by an increasing number of security frameworks (SLSA, Executive Order 14028, EU Cyber Resilience Act) and are useful for offline compliance reviews.
|
||||
|
||||
From the scan results drawer, click **Download SBOM** and choose a format:
|
||||
|
||||
| Format | Use case |
|
||||
|--------|----------|
|
||||
| **SPDX JSON** | Widely supported, best for tooling integration. |
|
||||
| **CycloneDX** | Richer dependency metadata, better for supply-chain analysis. |
|
||||
|
||||
The download starts immediately and uses the image's digest (when available) in the filename.
|
||||
|
||||
## Secret detection
|
||||
|
||||
<Note>
|
||||
Secret detection requires a **Skipper** or **Admiral** license.
|
||||
</Note>
|
||||
Full scans ask Trivy to walk the image filesystem for patterns that look like hardcoded credentials, API tokens, cloud access keys, or private keys. Detection rules cover the common providers (AWS, GCP, GitHub, Slack, Stripe) plus generic high-entropy strings.
|
||||
|
||||
Full scans ask Trivy to walk the image filesystem for patterns that look like hardcoded credentials, API tokens, cloud access keys, or private keys. Detection rules cover common providers (AWS, GCP, GitHub, Slack, Stripe) plus generic high-entropy strings.
|
||||
|
||||
To run a full scan, click the shield icon in the Resources Hub and pick **Full scan (vulnerabilities + secrets)**. Findings appear on the **Secrets** tab of the scan drawer:
|
||||
Trigger a full scan from the shield-icon menu on any image row with **Full scan (vulnerabilities + secrets)**. Findings appear on the **Secrets** tab of the drawer:
|
||||
|
||||
| Column | Description |
|
||||
|--------|-------------|
|
||||
@@ -240,50 +224,52 @@ To run a full scan, click the shield icon in the Resources Hub and pick **Full s
|
||||
| **Title** | A short description of what was detected. The second line shows a redacted excerpt of the match. |
|
||||
| **Target** | The file path inside the image filesystem, including the line number range when available. |
|
||||
|
||||
Only the first eight characters of any matched secret are stored, followed by an ellipsis. The full value is never written to the database, so exporting a scan drawer or comparing scans cannot leak the underlying credential.
|
||||
Only the first eight characters of any matched secret are persisted. The full value is never written to the database, so exporting a scan or comparing two scans cannot leak the underlying credential.
|
||||
|
||||
Full scans take longer than vulnerability-only scans because Trivy reads every file in the image. If runtime is a concern, schedule full scans overnight and keep deploy-time scans on the default vulnerability-only setting.
|
||||
|
||||
## Compose misconfiguration scanning
|
||||
|
||||
<Note>
|
||||
Compose misconfiguration scanning requires a **Skipper** or **Admiral** license.
|
||||
</Note>
|
||||
Beyond package CVEs, Sencho can run `trivy config` against a stack's Compose file to flag insecure defaults before you deploy. Typical checks cover privileged containers, missing resource limits, host networking, mounted Docker sockets, and overly broad capabilities.
|
||||
|
||||
Beyond package CVEs, Sencho can run `trivy config` against a stack's Compose file to flag insecure defaults before you deploy. Typical checks include containers running as root, missing resource limits, privileged mode, host network, and mounted Docker sockets.
|
||||
From any stack page, open the **More actions** overflow menu next to the Update button and select **Scan config**. Sencho runs the scanner against the stack's working directory and opens the scan drawer on the **Misconfigs** tab.
|
||||
|
||||
From any stack page, click **Scan config** next to the Deploy controls. Sencho runs the scanner against the stack's working directory and opens the scan drawer on the **Misconfigs** tab:
|
||||
<Frame>
|
||||
<img src="/images/vulnerability-scanning/scan-config-button.png" alt="Plex stack page with the More actions overflow menu open, showing the Scan config item with a shield icon and the Delete item below" />
|
||||
</Frame>
|
||||
|
||||
The Misconfigs tab columns:
|
||||
|
||||
| Column | Description |
|
||||
|--------|-------------|
|
||||
| **Severity** | Rule severity (CRITICAL/HIGH/MEDIUM/LOW). |
|
||||
| **Check** | Rule ID or AVD identifier for the violated check. |
|
||||
| **Severity** | Rule severity (CRITICAL / HIGH / MEDIUM / LOW). |
|
||||
| **Check** | Rule id or AVD identifier for the violated check. |
|
||||
| **Title** | Short summary, linked to the upstream advisory when available. The second line shows Trivy's detailed message. |
|
||||
| **Target** | The file that triggered the finding. |
|
||||
| **Fix** | The recommended resolution. |
|
||||
|
||||
Config scans are stored in the same history as image scans with an `image_ref` of `stack:<name>`, so they appear on the Scan history page and can be exported as CSV.
|
||||
Config scans are stored in the same history as image scans with an `image_ref` of `stack:<name>`, so they appear on the Scan history sheet and can be exported as CSV.
|
||||
|
||||
## Misconfig acknowledgements
|
||||
|
||||
Some misconfigurations are intentional. A reverse-proxy stack legitimately needs root to bind privileged ports; a network monitor might require host networking; an `--privileged` Docker socket mount might be exactly what your janitor service expects. Sencho lets admins acknowledge a rule so it stops triggering alerts without lowering the policy bar for every other stack.
|
||||
Some misconfigurations are intentional. A reverse proxy legitimately needs root to bind privileged ports; a network monitor might require host networking; an `--privileged` mount of the Docker socket might be exactly what your janitor service expects. Sencho lets admins acknowledge a rule so it stops triggering alerts on that stack without lowering the policy bar for every other stack.
|
||||
|
||||
Acknowledgements never modify stored finding rows. They are applied at read time, so deleting an acknowledgement immediately resurfaces the finding wherever it appears.
|
||||
Acknowledgements are applied at read time and never modify stored findings. Deleting an acknowledgement immediately resurfaces the underlying finding wherever it appears.
|
||||
|
||||
### Acknowledging from a scan result
|
||||
|
||||
1. Open a stack config scan that contains the finding.
|
||||
2. Click the shield-check icon at the right edge of the misconfig row. The dialog opens with the rule id prefilled and the stack pattern set to the current stack name (so a single click acknowledges *only this stack* — narrowest possible scope by default).
|
||||
3. Add a reason explaining why the misconfiguration is accepted. The reason is stored locally and replicated fleet-wide; it never appears in audit-log summaries to avoid leaking incident-tracker IDs or vendor secrets.
|
||||
1. Open a stack-config scan that contains the finding.
|
||||
2. Click the shield-check icon at the right edge of the misconfig row. The dialog opens with the rule id prefilled and the stack pattern set to the current stack name, so a single click acknowledges *only this stack*. That is the narrowest possible scope by default.
|
||||
3. Add a reason explaining why the misconfiguration is accepted. The reason is stored locally and replicates fleet-wide; it never appears in audit-log summaries to avoid leaking incident-tracker IDs or vendor secrets.
|
||||
4. Optionally set an expiry in days. After expiry the acknowledgement stops applying and the finding resurfaces.
|
||||
|
||||
The acknowledged row renders dimmed with a strikethrough title; hovering surfaces the acknowledgement reason.
|
||||
Acknowledged rows render dimmed with a strikethrough title; hovering surfaces the recorded reason.
|
||||
|
||||
### Managing acknowledgements
|
||||
|
||||
**Settings > Security** has a panel listing every acknowledgement on this control: rule id, optional stack pattern (glob), creator, expiry date, and a delete button. The same `replicated` badge that appears on CVE suppressions appears here for rows pushed from the control to a replica.
|
||||
**Settings → Security** has a Misconfig Acknowledgements panel listing every acknowledgement on this control: rule id, optional stack pattern (glob), creator, expiry date, and a delete button. The same `replicated` badge that appears on CVE suppressions appears here for rows pushed from the control to a replica.
|
||||
|
||||
Replicas show the panel read-only — write operations return 403 with a "managed by control" message so configuration drift cannot accumulate on the leaf nodes.
|
||||
Replicas show the panel read-only; write operations return 403 with a "managed by control" message so configuration drift cannot accumulate on the leaf nodes.
|
||||
|
||||
### Scope and matching
|
||||
|
||||
@@ -297,7 +283,24 @@ When more than one acknowledgement could match, Sencho picks the most specific:
|
||||
|
||||
### SARIF emission
|
||||
|
||||
Acknowledged misconfigs are emitted in the SARIF export with a `suppressions` entry of kind `external` and status `accepted`, mirroring CVE suppressions. Code-scanning dashboards that respect SARIF suppressions will dismiss them with the recorded justification.
|
||||
Acknowledged misconfigs are emitted in the SARIF export with a `suppressions` entry of kind `external` and status `accepted`, mirroring CVE suppressions. Code-scanning dashboards that respect SARIF suppressions dismiss them with the recorded justification.
|
||||
|
||||
## SBOM generation
|
||||
|
||||
<Note>
|
||||
SBOM generation requires a **Skipper** or **Admiral** license.
|
||||
</Note>
|
||||
|
||||
A Software Bill of Materials (SBOM) is a machine-readable inventory of every package in a container image. SBOMs satisfy security frameworks (SLSA, Executive Order 14028, EU Cyber Resilience Act) and support offline supply-chain analysis.
|
||||
|
||||
From the scan results drawer, click **SBOM** below the summary and choose a format:
|
||||
|
||||
| Format | Use case |
|
||||
|--------|----------|
|
||||
| **SPDX JSON** | Widely supported, best for tooling integration. |
|
||||
| **CycloneDX** | Richer dependency metadata, better for supply-chain analysis. |
|
||||
|
||||
The download starts immediately and uses the image's digest (when available) in the filename.
|
||||
|
||||
## SARIF export
|
||||
|
||||
@@ -305,19 +308,19 @@ Acknowledged misconfigs are emitted in the SARIF export with a `suppressions` en
|
||||
SARIF export requires a **Skipper** or **Admiral** license.
|
||||
</Note>
|
||||
|
||||
SARIF (Static Analysis Results Interchange Format) is the standard format supported by GitHub code scanning, Microsoft Defender for Cloud, and most security dashboards. Sencho generates SARIF 2.1.0 documents directly from the stored scan results so the download matches what you see in the drawer (same findings, same suppression state) without re-running Trivy.
|
||||
SARIF (Static Analysis Results Interchange Format) is the standard format supported by GitHub code scanning, Microsoft Defender for Cloud, and most security dashboards. Sencho generates SARIF 2.1.0 documents from the stored scan results, so the download matches what you see in the drawer (same findings, same suppression state) without re-running Trivy.
|
||||
|
||||
From the scan drawer header, click **SARIF** to download the report. The file is named after the image reference with a `.sarif.json` extension.
|
||||
|
||||
What the export contains:
|
||||
|
||||
- **Vulnerabilities**: one SARIF result per CVE, with `security-severity` scored 9.8 (CRITICAL), 7.5 (HIGH), 5.0 (MEDIUM), 2.5 (LOW), or 0.0 (UNKNOWN). The affected package appears as a logical location (`<pkg>@<version>`).
|
||||
- **Secrets**: rule IDs are namespaced as `SECRET:<rule>`. Results point at the file and line number where the match was found.
|
||||
- **Misconfigs**: rule IDs are namespaced as `MISCONFIG:<rule>`. Results point at the Compose file that triggered the check.
|
||||
- **Suppressions**: CVEs you have suppressed in Sencho are emitted with a SARIF `suppressions` entry of kind `external` and status `accepted`, so code-scanning dashboards can dismiss them with the justification you recorded.
|
||||
- **Secrets**: rule ids are namespaced as `SECRET:<rule>`. Results point at the file and line number where the match was found.
|
||||
- **Misconfigs**: rule ids are namespaced as `MISCONFIG:<rule>`. Results point at the Compose file that triggered the check.
|
||||
- **Suppressions**: CVEs you have suppressed in Sencho are emitted with a SARIF `suppressions` entry of kind `external` and status `accepted`, so code-scanning dashboards dismiss them with the justification you recorded.
|
||||
- **Acknowledged misconfigs**: misconfigs you have acknowledged in Sencho are emitted with the same `suppressions` shape so dashboards apply the same dismissal logic.
|
||||
|
||||
The export caps each finding type at 5000 rows to bound memory and serialisation time on pathological scans. When any type trips the cap, the SARIF run carries `properties.truncated = true`, `properties.row_limit`, and a `properties.totals` object with the original counts so downstream tooling can flag the export as partial.
|
||||
Each finding type is capped at 5000 rows to bound memory and serialisation time on pathological scans. When any type trips the cap, the SARIF run carries `properties.truncated = true`, `properties.row_limit`, and a `properties.totals` object with the original counts so downstream tooling can flag the export as partial.
|
||||
|
||||
Typical upload flow for GitHub code scanning:
|
||||
|
||||
@@ -330,22 +333,19 @@ Typical upload flow for GitHub code scanning:
|
||||
|
||||
## Scan history
|
||||
|
||||
Every scan Sencho runs is stored with its full vulnerability detail. Scan records are automatically pruned after 90 days to keep the database compact. The history is used to power:
|
||||
Every scan Sencho runs is stored with its full vulnerability detail. Scan records are pruned automatically after 90 days to keep the database compact. The history powers two things: digest caching (skip re-scanning a digest already scanned within 24 hours) and trend insight (compare a new scan to its predecessor to see what changed).
|
||||
|
||||
- **Digest caching**: skip re-scanning an image that has already been scanned within 24 hours.
|
||||
- **Trend badges**: surface whether the latest scan added or resolved vulnerabilities compared to the previous scan for the same image.
|
||||
|
||||
Click **Scan history** from the top of the Resources Hub to open a right-side sheet layered over the current page. The sheet lists completed scans grouped by image, lets you search by image reference, and lets you pick two scans to compare. Close the sheet by pressing Escape, clicking the overlay, or clicking the close button in the header.
|
||||
Click **Scan history** from the top of the Resources Hub to open the scan history sheet over the current page. The sheet lists completed scans grouped by image, lets you search by image reference, and lets you tick two scans to compare. Close the sheet with Escape, by clicking the overlay, or by clicking the close button in the header.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/vulnerability-scanning/scan-history-sheet.png" alt="Scan history sheet overlaid on the Resources Hub, with search box, pagination, and scan rows grouped by image" />
|
||||
<img src="/images/vulnerability-scanning/scan-history-sheet.png" alt="Scan history sheet over the Resources Hub showing 361 scans across 29 images, the Compare primary action enabled with two scans ticked, a search box, pagination, and scans grouped by image reference" />
|
||||
</Frame>
|
||||
|
||||
## Comparing scans <Badge>Skipper</Badge>
|
||||
## Comparing scans
|
||||
|
||||
Compare any two completed scans for an image to see what changed between them.
|
||||
Compare any two completed scans for an image to see what changed.
|
||||
|
||||
**From the Scan history page**: select two scans via the checkboxes (one baseline, one newer) and click **Compare**. Selecting a third scan replaces the oldest selection.
|
||||
**From the scan history sheet**: tick two scans (one baseline, one newer) and click **Compare**. Selecting a third scan replaces the oldest selection.
|
||||
|
||||
**From an open scan**: click **Compare** in the drawer header, then pick a baseline scan from the dropdown. Only completed scans for the same image appear.
|
||||
|
||||
@@ -353,21 +353,21 @@ The comparison sheet shows:
|
||||
|
||||
- A **delta ribbon** summarizing the net change per severity (CRITICAL, HIGH, MEDIUM, LOW). Net-positive deltas on CRITICAL render in destructive red so a regression on the worst tier is immediately visible.
|
||||
- Filter pills to switch between **Added** (new findings since the baseline), **Removed** (resolved findings), and **Unchanged** (findings present in both).
|
||||
- A sorted table of CVEs with severity, affected package, and direct links to the upstream advisory. CVE-prefixed identifiers resolve to [cve.org](https://www.cve.org); GHSA identifiers resolve to the GitHub Advisory Database. Critical and high rows carry the same left-rail tint as the scan results drawer for visual continuity.
|
||||
- A sorted table of CVEs with severity, affected package, and direct links to the upstream advisory. CVE-prefixed identifiers resolve to [cve.org](https://www.cve.org); GHSA identifiers resolve to the [GitHub Advisory Database](https://github.com/advisories). Critical and high rows carry the same left-rail tint as the scan drawer for visual continuity.
|
||||
|
||||
Comparisons are scoped to a single node; scans taken on different nodes cannot be compared against each other.
|
||||
|
||||
Cross-image comparisons (picking scans from two different image references) are allowed but flagged with a warning, since package-level changes may reflect image differences rather than CVE drift. In that mode, the **Unchanged** pill is relabeled to **Shared** to reflect that same CVE + package matches across different images are not necessarily the same finding.
|
||||
|
||||
Up to 1000 findings per scan are loaded for comparison. When a scan exceeds this limit, the sheet shows a banner indicating the comparison may be incomplete.
|
||||
Up to 1000 findings per scan are loaded for comparison. When a scan exceeds this limit, the sheet shows a truncation banner indicating the comparison may be incomplete.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/vulnerability-scanning/scan-compare-sheet.png" alt="Compare scans sheet with delta ribbon, Added/Removed/Unchanged filter pills, and a diff table tinted by severity" />
|
||||
<img src="/images/vulnerability-scanning/scan-compare-sheet.png" alt="Compare scans sheet diffing two profilarr scans, with baseline and current row, a truncation banner, a delta ribbon (HIGH -79, MEDIUM +79), Added/Removed/Unchanged filter pills, and a diff table tinted by severity" />
|
||||
</Frame>
|
||||
|
||||
## How it works
|
||||
|
||||
1. On startup, Sencho looks for the `trivy` binary on `PATH` and caches its availability.
|
||||
1. On startup, Sencho looks for the `trivy` binary on `PATH` (and the path configured by `TRIVY_BIN` if set) and caches its availability.
|
||||
2. When a scan is triggered, Sencho resolves the image digest via Docker and checks the 24-hour cache. If a completed scan exists for that digest, it is returned instantly.
|
||||
3. Otherwise, Sencho spawns `trivy image --format json --quiet <image-ref>` and parses the JSON output into the vulnerability database.
|
||||
4. Private registry credentials are forwarded automatically by writing a temporary `DOCKER_CONFIG` for the Trivy subprocess, then deleting it when the scan completes.
|
||||
@@ -375,101 +375,77 @@ Up to 1000 findings per scan are loaded for comparison. When a scan exceeds this
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### Scan button is not visible
|
||||
<AccordionGroup>
|
||||
<Accordion title="Scan button is not visible">
|
||||
Sencho hides scanning UI when the Trivy binary is not detected. Check **Settings → Security** for the scanner status, then follow [Installing Trivy](/operations/trivy-setup) if it is missing.
|
||||
</Accordion>
|
||||
<Accordion title="Scans time out">
|
||||
The default scan timeout is 5 minutes. Very large images (2 GB or more) over a slow connection may exceed this; pre-pulling the image to the host speeds up the scan significantly because Trivy then works against the local image store.
|
||||
</Accordion>
|
||||
<Accordion title="Private registry images fail to scan">
|
||||
Sencho forwards the same registry credentials configured under **Settings → Registries** to Trivy. If a pull works in Sencho but a scan fails, make sure the image has been pulled at least once so Trivy can work against the cached local image.
|
||||
</Accordion>
|
||||
<Accordion title="Badge is out of date after an image update">
|
||||
Post-deploy scanning only runs on deploy actions. For long-running images that aren't redeployed, schedule a recurring scan (Skipper or Admiral) or click the shield icon in the Resources Hub to re-scan on demand.
|
||||
</Accordion>
|
||||
<Accordion title="A scan shows in progress for a long time">
|
||||
Scans have a 5-minute internal timeout. The scheduler sweeps every tick and marks any scan stuck `in_progress` for more than 15 minutes as failed, so the UI always recovers on its own. Wait for the sweep, then click the shield icon again to trigger a fresh scan.
|
||||
</Accordion>
|
||||
<Accordion title="Trivy is not detected after installing it">
|
||||
Sencho re-checks for the Trivy binary every ten minutes from the scheduler. Install Trivy on the host, wait for the next window, and the scanning UI lights up automatically.
|
||||
</Accordion>
|
||||
<Accordion title="Collecting diagnostic logs for support">
|
||||
Enable **Developer Mode** under **Settings → Developer** and trigger the failing scan again. The backend logs verbose `[Trivy:diag]` entries covering detection timing, cache hits, Trivy invocation, and parse statistics. Attach these lines when filing a support issue. Turn Developer Mode off once you have captured the output.
|
||||
</Accordion>
|
||||
<Accordion title="Post-deploy scan failure notifications">
|
||||
When a post-deploy scan fails for a specific image (for example because Trivy could not resolve a private registry pull), Sencho dispatches a warning-level alert through your configured notification channels. The deploy itself is never blocked by a scan failure.
|
||||
</Accordion>
|
||||
<Accordion title="A deploy was blocked by a policy I did not expect">
|
||||
The block dialog names the policy that fired and lists every image that violated the threshold. Open **Settings → Security** and review the matching policy: check the stack pattern glob and the max severity. If the policy should not apply, tighten the pattern (for example `staging-*` instead of `*`) or turn **Block on deploy** off to keep the evaluation in alert-only mode. Admins can bypass a single deploy with the **Deploy anyway** button; every bypass is recorded in the [Audit Log](/features/audit-log) with the actor, policy, and violation list.
|
||||
</Accordion>
|
||||
<Accordion title="Trivy is not installed and a deploy with a block policy went through">
|
||||
Sencho fails open when Trivy is not installed on the target node, so operators are never locked out by tooling state. A warning alert is dispatched through your configured notification channels with the message `Pre-deploy scan for "<stack>" skipped: Trivy not installed on this node`. Install Trivy from **Settings → Security** to enforce the policy; see [Installing Trivy](/operations/trivy-setup) for options.
|
||||
</Accordion>
|
||||
<Accordion title="Compare button is disabled in the scan history sheet">
|
||||
The Compare primary action enables only after exactly two scans are ticked. Selecting zero, one, or three scans leaves it disabled. If you have only one scan for an image, trigger a second scan from the Resources Hub (or wait for a scheduled scan), then return to Scan history and tick both.
|
||||
</Accordion>
|
||||
<Accordion title="Comparison shows unexpected results">
|
||||
Two common causes:
|
||||
|
||||
Sencho hides scanning UI when the Trivy binary is not detected. Check **Settings → Support** for the Trivy availability status, then follow [Installing Trivy](/operations/trivy-setup) if it is missing.
|
||||
|
||||
### Scans time out
|
||||
|
||||
The default scan timeout is 5 minutes. Very large images (2+ GB) over a slow connection may exceed this; pre-pulling the image to the host speeds up the scan significantly because Trivy works against the local image store.
|
||||
|
||||
### Private registry images fail to scan
|
||||
|
||||
Sencho forwards the same registry credentials configured under **Settings → Registries** to Trivy during a scan. If a pull works in Sencho but a scan fails, make sure the image has been pulled at least once (Trivy can then work against the cached local image).
|
||||
|
||||
### Badge is out of date after an image update
|
||||
|
||||
Post-deploy scanning only runs on deploy actions. For long-running images that aren't redeployed, schedule a recurring scan (Skipper+) or click the shield icon in the Resources Hub to re-scan on demand.
|
||||
|
||||
### A scan shows "in progress" for a long time
|
||||
|
||||
Scans have a 5-minute internal timeout. The scheduler also sweeps every tick and marks any scan that has been `in progress` for more than 15 minutes as failed, so the UI always recovers on its own. Wait for the sweep to run, then click the shield icon again to start a fresh scan.
|
||||
|
||||
### Trivy is not detected after installing it
|
||||
|
||||
Sencho re-checks for the Trivy binary every ten minutes from the scheduler. Install Trivy on the host, wait for the next window, and the scanning UI lights up automatically.
|
||||
|
||||
### Collecting diagnostic logs for support
|
||||
|
||||
Enable **Developer Mode** under **Settings → Developer** and trigger the failing scan again. The backend logs verbose `[Trivy:diag]` entries covering detection timing, cache hits, Trivy invocation, and parse statistics. Attach these lines when filing a support issue. Turn Developer Mode off once you have captured the output.
|
||||
|
||||
### Post-deploy scan failure notifications
|
||||
|
||||
When a post-deploy scan fails for a specific image (for example because Trivy could not resolve a private registry pull), Sencho dispatches a warning-level alert through your configured notification channels. The deploy itself is never blocked by a scan failure.
|
||||
|
||||
### A deploy was blocked by a policy I did not expect
|
||||
|
||||
The block dialog names the policy that fired and lists every image that violated the threshold. Open **Settings → Security → Scan Policies** and review the matching policy: check the stack pattern glob and the max severity. If the policy should not apply, tighten the pattern (for example `staging-*` instead of `*`) or turn **Block on deploy** off to keep the evaluation in alert-only mode. Admins can also bypass a single deploy with the **Deploy anyway** button; every bypass is recorded in the [Audit Log](/features/audit-log) with the actor, policy, and violation list.
|
||||
|
||||
### Trivy is not installed and a deploy with a block policy went through
|
||||
|
||||
Sencho fails open when Trivy is not installed on the target node, so users are never locked out by tooling state. A warning alert is dispatched through your configured notification channels with the message `Pre-deploy scan for "<stack>" skipped: Trivy not installed on this node`. Install Trivy from **Settings → Security** to enforce the policy; see [Installing Trivy](/operations/trivy-setup) for options.
|
||||
|
||||
### Compare button is disabled
|
||||
|
||||
Scan comparison is a Skipper feature; on Community, the Compare button stays disabled with a tooltip explaining the upgrade path. If your license is Skipper or Admiral, make sure you have ticked exactly two completed scans: selecting zero, one, or three scans leaves the button disabled. If you have only one scan for an image, trigger a second scan from the Resources Hub (or wait for a scheduled scan), then return to the Scan history page and tick both.
|
||||
|
||||
### Comparison shows unexpected results
|
||||
|
||||
Two common causes:
|
||||
|
||||
- **Cross-image comparison.** If the Baseline and Current rows at the top of the sheet point at different image references, the warning banner is shown and the "Unchanged" pill is labeled "Shared". Items in that bucket match on CVE + package name but may not be the same finding across two distinct images. Stick to scans of the same image reference for apples-to-apples drift analysis.
|
||||
- **Truncated scans.** When either scan has more than 1000 stored findings, the sheet shows a truncation banner. In that case, items past the 1000-row cap do not contribute to the Added / Removed / Unchanged buckets and the totals may be misleading. Re-run the scan with a tighter image (or scope the investigation to the most severe findings) to avoid truncation.
|
||||
|
||||
### An older scan is missing from Scan history
|
||||
|
||||
The Scan history page uses server-driven pagination. If you know the scan exists but cannot see it, use the search box to filter by image reference, or page forward with the arrows in the card header. Scans older than 90 days are pruned automatically to keep the database compact.
|
||||
|
||||
### Scan policies are missing on one of my nodes
|
||||
|
||||
Scan policies are managed from the control Sencho instance and replicate to every remote. When you view **Settings → Security** on a remote Sencho (a replica), you will see a banner explaining that rules are managed upstream. See [Fleet Sync](/features/fleet-sync) for how replication works and how to investigate push failures.
|
||||
|
||||
### I suppressed a CVE but the scan badge count is unchanged
|
||||
|
||||
Badge counts reflect the raw findings so alerting and policy evaluation stay accurate. Open the scan drawer to confirm the row is dimmed with a shield-off icon. See [CVE Suppressions](/features/cve-suppressions) for how the filter is applied across the drawer, compare sheet, and other read surfaces.
|
||||
|
||||
### The Secrets tab is empty on an image I expect to contain credentials
|
||||
|
||||
Secret detection matches against Trivy's built-in rule set, which focuses on well-known provider patterns. Plain text passwords, custom token formats, or values that do not match any published rule will not appear. Ensure you picked **Full scan (vulnerabilities + secrets)** from the shield-icon menu; a plain vulnerability scan does not walk the filesystem.
|
||||
|
||||
### Scan config button is disabled on a stack
|
||||
|
||||
The button is only shown when Trivy is available on the stack's node, the current user is an admin, and the license is Skipper or Admiral. If all three conditions are met but the button stays disabled, another stack action (deploy, update, rollback) is still in progress; wait for it to finish.
|
||||
|
||||
### Compose misconfiguration scan returns 404
|
||||
|
||||
The scanner needs to locate a Compose file in the stack directory. If the stack was created outside Sencho and the working directory does not contain a file named `compose.yml`, `compose.yaml`, `docker-compose.yml`, or `docker-compose.yaml`, the scan returns 404. Name the file accordingly or keep the stack under Sencho's managed compose directory.
|
||||
|
||||
### SARIF download returns 409 "Scan not complete"
|
||||
|
||||
SARIF export requires a completed scan. If a scan failed, timed out, or is still running, the button downloads nothing and the server returns a 409. Trigger a fresh scan from the Resources Hub or the stack page, wait for the drawer to populate, then export again.
|
||||
|
||||
### SARIF download is missing findings I see in the drawer
|
||||
|
||||
Each finding type (vulnerabilities, secrets, misconfigs) is capped at 5000 rows in the SARIF export. When any type trips the cap, the SARIF run includes `properties.truncated: true` along with the original counts, and the server logs a warning. The drawer paginates beyond 5000 so it shows everything, but the export is bounded for memory safety. If you need every row, narrow the scope (per-stack scan, per-image scan) before exporting.
|
||||
|
||||
### Acknowledge button is missing on a misconfig finding
|
||||
|
||||
The button is admin-only and hidden on replica nodes. Replicas read acknowledgements from the control via fleet sync; mutations must happen on the control. If you are an admin on the control and still do not see the button, the row is likely already acknowledged — look for the dimmed/strikethrough rendering and hover for the recorded reason.
|
||||
|
||||
### Findings reappeared after deleting a CVE suppression or misconfig acknowledgement
|
||||
|
||||
Suppressions and acknowledgements are applied at read time and never modify the stored finding rows. Removing one immediately resurfaces the underlying finding everywhere it appears (drawer, compare sheet, badge counts, SARIF export). The behaviour is intentional: deleting an acknowledgement is meant to revert the operator decision, not to mask history.
|
||||
|
||||
### Outbound traffic to ghcr.io / aquasecurity from the scanner host
|
||||
|
||||
Trivy itself fetches its CVE and secret-rule database from public registries on first scan and refreshes periodically; that egress is required for vulnerability scanning to work. Sencho does not emit telemetry of its own. If your environment forbids egress, see Trivy's [air-gapped scanning guide](https://aquasecurity.github.io/trivy/latest/docs/advanced/air-gap/) for how to pre-seed the database and run scans with `--offline-scan`.
|
||||
|
||||
### Compose stack scan returns 409 "Already scanning this stack"
|
||||
|
||||
Sencho deduplicates concurrent scans of the same stack so two simultaneous calls cannot double-process the result. Wait for the in-flight scan to finish (its row appears in Scan history with status `in_progress` and flips to `completed` or `failed` shortly after) and trigger again.
|
||||
- **Cross-image comparison.** When the Baseline and Current rows at the top of the sheet point at different image references, a warning banner appears and the "Unchanged" pill is labeled "Shared". Items in that bucket match on CVE + package name but may not be the same finding across two distinct images. Stick to scans of the same image reference for apples-to-apples drift analysis.
|
||||
- **Truncated scans.** When either scan has more than 1000 stored findings, the sheet shows a truncation banner. Items past the 1000-row cap do not contribute to the Added / Removed / Unchanged buckets and totals may be misleading. Re-run the scan with a tighter image or scope the investigation to the most severe findings to avoid truncation.
|
||||
</Accordion>
|
||||
<Accordion title="An older scan is missing from Scan history">
|
||||
The Scan history sheet uses server-driven pagination. If you know the scan exists but cannot see it, use the search box to filter by image reference, or page forward with the arrows in the card header. Scans older than 90 days are pruned automatically to keep the database compact.
|
||||
</Accordion>
|
||||
<Accordion title="Scan policies are missing on one of my nodes">
|
||||
Scan policies are managed from the control Sencho instance and replicate to every remote. On a replica, **Settings → Security** shows a banner explaining that rules are managed upstream. See [Fleet Sync](/features/fleet-sync) for how replication works and how to investigate push failures.
|
||||
</Accordion>
|
||||
<Accordion title="I suppressed a CVE but the scan badge count is unchanged">
|
||||
Badge counts reflect raw findings so alerting and policy evaluation stay accurate. Open the scan drawer to confirm the row is dimmed with a shield-off icon. See [CVE Suppressions](/features/cve-suppressions) for how the filter is applied across the drawer, compare sheet, and other read surfaces.
|
||||
</Accordion>
|
||||
<Accordion title="The Secrets tab is empty on an image I expect to contain credentials">
|
||||
Secret detection matches against Trivy's built-in rule set, which focuses on well-known provider patterns. Plain text passwords, custom token formats, or values that do not match any published rule will not appear. Make sure you picked **Full scan (vulnerabilities + secrets)** from the shield-icon menu; a plain vulnerability scan does not walk the filesystem.
|
||||
</Accordion>
|
||||
<Accordion title="Compose misconfiguration scan returns 404">
|
||||
The scanner needs a Compose file in the stack directory. If the stack was created outside Sencho and the working directory does not contain a file named `compose.yml`, `compose.yaml`, `docker-compose.yml`, or `docker-compose.yaml`, the scan returns 404. Name the file accordingly or keep the stack under Sencho's managed compose directory.
|
||||
</Accordion>
|
||||
<Accordion title="SARIF download returns 409 Scan not complete">
|
||||
SARIF export requires a completed scan. If a scan failed, timed out, or is still running, the button downloads nothing and the server returns a 409. Trigger a fresh scan from the Resources Hub or the stack page, wait for the drawer to populate, then export again.
|
||||
</Accordion>
|
||||
<Accordion title="SARIF download is missing findings I see in the drawer">
|
||||
Each finding type (vulnerabilities, secrets, misconfigs) is capped at 5000 rows in the SARIF export. When any type trips the cap, the SARIF run includes `properties.truncated: true` along with the original counts, and the server logs a warning. The drawer paginates beyond 5000 so it shows everything, but the export is bounded for memory safety. If you need every row, narrow the scope (per-stack scan, per-image scan) before exporting.
|
||||
</Accordion>
|
||||
<Accordion title="Acknowledge button is missing on a misconfig finding">
|
||||
The button is admin-only and hidden on replica nodes. Replicas read acknowledgements from the control via fleet sync; mutations happen on the control. If you are an admin on the control and still do not see the button, the row is likely already acknowledged. Look for the dimmed strikethrough rendering and hover for the recorded reason.
|
||||
</Accordion>
|
||||
<Accordion title="Findings reappeared after deleting a CVE suppression or misconfig acknowledgement">
|
||||
Suppressions and acknowledgements are applied at read time and never modify stored finding rows. Removing one immediately resurfaces the underlying finding everywhere it appears (drawer, compare sheet, badge counts, SARIF export). The behavior is intentional: deleting an acknowledgement is meant to revert the operator decision, not to mask history.
|
||||
</Accordion>
|
||||
<Accordion title="Outbound traffic to ghcr.io and aquasecurity from the scanner host">
|
||||
Trivy fetches its CVE and secret-rule database from public registries on first scan and refreshes periodically; that egress is required for vulnerability scanning to work. Sencho does not emit telemetry of its own. If your environment forbids egress, see Trivy's [air-gapped scanning guide](https://trivy.dev/latest/docs/advanced/air-gap/) for how to pre-seed the database and run scans with `--offline-scan`.
|
||||
</Accordion>
|
||||
<Accordion title="Compose stack scan returns 409 Already scanning this stack">
|
||||
Sencho deduplicates concurrent scans of the same stack so two simultaneous calls cannot double-process the result. Wait for the in-flight scan to finish (its row appears in Scan history with status `in_progress` and flips to `completed` or `failed` shortly after) and trigger again.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
@@ -4,62 +4,68 @@ description: Trigger stack actions from CI/CD pipelines via HTTP webhooks with H
|
||||
---
|
||||
|
||||
<Note>
|
||||
Webhooks require a **Sencho Admiral** license.
|
||||
Webhooks require a **Skipper** or **Admiral** license. Creating, editing, and deleting webhooks is restricted to admin users; non-admins on a paid tier can view the list but cannot manage it.
|
||||
</Note>
|
||||
|
||||
Sencho webhooks let external systems trigger stack actions over HTTP. The typical use case: your CI pipeline builds a new image, then calls a Sencho webhook to deploy the updated stack, with no manual intervention required.
|
||||
|
||||
## How it works
|
||||
|
||||
1. You create a webhook in **Settings > Webhooks**, targeting a specific stack and action
|
||||
2. Sencho generates a unique secret for HMAC-SHA256 signature validation
|
||||
3. Your CI/CD system sends a `POST` request to the trigger URL with the correct signature
|
||||
4. Sencho validates the signature and executes the action asynchronously
|
||||
1. You create a webhook in **Settings → Alerts → Webhooks**, targeting a specific stack and action.
|
||||
2. Sencho generates a unique secret for HMAC-SHA256 signature validation.
|
||||
3. Your CI/CD system sends a `POST` request to the trigger URL with the correct signature.
|
||||
4. Sencho validates the signature and executes the action asynchronously.
|
||||
|
||||
## Creating a webhook
|
||||
|
||||
Go to **Settings > Webhooks** and click **Create Webhook**. Fill in:
|
||||
Open **Settings → Alerts → Webhooks** and click **Create webhook**. The page is part of the global settings and is reachable while the **Local** node is active in the node switcher; webhooks created from this page execute against the local Sencho instance.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/webhooks/webhooks-create-form.png" alt="Webhook creation form with Name, Stack, and Action fields" />
|
||||
<img src="/images/webhooks/webhooks-create-form.png" alt="New webhook form with Name, Stack, Node, and Action fields. The Action dropdown is open showing Deploy (down + up), Restart, Stop, Start, Pull and Update, and Git source sync." />
|
||||
</Frame>
|
||||
|
||||
| Field | Description |
|
||||
|-------|-------------|
|
||||
| **Name** | A display name (e.g. "Deploy on push") |
|
||||
| **Stack** | The target stack to act on |
|
||||
| **Action** | One of the actions listed below |
|
||||
| **Name** | A display name shown in the configured-webhooks list and execution history. |
|
||||
| **Stack** | The target stack to act on. Picked from the stacks on the active node. |
|
||||
| **Node** | Read-only. Reflects the currently active node; webhook execution is pinned to that node. |
|
||||
| **Action** | One of the actions listed below. |
|
||||
|
||||
After creation, Sencho shows the webhook secret **once** in a green banner. Copy it immediately; it cannot be retrieved later. The webhook list only shows the last four characters of the secret.
|
||||
After you click **Create**, the new webhook's secret appears once in a green callout titled **Webhook created. Copy your secret now.** Copy it before dismissing the callout; once you navigate away, it cannot be retrieved. The configured-webhooks list always shows the secret in masked form (`********` plus the last four characters).
|
||||
|
||||
<Frame>
|
||||
<img src="/images/webhooks/webhooks-secret-reveal.png" alt="Secret reveal banner after creating a webhook, showing the full secret and a copy button" />
|
||||
<img src="/images/webhooks/webhooks-secret-reveal.png" alt="Green success callout after creating a webhook, showing the full secret in a monospace block, a Copy button, and a Dismiss button. Below the callout, the configured-webhooks list shows the new webhook card with its trigger URL and masked secret." />
|
||||
</Frame>
|
||||
|
||||
### Available actions
|
||||
|
||||
| Action | What it does |
|
||||
|--------|-------------|
|
||||
| **Deploy (down + up)** | Stops and redeploys the entire stack |
|
||||
| **Restart** | Restarts the stack's running containers |
|
||||
| **Stop** | Stops the stack |
|
||||
| **Start** | Starts a stopped stack |
|
||||
| **Pull & Update** | Pulls the latest images and recreates changed containers |
|
||||
| **Deploy (down + up)** | Stops and redeploys the entire stack. |
|
||||
| **Restart** | Restarts the stack's running containers. |
|
||||
| **Stop** | Stops the stack. |
|
||||
| **Start** | Starts a stopped stack. |
|
||||
| **Pull & Update** | Pulls the latest images and recreates changed containers. |
|
||||
| **Git source sync** | Pulls the latest commit from the stack's Git source, then deploys. Requires a [Git source](/features/git-sources) attached to the stack. |
|
||||
|
||||
## Managing webhooks
|
||||
|
||||
Each webhook appears as a card showing the webhook name, an action badge, a stack badge, the trigger URL with a copy button, and the masked secret.
|
||||
Each configured webhook renders as a card showing the webhook name, an **action** badge, a **stack** badge, a **node** badge, an enable toggle, and a delete button. Below the header are a labeled **Trigger URL** row with a copy button, a labeled **Secret** row showing the masked secret, and a **Recent executions** disclosure that expands inline.
|
||||
|
||||
The masthead above the section reports total **Webhooks** and how many are **Enabled**.
|
||||
|
||||
<Frame>
|
||||
<img src="/images/webhooks/webhooks-settings.png" alt="Webhooks settings showing a configured webhook with trigger URL, masked secret, and controls" />
|
||||
<img src="/images/webhooks/webhooks-settings.png" alt="Configured webhooks section with two webhook cards. The first card has Recent executions expanded, showing several deploy entries with timestamps and durations and one restart entry that failed with an error message. The second card is collapsed. The masthead at the top shows WEBHOOKS 2 and ENABLED 2." />
|
||||
</Frame>
|
||||
|
||||
From the webhook card you can:
|
||||
From a webhook card you can:
|
||||
|
||||
- **Toggle** the webhook on or off with the switch
|
||||
- **Delete** the webhook with the trash icon
|
||||
- **Copy** the trigger URL with the clipboard button
|
||||
- **View execution history** by clicking **Recent executions**
|
||||
- Toggle the webhook on or off with the **On / Off** switch.
|
||||
- Delete the webhook with the trash icon.
|
||||
- Copy the trigger URL to the clipboard with the copy button.
|
||||
- Expand **Recent executions** to view the inline history.
|
||||
|
||||
The card has no in-place edit affordance: to change the name, stack, or action, delete the webhook and create a new one. The signing secret is bound to a single webhook for its lifetime; deleting and recreating rotates it.
|
||||
|
||||
## Triggering a webhook
|
||||
|
||||
@@ -71,7 +77,7 @@ Content-Type: application/json
|
||||
X-Webhook-Signature: sha256={hmac}
|
||||
```
|
||||
|
||||
The signature is computed as `HMAC-SHA256(request_body, webhook_secret)` and sent with a `sha256=` prefix.
|
||||
The signature is computed as `HMAC-SHA256(raw_request_body, webhook_secret)` over the exact bytes of the request body (UTF-8 if you send JSON) and sent with a `sha256=` prefix. Sencho uses a constant-time comparison.
|
||||
|
||||
### Example with curl
|
||||
|
||||
@@ -88,21 +94,22 @@ curl -X POST https://your-sencho.example.com/api/webhooks/3/trigger \
|
||||
|
||||
### Overriding the action
|
||||
|
||||
By default, the webhook executes the action configured at creation time. You can override it per-request by including an `action` field in the body:
|
||||
By default, the webhook executes the action configured at creation time. You can override it per request by including an `action` field in the body:
|
||||
|
||||
```json
|
||||
{ "action": "restart" }
|
||||
```
|
||||
|
||||
Valid override values are: `deploy`, `restart`, `stop`, `start`, `pull`.
|
||||
Valid override values are: `deploy`, `restart`, `stop`, `start`, `pull`, `git-pull`. An unknown value causes the execution to fail and appear with an error in **Recent executions**.
|
||||
|
||||
### Response
|
||||
### Responses
|
||||
|
||||
A successful trigger returns `202 Accepted` immediately. The action executes asynchronously in the background.
|
||||
|
||||
```json
|
||||
{ "message": "Webhook accepted", "action": "deploy" }
|
||||
```
|
||||
| Status | Body | Meaning |
|
||||
|--------|------|---------|
|
||||
| `202 Accepted` | `{ "message": "Webhook accepted", "action": "deploy" }` | Signature is valid; the action is now running asynchronously. A 202 means accepted, not finished. |
|
||||
| `401 Unauthorized` | `{ "error": "Missing X-Webhook-Signature header" }` | The request omitted the signature header. |
|
||||
| `401 Unauthorized` | `{ "error": "Invalid signature" }` | The provided signature did not match the expected HMAC. |
|
||||
| `404 Not Found` | `{ "error": "Webhook not found or disabled" }` | The webhook id is unknown or its enable toggle is off. |
|
||||
|
||||
## CI/CD integration examples
|
||||
|
||||
@@ -135,20 +142,48 @@ deploy:
|
||||
|
||||
## Execution history
|
||||
|
||||
Each webhook tracks its recent executions. Click **Recent executions** on any webhook card to expand the history, which shows:
|
||||
Click **Recent executions** on any webhook card to expand its inline history. Each row shows:
|
||||
|
||||
- **Status** icon (green check for success, red X for failure)
|
||||
- **Action** that was performed
|
||||
- **Timestamp** of the execution
|
||||
- **Duration** in seconds
|
||||
- **Error message** if the execution failed
|
||||
- A **status** icon (green check for success, red X for failure).
|
||||
- The **action** that ran (the configured action, or the override if one was supplied).
|
||||
- A localized **timestamp**.
|
||||
- The **duration** in seconds when available.
|
||||
- A truncated **error** message on failures; hover the row to see the full message in a tooltip.
|
||||
|
||||
Sencho retains the last 100 executions per webhook and surfaces the 20 most recent in the inline view.
|
||||
|
||||
## Security
|
||||
|
||||
- Secrets are shown only once at creation. The webhook list displays only the last four characters.
|
||||
- Trigger endpoints do not require a session cookie, but are protected by HMAC signature validation.
|
||||
- Each webhook targets a single stack. There is no way to execute arbitrary commands through a webhook.
|
||||
- Webhook secrets are returned only once, at the moment of creation. Every subsequent view of the webhook (in the list, in the API, anywhere) shows the masked form `********` plus the last four characters.
|
||||
- Trigger endpoints are public in the sense that they do not require a Sencho session cookie, but every request must carry a valid HMAC-SHA256 signature; signatures are compared with a constant-time check.
|
||||
- Each webhook targets a single stack and runs a single action, optionally narrowed by the supplied override. There is no path to execute arbitrary commands through a webhook.
|
||||
- Webhook execution is pinned to the node selected at creation time.
|
||||
|
||||
<Note>
|
||||
Webhooks only operate on the **local node**. You cannot trigger actions on remote nodes through webhooks.
|
||||
</Note>
|
||||
## Troubleshooting
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="I don't see Webhooks in Settings.">
|
||||
The page requires a **Skipper** or **Admiral** license. If you are on a paid tier but the node switcher in the top-left shows a remote node, switch to **Local** to reveal the page.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="My trigger returns '401 Invalid signature'.">
|
||||
The signature must be computed over the **exact raw bytes** of the request body. Common causes:
|
||||
|
||||
- The CI step JSON-encodes the body before signing but sends a different body to curl (or vice versa). Sign the same string you send.
|
||||
- The shell adds a trailing newline (use `echo -n` rather than `echo`).
|
||||
- The header is missing the `sha256=` prefix; the full value is `sha256={hex}`.
|
||||
- The webhook secret used to sign does not match the secret stored in Sencho. The secret is shown once at creation; if you lost it, delete the webhook and create a new one.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="My trigger returns '404 Webhook not found or disabled'.">
|
||||
The id in the URL is wrong, or the webhook's enable toggle is off. Copy the trigger URL from the card and check the **On / Off** switch on the webhook in **Settings → Alerts → Webhooks**.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="The action returns 202 but the stack does not change.">
|
||||
A 202 means the action was accepted, not that it produced a visible change. **Pull & Update** is a no-op if the registry has no newer image for any service in the stack; **Start** is a no-op on an already-running stack. Check the stack's **Activity** sheet to confirm what actually ran, and the webhook's **Recent executions** for the action's status and duration.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="I cannot create a Git source sync webhook for this stack.">
|
||||
The **Git source sync** action can only be selected for a stack that already has a [Git source](/features/git-sources) attached. If you try to create one without it, the form returns a validation error. Attach a Git source to the stack in its Compose editor, then create the webhook.
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
@@ -70,7 +70,7 @@ Stacks on different nodes need to call each other often enough that doing it by
|
||||
|
||||
- The [dashboard](/features/dashboard) hits you with a status line, a unified gauge strip, a stack-health table sorted by load, and a fleet heartbeat.
|
||||
- [Global observability](/features/global-observability) streams logs from every container on every node into one searchable view.
|
||||
- [Alerts](/features/alerts-notifications) on CPU, memory, disk, restart count, and similar thresholds, [routed](/features/notification-routing) to Discord, Slack, or any webhook by rules.
|
||||
- [Alerts](/features/alerts-notifications) on CPU, memory, disk, restart count, and similar thresholds, [routed](/features/alerts-notifications#notification-routing) to Discord, Slack, or any webhook by rules.
|
||||
- [Audit log](/features/audit-log) keeps a searchable trail of every mutating action with actor and node attribution.
|
||||
|
||||
### Resources you can actually clean up
|
||||
|
||||
|
After Width: | Height: | Size: 114 KiB |
|
Before Width: | Height: | Size: 151 KiB After Width: | Height: | Size: 146 KiB |
|
After Width: | Height: | Size: 175 KiB |
|
After Width: | Height: | Size: 203 KiB |
|
After Width: | Height: | Size: 117 KiB |
|
Before Width: | Height: | Size: 52 KiB After Width: | Height: | Size: 196 KiB |
|
Before Width: | Height: | Size: 115 KiB After Width: | Height: | Size: 112 KiB |
|
After Width: | Height: | Size: 106 KiB |
|
After Width: | Height: | Size: 102 KiB |
|
After Width: | Height: | Size: 123 KiB |
|
Before Width: | Height: | Size: 38 KiB After Width: | Height: | Size: 56 KiB |
|
After Width: | Height: | Size: 170 KiB |
|
Before Width: | Height: | Size: 147 KiB After Width: | Height: | Size: 162 KiB |
|
Before Width: | Height: | Size: 151 KiB After Width: | Height: | Size: 174 KiB |
|
Before Width: | Height: | Size: 128 KiB After Width: | Height: | Size: 148 KiB |
|
Before Width: | Height: | Size: 21 KiB After Width: | Height: | Size: 20 KiB |
|
After Width: | Height: | Size: 4.8 KiB |
|
After Width: | Height: | Size: 8.5 KiB |
|
After Width: | Height: | Size: 36 KiB |
|
Before Width: | Height: | Size: 116 KiB After Width: | Height: | Size: 132 KiB |
|
Before Width: | Height: | Size: 133 KiB After Width: | Height: | Size: 29 KiB |
|
After Width: | Height: | Size: 8.1 KiB |
|
Before Width: | Height: | Size: 97 KiB After Width: | Height: | Size: 32 KiB |
|
Before Width: | Height: | Size: 106 KiB After Width: | Height: | Size: 35 KiB |
|
Before Width: | Height: | Size: 118 KiB After Width: | Height: | Size: 63 KiB |
|
Before Width: | Height: | Size: 148 KiB After Width: | Height: | Size: 56 KiB |
|
After Width: | Height: | Size: 50 KiB |
|
After Width: | Height: | Size: 18 KiB |
|
After Width: | Height: | Size: 20 KiB |
|
Before Width: | Height: | Size: 129 KiB After Width: | Height: | Size: 168 KiB |
|
After Width: | Height: | Size: 36 KiB |
|
After Width: | Height: | Size: 150 KiB |
|
Before Width: | Height: | Size: 200 KiB After Width: | Height: | Size: 133 KiB |
|
Before Width: | Height: | Size: 166 KiB |
|
After Width: | Height: | Size: 36 KiB |
|
After Width: | Height: | Size: 50 KiB |
|
After Width: | Height: | Size: 24 KiB |
|
Before Width: | Height: | Size: 163 KiB After Width: | Height: | Size: 29 KiB |
|
Before Width: | Height: | Size: 141 KiB After Width: | Height: | Size: 173 KiB |
|
After Width: | Height: | Size: 13 KiB |
|
After Width: | Height: | Size: 46 KiB |
|
After Width: | Height: | Size: 19 KiB |
|
After Width: | Height: | Size: 30 KiB |
|
After Width: | Height: | Size: 33 KiB |
|
After Width: | Height: | Size: 28 KiB |
|
After Width: | Height: | Size: 30 KiB |
|
After Width: | Height: | Size: 34 KiB |
|
After Width: | Height: | Size: 7.4 KiB |
|
After Width: | Height: | Size: 27 KiB |
|
After Width: | Height: | Size: 168 KiB |
|
After Width: | Height: | Size: 26 KiB |
|
After Width: | Height: | Size: 114 KiB |
|
After Width: | Height: | Size: 39 KiB |
|
Before Width: | Height: | Size: 61 KiB After Width: | Height: | Size: 189 KiB |
|
Before Width: | Height: | Size: 65 KiB After Width: | Height: | Size: 175 KiB |
|
Before Width: | Height: | Size: 14 KiB After Width: | Height: | Size: 102 KiB |
|
Before Width: | Height: | Size: 87 KiB After Width: | Height: | Size: 132 KiB |
|
After Width: | Height: | Size: 152 KiB |
|
After Width: | Height: | Size: 78 KiB |
|
After Width: | Height: | Size: 143 KiB |
|
After Width: | Height: | Size: 146 KiB |
|
After Width: | Height: | Size: 103 KiB |