mirror of
https://github.com/Studio-Saelix/sencho.git
synced 2026-08-13 12:17:34 +00:00
9ff678a7bb
* docs(introduction): refresh for the redesigned UI and replace screenshots Bring the Getting Started Introduction page in line with the current product: - Add the Security top-level view to the navigation list and a dedicated Security section with a new screenshot. - Correct the Fleet tab names (Snapshots, Status, Map, Deployments, Routing, Federation, Actions, Secrets). - Split Settings out from security and list the current nine setting groups (Security graduated to its own view). - Refine the navigation paragraph so role, tier, and local-vs-remote context read accurately. Replace all four existing screenshots (Home, stack workspace, Fleet, Resources) with fresh captures of the redesigned UI and add a Security overview screenshot. * docs(configuration): document advanced env vars and clarify deployment vs runtime config Add an Advanced environment variables section (TRIVY_BIN, SENCHO_MESH_SUBNET, GITSOURCE_MAX_CLONE_BYTES, SENCHO_PUBLIC_URL, SENCHO_COMPOSE_STALL_TIMEOUT_MS) and reframe the intro to separate deployment-time configuration from the runtime settings that live in the in-app Settings Hub. Cross-link the pilot-agent variables to the Pilot Agent page instead of duplicating them. * docs(sso): refresh SSO Setup Guide and SSO & LDAP reference for the redesigned UI Refresh both SSO documentation pages against the current product and the redesigned settings UI. - Correct the navigation path to Settings -> Access -> SSO on both pages. - Fix the "Require 2FA on SSO sign-in" toggle location to Settings -> Personal -> Account. - Describe the login-page experience (the Local / LDAP toggle and the branded OIDC buttons under the "Or continue with" divider) and the SSO panel masthead (SCOPE, PROVIDERS, ENABLED). - Replace all six SSO screenshots with fresh captures of the redesigned UI. * docs(features): refresh the Features Overview page for the redesigned UI Rewrite docs/features/overview.mdx to mirror the current Features navigation grouping (Stacks, Deployment, Resources, Observability, Fleet, Automation, Security & Identity) and add the recently shipped capabilities surfaced in the redesign: Stack Dossier, Drift Detection, Compose Doctor, Compose Networking, Environment & secrets guardrails, Storage portability, Health-Gated Updates, Fleet Dossier, and the dedicated Security page. Correct stale claims (the file explorer now gates writes on stack edit permission, not an admin role; downloads are a read action; bulk label assign now spans nodes) and standardize the tier callouts so partly paid features read as "Admiral adds X". Replace the three pre-redesign screenshots and add a Security overview banner, all captured from a populated fleet. * docs(features): refresh the Appearance page for the redesigned UI Add fresh screenshots and a troubleshooting section to the Appearance page, verified against the live product. - Add four screenshots: the Theme card (live preview, mode, accent, and fine-tune sliders), the top-bar quick switcher, the Typography card, and the Display card. - Refresh the Density screenshot used by the Settings reference page. - State that the quick switcher also covers text size, and that the contrast, border, and glow sliders stay in Settings. - Add a Troubleshooting accordion covering per-browser persistence, resets to defaults, cross-operator scope, and the quick-switcher versus full-Settings split. * docs(introduction): refresh screenshots and correct stale content * docs(reference): refresh the Settings Reference page for the redesigned UI Replace all seven stale screenshots with fresh 1920x1080 captures. Add five new screenshots for the sections that previously had none. Content changes: - Sidebar table: rename Infrastructure "Fleet Mesh" entry to "Fleet"; add "Image update checks" to the Automation group list - Fleet section: rename heading to match registry label; add the Documentation snapshots subsection (snapshot_documentation toggle) - Container Alerts: add screenshot - Image update checks: add the full section (Registry checks table, scheduling mode, interval presets, cron expression support) - Stacks / Deploy Guardrails: add screenshot - Recovery: add the full section (System health snapshot, Environment preflight checks, Safe actions, Command-line recovery table) * docs(sso): refresh screenshots for SSO quickstart and feature pages * docs: refresh Features Overview screenshots and content Replace all 4 hero screenshots with fresh 1920x1080 production captures. Correct security posture state names (Action needed / Monitoring / Secure), add the Policies tab to the Security section tab list, mention the Simple mode in Scheduled operations, and update all alt text to match the new screenshots. * docs: refresh Appearance page screenshots and correct quick-switcher scope Replace all four Appearance screenshots with fresh production captures. Fix the quick-switcher control list: remove fonts (not present in the popover), add visual style and readability which are. Add Log chip color to the Display section. Update all screenshot alt text to match new captures. * docs: refresh stack management page with current UI and anatomy tabs * docs: fix convert-tab-error screenshot with fully visible error toast * docs: convert troubleshooting section to AccordionGroup format * docs(quickstart): refresh screenshots and align dashboard description Replace all three first-boot and dashboard screenshots with current UI. Add Security to the top navigation list, update gauge and Stack health descriptions to reflect sparklines and column detail, and align Configuration Status wording with the Introduction page. * docs(editor): rewrite anatomy panel, replace all screenshots - Correct the anatomy panel tab inventory: the panel has eight tabs (Anatomy, Activity, Dossier, Drift always; Environment, Networking, Doctor, Storage when the node advertises the matching capability), not three as previously documented - Add table describing all eight tabs with capability gates and links to dedicated feature pages - Add anatomy-tabs.png screenshot showing the scrollable tab row - Note the Doctor severity dot (red for blocker, amber for high-risk) - Remove the stale Markdown-export subsection; Dossier and Activity are now covered in the tab table - Replace all six stale screenshots with fresh 1920x1080 captures - Replace the compose diff preview screenshot * docs(files): refresh Files & Volumes screenshots and fix context-menu alt text Replace all 9 stale screenshots on the Files & Volumes page with fresh captures from the production node. Fix three alt-text strings that did not match the live UI: removed hardcoded octal value 644, and added the Duplicate, Copy to, and Move to entries missing from the context-menu alt text. * docs: rewrite Stack Activity page with full event categories and fresh screenshots Expands the event category table from 5 to 10 entries to cover drift detected, drift resolved, update started, health gate passed, and health gate failed. Adds a live-disconnected-state section, a background-actor attribution table, and a corrected troubleshooting accordion covering the WebSocket reconnect case. Replaces both stale screenshots with fresh 1920x1080 captures from the production node. * docs(drift): rewrite drift detection page with screenshots and full coverage Full rewrite of the Drift Detection feature page. Adds two previously undocumented finding types (network-undeclared, network-missing), expands the temporal section to distinguish the raw-file hash from the parsed-model hash, documents the two-layer spatial-engine and ledger architecture, explains when the ledger is reconciled (post-deploy vs manual re-check vs tab open), adds Activity timeline integration note, introduces a Limitations section (no background scanner, port-range caveat, history cap, advisory-only enforcement), expands Troubleshooting from five entries to seven using the AccordionGroup convention, and adds four production screenshots. * docs(drift): use CardGroup for Related section * docs(dossier): rewrite Stack Dossier page with full feature coverage * docs(networking): rewrite Compose Networking page with full feature coverage * docs(doctor): rewrite Compose Doctor with full 30-rule reference, screenshots, and cross-links * docs(networking): add production screenshots and correct alt text Adds 7 production screenshots for all sections of the Compose Networking page and updates the four placeholder alt texts written before screenshots were taken to match what the actual images show (arr-net external badge, swag service with 443/tcp and 80/tcp, single-service exposure intent row). Also adds the full-panel overview image at the top of the page. * docs(environment-guardrails): rewrite with project env file, env file status, and screenshots * docs(storage): rewrite Storage Portability page with screenshots and full coverage Rewrites compose-storage.mdx from a 61-line sketch into a complete reference page. Key additions: Where to find it section with screenshot, full storage inventory section documenting all mount type/access/status chips and the Linux owner display, expanded portability verdict section with per-reason detail and edge-case caveats (read-only binds, symlink escapes, anonymous volume risks), snapshot coverage section with admin scope and remote-node behavior, Findings in Doctor cross-reference, and six troubleshooting accordions covering tab visibility, bind status, external named volumes, render errors, and snapshot coverage states. Adds two production screenshots: storage-tab.png and storage-node-bound.png. * docs(stack-labels): rewrite with accurate permissions, capability gate, dry run, live preview, and color conflict docs * docs: rewrite Stack Sidebar page with accurate feature coverage Rewrites the Stack Sidebar documentation page to match the current UI. Key changes: - Fix branding header description (shows logo + version, not just version) - Fix bulk mode icon description (stacked-rows, not square) - Add cross-node search section (fan-out behavior, Other nodes section, unreachable-node warnings, click-to-switch navigation) - Update Labels submenu description (inline New label creation, Manage labels link) - Note that Delete only appears when the user has delete permission - Remove the auto-update implication from Schedule task description - Rewrite the Activity ticker section with the full 6-state priority cascade table; remove the non-existent IDLE state; correct pulsing-dot behavior - Replace all 7 stale screenshots with fresh production screenshots - Add new sidebar-cross-node-search.png screenshot * docs(atomic-deployments): refresh screenshot and document project env files, rollback readiness, and recovery actions * docs(atomic-deployments): fix rollback permission visibility and banner string accuracy The Rollback menu entry is hidden by the frontend when the user lacks stack:deploy; it never appears and does not 403. Fixed the step-4 narrative and troubleshooting accordion to match. The rollback-failure banner emitted by ComposeService is '=== Rollback failed. Manual intervention may be required ===' (period, capital M). Fixed both occurrences in the page. Updated the Settings navigation path from the nonexistent 'Roles & Access' to the real 'Access'. * docs(deploy-progress): rewrite with health gate, inline style, and 9 fresh screenshots Add health gate section covering all four states (observing, passed, failed, unknown) with exact UI banner text and the configurable observation window. Expand the inline style section with full band content, 4s auto-dismiss, and pill handoff. Add Scanning as a supported entry point. Replace all 6 existing screenshots and add 3 new ones (modal-health-gate, inline-banner, setting-style). Add two health gate troubleshooting accordions. Add Related CardGroup linking to health-gated-updates, stack-activity, deploy-enforcement, and atomic-deployments. * docs(health-gated-updates): refresh screenshots and correct signal row order and label * docs(deploy-enforcement): rewrite with fleet replication, honor suppressions location, scan-failed dialog state, and fresh screenshots Adds the Fleet policy replication section covering control/replica behavior, Managed by control node banner, and Demote to control. Documents the exact location of the Honor suppressions toggle (bottom of Policies tab). Expands the block dialog section with the scan-failed row state. Updates all three screenshots to the current visual design. Restores the Admiral license note and corrects the policy-card scope description. * docs(app-store): rewrite with mobile layout, fresh screenshots, and registry admin note - Replace all 5 stale screenshots with 1920x1080 production captures - Add app-store-mobile.png showing the status masthead layout - Document mobile single-column layout in a new Mobile subsection - Note that the featured hero has its own Deploy button - Mark the category rail as desktop only with a cross-link to Mobile - Add admin-account requirement to the custom registry section - Add Related CardGroup linking vulnerability scanning, deploy progress, deploy enforcement, and resources
120 lines
12 KiB
Plaintext
120 lines
12 KiB
Plaintext
---
|
|
title: Health-Gated Updates
|
|
description: Check update readiness before applying, watch container health after the update lands, and know exactly what a rollback can restore.
|
|
---
|
|
|
|
Updating a stack is the highest-anxiety operation in a homelab: a pull-and-recreate can break containers, networking, env configuration, or storage assumptions, and `docker compose` alone gives you no answer to "is it safe to update this right now?". Sencho makes updates deliberate and reversible in three layers: a readiness check before the update, a health gate after it, and an always-honest rollback readiness report on the stack.
|
|
|
|
None of these layers blocks you. Readiness is advisory, the health gate is purely observational, and rollback is always an explicit action you confirm. The only thing that hard-blocks an update is a [deploy enforcement policy](/features/deploy-enforcement).
|
|
|
|
## Update readiness
|
|
|
|
When you trigger **Update** from the stack editor toolbar or the sidebar menu, Sencho first opens a readiness dialog with a single verdict and the signals behind it:
|
|
|
|
| Verdict | Meaning |
|
|
|---------|---------|
|
|
| Ready | Nothing stands out; the update can proceed. |
|
|
| Ready with warnings | Proceed, but read the warnings first. |
|
|
| Review required | Something needs a look: a high-risk preflight finding, an already-unhealthy container, a pending major version bump, or low disk. |
|
|
| Blocked | A blocker was found, such as a Compose Doctor blocker or a scan policy that will stop the update. |
|
|
| Unknown | Readiness could not be verified, for example when the node or Docker is unreachable. |
|
|
|
|
The verdict is computed from signals Sencho already tracks, so the check is fast and never re-runs anything heavy:
|
|
|
|
- **Compose Doctor**: the stored result of the last [preflight run](/features/compose-doctor). A blocker finding makes the verdict Blocked; high-risk findings ask for review. If Compose Doctor has never run, the dialog says so without dragging the verdict down.
|
|
- **Drift**: open [drift findings](/features/stack-drift) warn you that the running state has diverged, so the rollback target may not match what is running.
|
|
- **Current containers**: a container that is already unhealthy, restarting, or crashed before the update makes the result hard to evaluate; the dialog asks you to look first.
|
|
- **Healthcheck coverage**: how many services define healthchecks, which determines how thoroughly the post-update health gate can verify the result.
|
|
- **Pending update**: the [update preview](/features/auto-update-policies) classifies the pending change. A major version bump asks for review; patch updates and same-tag refreshes pass quietly.
|
|
- **Rollback backup**: whether a backup slot exists and how old it is. A missing backup is only a note, because the update itself creates a fresh one when it starts.
|
|
- **Node disk**: disk usage near or above the node's alert threshold warns you before a large pull fills the disk.
|
|
|
|
Admins also see whether a [fleet snapshot](/features/fleet-backups) covers this stack and can tick **Create a fleet snapshot before updating** to capture one as part of proceeding. If the snapshot fails, the update does not start.
|
|
|
|
Every verdict keeps the **Update now** button enabled. The dialog informs the decision; it does not make it for you.
|
|
|
|
<Frame>
|
|
<img src="/images/health-gated-updates/readiness-dialog.png" alt="Update readiness dialog showing the verdict chip, signal rows for Compose Doctor, Drift, Current containers, Healthcheck coverage, Pending update, Rollback backup and Node disk, the fleet snapshot row with its checkbox, and the Cancel and Update now buttons" />
|
|
</Frame>
|
|
|
|
## The post-update health gate
|
|
|
|
The 3-second crash probe from [atomic deployments](/features/atomic-deployments) catches instant failures. The health gate is the longer observational layer behind it: after a deploy or update succeeds, Sencho watches the stack's containers for an observation window (90 seconds by default) and records a verdict.
|
|
|
|
During the window, the gate checks every container that came up with the stack:
|
|
|
|
- all expected containers stay **running**,
|
|
- Docker healthchecks report **healthy** wherever a service defines one,
|
|
- no container **exits**, **disappears**, or enters a **restart loop**.
|
|
|
|
A clear failure ends the observation immediately. Passing requires the full window, and a healthcheck still in its start period when the window ends records the verdict as unknown rather than claiming success.
|
|
|
|
The deploy progress modal treats the gate verdict as the real result: while the gate observes, the status reads **Verifying health** rather than claiming success, and the auto-close waits. A passed gate shows the success state; a failed or unknown gate becomes the headline result with its reason. Closing the modal never stops the observation; the gate runs server-side and its result appears on the stack timeline either way, as **health gate passed** or **health gate failed** events alongside the **update started** marker.
|
|
|
|
The live verifying and recovery view is part of the deploy progress panel. If you have turned that panel off, the in-browser gate view does not appear, but the gate still runs on the node and records its verdict on the stack timeline.
|
|
|
|
<Frame>
|
|
<img src="/images/health-gated-updates/modal-verifying.png" alt="Deploy progress modal after an update succeeded, showing the Verifying health status and the health gate banner observing containers" />
|
|
</Frame>
|
|
|
|
<Frame>
|
|
<img src="/images/health-gated-updates/modal-gate-failed.png" alt="Deploy progress modal with the Health gate failed headline and the banner identifying the unhealthy container, with the hint that rollback options are available on the stack" />
|
|
</Frame>
|
|
|
|
When the gate fails, the stack page surfaces the same [recovery actions](/features/deploy-progress#recovery-actions) as a failed update: retry, restart, roll back when a backup exists, refresh the container state, or copy diagnostics. Rolling back is always your call; the gate never rolls anything back on its own.
|
|
|
|
The gate also runs for updates you did not click: scheduled image updates, webhook-triggered deploys and pulls, bulk updates, and Git source applies all record gate verdicts on the stack timeline. Rollbacks, App Store installs, and Sencho's own automation loops are deliberately not gated.
|
|
|
|
### Tuning or disabling the gate
|
|
|
|
Open **Settings > Infrastructure > Stacks > Deploy Guardrails** on the node you want to configure:
|
|
|
|
- **Observe health after updates** turns the gate on or off for that node. On by default; turning it off only stops the observation and its timeline events, never the update itself.
|
|
- **Observation window** sets how long containers are watched, from 15 to 600 seconds. The default is 90 seconds; raise it for stacks that take a while to settle, since a healthcheck still starting at the end of the window records an unknown verdict instead of a pass.
|
|
|
|
<Frame>
|
|
<img src="/images/health-gated-updates/settings.png" alt="Stacks settings page showing the Deploy Guardrails subsection, with the Observe health after updates toggle set to On and the Observation window field set to 90 seconds" />
|
|
</Frame>
|
|
|
|
## Rollback readiness
|
|
|
|
The Stack Dossier carries a **Rollback readiness** section that answers one question honestly: if this update goes wrong, what can a rollback actually restore? It reports an overall state of Ready, Partial, or Not ready, built from:
|
|
|
|
- **Previous compose file**: whether a backup slot exists and how old it is.
|
|
- **Previous env file**: whether the backup contains the stack's env file. Sencho lists how many variable names are covered; the values themselves are restored with the file and never displayed.
|
|
- **Previous image tag**: the known rollback target from the update preview. When the compose file uses a moving tag, restoring files alone does not revert the image, so the report names the exact tag you would pin to be precise.
|
|
- **Last successful deploy**: whether the backup reflects a configuration that actually deployed successfully.
|
|
- **Healthchecks**: whether a rollback can be verified beyond run state.
|
|
- **Application data**: always reported as not covered. Named volumes and bind-mounted data are not included in file backups; a rollback restores compose and env files only, and your application data keeps its current state. This row exists so the limit is stated where you decide, not discovered during an incident.
|
|
|
|
<Frame>
|
|
<img src="/images/health-gated-updates/dossier-rollback-readiness.png" alt="Rollback readiness section in the Stack Dossier showing the overall state chip and the six rows: Previous compose file, Previous env file, Previous image tag, Last successful deploy, Healthchecks, and the Application data row marked not covered" />
|
|
</Frame>
|
|
|
|
## Classified failures
|
|
|
|
When a deploy or update fails, Sencho classifies the failure from the compose output and shows the cause with a suggested next step in the recovery panel: an image pull failure, a missing environment variable, a host port conflict, a missing bind-mount path, a permission problem, a crashed container, a failed healthcheck, an unavailable dependency, an unreachable node or Docker daemon, or an invalid compose file. The classification also lands in **Copy details**, so a bug report carries the cause, not just the raw output.
|
|
|
|
## Troubleshooting
|
|
|
|
<AccordionGroup>
|
|
<Accordion title="The readiness dialog shows Unknown">
|
|
Unknown means one of the verdict-affecting signals could not be verified, most often because Docker or the node did not answer in time. The update path is never blocked by an unknown verdict; check the node's connection if it persists, and proceed when you are confident in the stack's state.
|
|
</Accordion>
|
|
<Accordion title="The health gate failed but my app seems fine">
|
|
The gate fails on the first clear problem it observes: a container exit, an unhealthy healthcheck, a restart loop, or a container disappearing mid-window. Check the named container's logs from the stack page. If the container legitimately restarts during startup (a migration step, for example), raise the observation window under **Settings > Infrastructure > Stacks > Deploy Guardrails** so the gate watches past the settle period.
|
|
</Accordion>
|
|
<Accordion title="The health gate result is unknown">
|
|
Unknown means the observation could not finish: Sencho restarted mid-window, a newer update superseded the observation, Docker became unreachable, a healthcheck was still starting when the window ended, or no containers appeared to observe. The stack timeline records the reason. Check the containers directly; an unknown verdict makes no claim either way.
|
|
</Accordion>
|
|
<Accordion title="The modal stays open saying the gate is observing">
|
|
The modal holds its auto-close while the gate observes so the verdict is not lost. You can close it at any time; the observation continues server-side and the verdict lands on the stack timeline. If the gate result repeatedly cannot be retrieved, the modal gives up with an unknown verdict instead of waiting forever.
|
|
</Accordion>
|
|
<Accordion title="Rollback readiness says my data is not covered">
|
|
That is by design and true for every stack: file backups cover compose and env files, never named volumes or bind-mounted data. For point-in-time copies of stack files across the fleet, use [fleet snapshots](/features/fleet-backups); for application data, use a backup tool appropriate to the workload (database dumps, volume backups) before risky updates.
|
|
</Accordion>
|
|
<Accordion title="Updates from the sidebar menu now show a dialog first">
|
|
The sidebar's per-stack **Update** action runs the same path as the editor toolbar, so it shows the same readiness dialog and deploy progress. One click on **Update now** proceeds. On nodes that do not advertise the capability, updates run directly without the dialog.
|
|
</Accordion>
|
|
</AccordionGroup>
|