Adds the full step-by-step content for the Configure Auto-Heal Policies
stub: an nginx+redis scenario stack, adding a service-scoped policy,
and a live verification that breaks a container's healthcheck,
confirms the policy restarts it, and recovers it.
Covers labeling a target node, authoring a stateless Blueprint,
walking through the create-then-approve rollout flow, verifying
from the Deployments tab and the audit log, and recovering from a
port-conflict deploy failure. Cross-links with Move a Blueprint
Deployment to a New Node in both directions.
Migrates a Blueprint-managed workload from one node to another using
pin and cordon, with the confirm-before-mutate rollout in between.
Corrects the published feature page's claim that pin requires the
global admin role; the code gates cordon and pin identically, scoped
to the target node.
Registers an OAuth client in a self-hosted identity provider (Keycloak
worked example), configures Sencho's Custom OIDC settings, tests the
connection, and verifies a real end-to-end login with auto-provisioning
from two independent surfaces.
* chore: bump brace-expansion and fast-uri via npm audit fix
Resolves GHSA-rgw5-rvv9-x895 (brace-expansion DoS via unbounded
intermediate arrays). Both transitive dev dependencies updated:
- brace-expansion 5.0.8 -> 5.0.9
- fast-uri 3.1.4 -> 3.1.5
* chore: also bump frontend deps via npm audit fix
Fixes brace-expansion and postcss in the frontend lockfile so
npm audit --audit-level=high passes on both packages.
* chore: bump ip-address transitive dep via npm audit fix
Resolves three new ip-address advisories (GHSA-mwp4-54f8-5fhr,
GHSA-4xrf-jv44-h6hh, GHSA-22jq-vg5j-6vgg) published between prior
push and CI run.
* feat: surface ZFS ARC reclaimable as dashboard context line
Add arcReclaimable to the HostMemory interface and MemoryWire shape so
the reclaimable ARC amount computed by readReclaimableArc() is exposed
through /api/system/stats and /api/fleet/overview. Show it as a context
line on the dashboard memory tile, matching the balloon pattern.
ARC continues to feed the gauge percentage as before; this is a
display-only addition for operator visibility.
* docs: clarify ARC line requires nonzero reclaimable, not just readable stats
* fix(ui): flatten single-container Update onto the service row
On multi-service stacks, put Update/Rebuild on the container card (left of
image source) when a service has one matching container, and keep the shared
header only for multi-replica services.
* fix(ui): show per-service Update only when an image update is confirmed
Registry services were always showing Update because eligibility checked
declaredImage/hasBuild only. Gate Update on a confirmed pending check so
the button clears after a successful recheck; keep Rebuild for build-backed
services.
* feat(fleet-secrets): graduate encrypted fleet-wide environment bundles to Community
* fix(fleet-secrets): update reachability test for Secrets community graduation
* fix(fleet-secrets): address review findings
* fix(fleet-secrets): add HTTP-level Community admin push/import tests and non-admin tab-hidden test
* feat: validate scheduled task target existence at creation time
Reject POST/PUT /api/scheduled-tasks with 400 when the target stack,
container, or node does not exist. Previously only structural format was
validated; a task targeting a deleted stack would return 201 and fail
forever at execution time with noisy error logs.
Node existence is now validated for every action that requires a node
(previously only scan/prune got this check). Stack and container
existence is validated on local nodes; remote nodes are skipped since
the check would require a proxy call (execution-time validation still
serves as the safety net there).
Permission checks run before existence checks, so unauthorized callers
receive 403 regardless of whether the target exists.
* fix: close node-existence oracle in scheduled task creation
Move validateActionNode (which contains a getNode DB lookup) after the
permission gate in both POST and PUT handlers, so unauthorized callers
always receive 403 regardless of whether a fleet or system target node
exists. Previously a viewer probing a nonexistent fleet node would get
400 ("node not found"), leaking node ID enumeration through the error
code difference.
The stack/container existence check (validateTargetExists) was already
correctly positioned after permission; this fix extends the same
discipline to the node-existence path.
* fix(fleet-secrets): invalidate stale preview when target inputs change
Clear the preview plan and push results whenever a target input
(selectedLabels, labelMode, stackName, or envFile) changes, so the
operator cannot review a diff computed from different inputs.
A version counter in a ref drops in-flight preview/push responses
when inputs change while a request is still pending.
* fix(fleet-secrets): always surface push results, add invalidation test
Remove the version guard from handlePush so a completed secret write
always reports its outcome to the operator, even if target inputs changed
while the push was in flight. The guard remains on handlePreview (read-only).
Add a component test covering: plan cleared on input change, stale
in-flight preview discarded, and push results always surfaced.
Add seedMfaUserWithToken helper that wraps seedMfaUser and returns a
signed JWT with the current token_version claim, removing the manual
sign-and-read-boilerplate from callers.
Add an integration test in the MFA reset describe block that exercises
the full invalidation cycle: a target user's pre-reset JWT is accepted
before reset and rejected with 401 'Session invalidated' after an admin
resets their MFA. The test also confirms a pre-minted admin JWT survives
the reset unchanged.
* feat(recovery): make rollback-recovery image lifecycle visible and controllable
GitHub discussion #1751 asked why Sencho creates sencho-rb/<id>/<service>:hold
images during automatic updates and how to clean them up. That surfaced a real
safety bug alongside the missing visibility: the manual single-image delete
route did not consult the held-image predicate every other deletion path
already honors, so a user could delete a rollback-protected image straight
through the Images tab and silently break automatic recovery for that update.
A short/truncated id also bypassed the predicate's full-id lookup.
Fixes:
- POST /images/delete now resolves the submitted id to its canonical form and
checks the unified held-image predicate before deleting, returning 409
IMAGE_HELD_FOR_ROLLBACK for a protected image.
- The Images tab no longer mislabels a protected image as plain "Unused"; a
fully-synthetic hold image is kept out of the generic inventory entirely and
surfaced instead in a new Resources -> Rollback tab, with an additive
"Rollback protected" badge for images that still carry a normal tag too.
New capability:
- Two settings (Deploy Guardrails): superseded-generation retention (days,
replaces a hardcoded 7) and a cap on retained generations per stack.
- A new Resources -> Rollback tab lists every generation (stack, short id,
state, retention) with an admin-gated manual release action, including
releasing the current generation with an explicit warning that automatic
rollback becomes unavailable until the next successful update. Release is
a single atomic, server-revalidated transition so a stale UI read can never
release a row that has since become ineligible.
Also consolidated three near-duplicate implementations of the held-image
predicate (two of which relied on a require() of a sibling .ts file that
silently failed to resolve under the test runner and was never actually
exercised by a real test before this change) into one shared module.
Known follow-up, not fixed here: an orphaned sencho-rb tag whose recovery row
no longer exists (DB restore, node re-add) is invisible in both the Images
and Rollback tabs with no UI path to reclaim it.
* fix(audit): add summary mapping for rollback generation release
* fix(security): sanitize prune target in log sinks and cover release RBAC
Closes two open js/log-injection findings on the system prune route by
applying the same inline sanitizeForLog barrier the rest of the file
already uses. The prune target is validated against an enum by
parsePruneTargets before reaching these sinks, so the findings were false
positives, but the barrier is cheap and removes the standing alerts on a
file this change already touches. Also wraps the generation id in the
release log line for consistency with the stack name beside it.
Adds coverage for gaps a QA pass identified:
- Release endpoint refuses a viewer and a deployer (Admin-only), leaving
the generation and its artifacts untouched.
- Viewer can still read the generations list, matching the sibling
Resources routes.
- The predicate the prune routes build reports full-stack rollback holds,
not just service-scoped ones, and re-reads per call so a hold taken
between plan and delete still gates the delete.
- After releasing the current generation, no rollback point is claimed
for the stack through any consumer of the current-generation lookup.
* feat(fleet): add Node details sheet to the node card kebab
Every Fleet node card now carries a "Node details" kebab item, open to
any role that can see the card (previously the kebab only rendered
for users with node-manage permissions, so plain viewers had none).
The sheet shows connectivity, live capacity, Compose workload,
version/capability compatibility, and governance info (labels, cordon
reason and date, default-node flag, Compose directory, registration
date) using data the Fleet page already fetches, plus one lazy call
to the existing node meta endpoint for capabilities. Wired into both
the desktop card and the mobile bespoke Fleet screen.
* fix(fleet): correct Node details sheet timestamp units and update-status fallback
QA against a live 3-node fleet found that last_successful_contact and
pilot_last_seen come back from the fleet-overview endpoint in Unix
seconds, but the sheet passed them straight into a milliseconds-only
formatter, rendering values like "20647d ago" instead of "just now".
Both are now converted before formatting.
The Compatibility section's update-status badge also fell through to
a confident "Up to date" whenever updateStatus was absent (e.g. on
mobile, which doesn't poll update status) instead of reflecting that
there was no data to back the claim; it now renders "Unknown" in that
case. The local node no longer shows a misleading "Last successful
contact: Never". Reworded the "read-only sheet" language in the docs
page to describe the sheet accurately, since the Governance section's
label picker stays editable for node managers by design.
acceptReverseLocal swapped a pre-connect error handler for permanent
close/error handlers inside the socket's 'connect' callback, briefly
leaving 'error' with the old handler removed and the new ones not yet
attached. A Node EventEmitter 'error' event with zero listeners
throws instead of being swallowed, and CI's runner timing hit this
window intermittently (backend-tests-pilot-tunnel-bridge-reverse-
route-events.test.ts's post-handshake-close test), surfacing as an
unhandled ECONNRESET exception even though every assertion passed.
Replace the swap with a single 'error' listener attached at socket
creation that branches on a connected flag, so the socket is never
without an error listener at any point in its lifecycle.
* feat: account for VM memory ballooning in host memory reporting
Extend hostMemory.ts with a readBalloonedMemory() function that parses the
Balloon: field from /proc/meminfo, following the same fail-open pattern as
the ZFS ARC integration. When a nonzero balloon is detected, effective
memory fields (effectiveUsed, effectiveFree, effectiveUsagePercent) are
computed and exposed through /api/system/stats and /api/fleet/overview.
All consumers that derive meaning from host memory now prefer effective
values when present: the dashboard gauge, Fleet card RAM bar, mobile
views, health verdict, health status bar stat tile, and host RAM alerts.
Backward compatible: missing /proc/meminfo or absent Balloon: line
preserves exact current behavior. Old remote nodes without the new fields
continue rendering normally.
* refactor: extract shared helpers for balloon memory wiring
Extract readCandidateFile() and logSelectedPath() in hostMemory.ts to
deduplicate ARC and balloon file-read logic. Add memoryToWire() to
centralize the optional-field spread used by /api/system/stats and
/api/fleet/overview. Add getNodeMemUsed()/getNodeMemTotal() helpers
in nodeUtils.ts for frontend byte-text consumers.
* fix: make desktop fleet masthead aggregate balloon-aware
The desktop fleet overview's memory aggregate in useFleetOverview.ts still
summed raw memory.used, while the mobile fleet aggregate and per-node cards
already used effective values. Update to use getNodeMemUsed/getNodeMemTotal
helpers.
* fix: revert balloon adjustment from alerting and health decisions
Ballooned memory is host-reclaimed (unlike ZFS ARC, which the guest can
reclaim on demand). The guest cannot get ballooned pages back until the
hypervisor deflates them, so treating ballooned memory as available for
alerting or health can mask real memory pressure.
Keep balloon parsing, wire fields, and the dashboard context line as
informational-only. The memory gauge, health verdict, and host RAM alerts
now use the standard ARC-adjusted working-set percentage regardless of
balloon. Updated configuration.mdx and dashboard.mdx to document that
balloon data is informational and does not influence alerting.
* fix: forward scoped node-admin permission for Settings writes through remote proxy
The proxy forwards only the user's global role via PROXY_ROLE_HEADER
to remote nodes. Scoped role assignments live only on the hub's
role_assignments table and are never transmitted, so a scoped Node
Admin could not save settings on their granted remote node through
the hub proxy.
Add a settings-write pre-authorization gate in runGatedProxy that
buffers the body, extracts required permission buckets from
SETTING_WRITE_PERMISSIONS, checks them hub-side, and elevates
PROXY_ROLE_HEADER to 'node-admin' when the scoped check passes.
The gate is fail-closed: empty or unparseable bodies require
checkNodeManage on the hub, matching the existing
requireSettingsWritePermission empty-keys branch.
Fixes the gate-parity gap where scoped node:manage worked locally
but not through the proxy for Settings writes.
* fix: remove unused UserRole import from remoteNodeProxy.ts
* test(self-update): poll instead of a fixed delay in triggerUpdate assertion
The 600ms sleep raced the route's 500ms post-response timer plus the
persist/watch work executeClaimedCommunityUpdate does before calling
triggerUpdate, leaving too little margin under CI's forked test pool.
Poll with vi.waitFor instead, matching the pattern already used
elsewhere in this suite.
The Labels section registry entry had no requiredPermission, leaving it
visible to every authenticated operator. The backend already enforces
stack:read on GET and stack:edit on write endpoints, and the frontend
component already hides edit controls behind can('stack:edit'). Adding
stack:read to the registry declares the contract explicitly.
All five built-in roles hold stack:read, so this has no observable
effect on current users. It becomes a functioning gate automatically
if a future role or scoped user type is introduced without the permission.
* feat: add target-aware RBAC authorization to Scheduled Operations
Replace the blanket requireAdmin gate on all 9 scheduled-tasks endpoints
with per-action permission checks derived from the centralized action
registry. Each scheduled action now declares the existing permission it
requires: stack lifecycle actions need stack:deploy on the target stack,
node-wide operations need node:manage on the target node, prune stays
admin-only via system:settings, and snapshot requires unscoped
node:manage.
Key changes:
- Backend registry: add permission field and resolveTaskPermissionScope
- Routes: replace requireAdmin with requireTaskPermission, filter GET
listing by permission, add two-phase PUT check
- Scheduler: revalidate creator permission at execution time, auto-disable
on revocation (TaskAuthorizationError)
- Database: new creator_user_id column with migration and backfill
- Frontend: canScheduleAction/canScheduleAny helpers, reachability gate
via scheduledOpsAccessible, stack context menu uses canDeploy
not isAdmin, permissions field on all ScheduledActionDefinitions
- Expose checkPermissionForSubject for in-process callers (scheduler)
No new PermissionAction values are added. Uses the existing stack:deploy,
node:manage, and system:settings matrix. Scoped Admiral grants work on the
exact (nodeId, stackName) target per SEN-438.
* test: add RBAC coverage for scheduled operations authorization
* test: update frontend tests for scheduled-ops RBAC gate changes
- useStackMenuItems: gate Schedule task on canDeploy not isAdmin;
add test for canDeploy=true, non-admin case
- buildNavigationModel: default scheduledOpsAccessible to true
(default ctx is admin, who can always schedule)
- useViewNavigationState: add stack:deploy to admin can() mocks
so canScheduleAny resolves correctly
* fix: keep checkPermission unchanged, add checkPermissionForSubject standalone
The earlier refactoring that made checkPermission delegate to
checkPermissionForSubject changed the call order of effectiveTier(req)
relative to the admin bypass, which subtly broke the Community tier
clamping in the audit-log route. Keep checkPermission byte-identical
to the original and expose checkPermissionForSubject as a standalone
function used only by the scheduler revalidation path.
* fix: prefix unused selectorType parameter in resolveTaskPermissionScope
* fix: address audit findings — existence oracle, revalidation test, action filtering
Three corrections from the independent PR audit:
RC-1 (action filtering): Wire canScheduleActionAnywhere into the action
picker in ScheduledOperationsView so actions the user can never schedule
(prune without system:settings, node:manage actions without a scoped
grant) are filtered from the picker entirely. Add canScheduleAction
check on the Create button against the currently selected target, so
the submit button is disabled when the caller cannot schedule the
chosen action on the chosen target.
RC-2 (existence oracle): Six by-ID endpoints (GET /:id, DELETE /:id,
PATCH /:id/toggle, POST /:id/run, GET /:id/runs/export, GET /:id/runs)
now return a uniform 404 when permission is denied on an existing task,
so an unauthorized caller cannot distinguish "task does not exist" from
"task exists but you are not authorized." Added
requireTaskExistsPermission helper for the 404 variant; POST create and
PUT merged-scope checks keep requireTaskPermission (403).
RC-3 (revalidation test): Fixed the orphan-creator test to assert the
post-execution task state (auto-disabled, explicit error message) rather
than wrapping executeTask in a try/catch with a no-op else branch, since
executeTask catches TaskAuthorizationError internally and returns
normally.
Added canScheduleActionAnywhere helper and AuthContext mock to
ScheduledOperationsView tests (38/38 pass).
* fix: remove unused TaskAuthorizationError import from test
* fix: use 404 on PUT phase-1 unauthorized-access check
The PUT two-phase check's first phase (ownership verification)
now uses requireTaskExistsPermission (404) instead of a manual
403, consistent with the six other by-ID endpoints. This prevents
a caller from probing task-ID existence via the PUT route.
* fix: address QA findings — reorder prechecks, surface permission reason
Two live-confirmed fixes from the 3-node fleet QA pass:
Finding #4 (offline-node error ordering): Swap the order of the node-
reachability precheck and the creator-permission revalidation in
SchedulerService.executeTask. Authorization now runs first, so a
revoked grant is always surfaced as an explicit auto-disable with a
clear error message, even when the target node is offline. Previously
the reachability check ran first, hiding the revocation behind a
misleading "target node is offline" error and leaving the task enabled
indefinitely while the node was down.
Finding #7 (unexplained Save dead-end): When the Create/Update button
is disabled because canScheduleAction denies the selected action-target
combination, a muted text line now appears below the button:
"You do not have permission to schedule this action on the selected
target." This gives scoped users meaningful feedback instead of a
silently disabled button with no explanation.
* chore: remove redundant !!formName guards
The earlier short-circuit conditions in isSaveDisabled and
saveDisabledReason already guarantee formName is truthy by the
time the canSaveWithCurrentTarget check runs. The !!formName
guard is a no-op — flagged by GitHub code-quality as 'Useless
conditional: This negation always evaluates to true.'
* fix(proxy): forward scoped stack evidence for alerts, auto-heal, and node-wide image refresh
Extend the remote proxy scoped-evidence mechanism beyond /stacks/*
routes. Three new gates in runGatedProxy:
- Alerts POST: reuse the already-buffered body from the existing
isAlertCreateRoute block, extract stack_name, check hub-side
scoped permission, and forward SCOPED_STACK_AUTH_EVIDENCE headers.
- Auto-heal POST: same pattern with new body buffering and encoding
rejection (no pre-existing buffering exists for this route).
- Node-wide image refresh: elevate PROXY_ROLE_HEADER to node-admin
when the user has a scoped node:manage grant on the target node,
matching the Settings pre-auth gate pattern.
Also extend classifyStackApiPath to recognize
/image-updates/refresh/:stackName as a named-stack route (stack:deploy),
ready for when PR #1743 adds the per-stack refresh endpoint.
Explicitly excluded: ID-based routes (DELETE /alerts/:id,
PATCH/DELETE /auto-heal/policies/:id) where the hub cannot resolve
remote-owned IDs to stack names; GET routes where stack:read is
globally granted to every role; and POST /auto-update/execute where
multi-stack/wildcard targets need a different evidence format.
* chore(proxy): add RBAC diagnostic logging to scoped permission path
Add developer_mode-gated diagnostic logs to checkPermission to expose
which check is failing when a scoped user is denied: effective tier,
DB query parameters, and node-scoped assignment lookups.
* chore(rbac): log effective tier and license status when scoped checks are blocked
Add an always-visible console.warn in checkPermission when the
effective tier prevents scoped role-assignment lookups, logging
both the resolved tier and the raw license_status DB value. This
surfaces the failure reason in container logs without requiring
developer_mode, so QA can diagnose why scoped users are denied.
* chore(rbac): sanitize scoped-tier log values to satisfy CodeQL log-injection check
* fix: gate alerts and auto-heal routes on stack:edit/stack:read permissions
Replace requireAdmin with requirePermission across backend/src/routes/alerts.ts
and backend/src/routes/autoHeal.ts, mirroring the stack:read/stack:edit model
already used by stacks, blueprints, git sources, and settings. Adds the
previously-missing permission gate on the auto-heal history route, and adds
ownership-aware deletion for alerts via a new DatabaseService.getStackAlert(id)
lookup.
* fix: gate image-update fleet, per-stack refresh, and auto-update execute on RBAC permissions
Replace requireAdmin with requirePermission/checkPermission across
backend/src/routes/imageUpdates.ts (imageUpdatesRouter and autoUpdateRouter),
mirroring the permission-aware model already used by alerts and auto-heal.
GET /fleet drops its admin gate to match the auth-only read model shared with
GET / and /detail. POST /fleet/refresh now requires node:manage. A new route,
POST /refresh/:stackName, lets a caller with stack:deploy on that stack trigger
a per-stack recheck, distinct from the node-wide POST /refresh. The auto-update
executor now pre-checks stack:deploy across every resolved target before any
work starts, so a denied stack in a bulk request fails the whole call instead
of partially executing; the "*" wildcard additionally requires global
stack:deploy up front since it expands to every stack on the node, including
the empty case where a per-stack check would otherwise have nothing to gate.
* fix: evaluate permission before checks-enabled state in auto-update execute
The checks-enabled short-circuit in autoUpdateRouter POST /execute ran before
target parsing and before any permission check, so a node with image-update
checks disabled returned 200 to any authenticated caller regardless of
stack:deploy grants. Move the checks-enabled check to run after the resolved
stackNames have cleared requireExactStacks, so permission is always evaluated
first.
Add coverage: a denied role still gets 403 PERMISSION_DENIED (not the
disabled-checks 200) while checks are disabled node-wide, and a scoped-only
user whose stack:deploy grant covers every stack on the node is still denied
target="*" (the wildcard requires global stack:deploy, per the earlier fix),
proving that tradeoff against a real on-disk stack rather than the always-
empty fresh test instance.
* fix: gate alerts, auto-heal, and image-update controls on frontend permission checks
Match the backend RBAC gates for alerts, auto-heal, and per-stack image
updates with matching frontend checks, replacing raw isAdmin/node:manage
gates with scoped can() calls:
- Alerts/Auto-Heal menu items and their keyboard shortcuts now gate on
stack:read (canViewMonitor), including the window-level keyboard
shortcut handler that previously bypassed the menu item gate entirely.
- Check updates now gates on stack:deploy (previously node:manage) and
calls the new per-stack POST /image-updates/refresh/:stackName
endpoint instead of the node-wide refresh. Since the endpoint runs the
recheck synchronously and returns the result directly, the old
node-wide /status polling loop is removed in favor of handling the
response inline.
- StackAlertSheet's alert and auto-heal policy mutation controls gate on
stack:edit instead of isAdmin.
- The Fleet Image Updates refresh button (mobile and desktop) gates on
node:manage, hidden rather than disabled to match the existing
convention for node:manage-gated affordances.
* fix: cover the stack:edit deny path for StackAlertSheet gates
The useAuth mock in StackAlertSheet.test.tsx returned can: () => true
unconditionally, so canEditAlerts, canEditAutoHeal, and PolicyRow's
canEdit prop were never exercised with a denial. Make the mock
per-test-controllable (matching the vi.fn() pattern already used in
NodeCard.test.tsx) and add one deny-path test per tab asserting the
mutation controls are absent while reads stay visible.
Also adds an aria-label to the alert row's delete button so the deny
test can assert on its absence, matching the aria-label convention
PolicyRow's own toggle/delete controls already use.
* fix: surface accurate warnings and loading feedback on stack update checks
checkUpdatesForStack ignored the backend's StackRecheckResult outcome
and always showed a success toast, even when verification failed or
an update is still present. It also gave no feedback while the
multi-second per-image registry probe was in flight.
Add a loading toast on request start, and branch the result toast on
outcome/warning instead of unconditional success. The backend reuses
its post-update reconciliation copy for this pre-update discovery
check, so the two generic "update command completed" strings are
replaced with accurate pre-update wording; a genuine stack-specific
warning (e.g. a compose render failure) is still shown as-is.
Also update docs/features/rbac.mdx: stack:edit now covers alert and
auto-heal management, stack:deploy covers per-stack image-update
checks, and the Deployer role description reflects both.
* fix: add per-stack cooldown rate limit for image-update recheck route
The per-stack POST /refresh/:stackName route bypassed the existing
node-wide manual-refresh cooldown. A caller with stack:deploy could
hammer the registry with unbounded concurrent recheck calls.
Add tryMarkStackRecheck in ImageUpdateService, sharing the same
2-minute cooldown window, keyed per (nodeId, stackName). The route
handler returns 429 when denied. The mark is written synchronously
before the first await so concurrent calls on the same tick are blocked.
Partial-auth (mfa_pending) and enroll-only (pilot_enroll) tokens skipped
the admin check on /ws because any set scope was treated as already gated.
Reject those scopes in the shared upgrade pipeline, and deny unknown scopes
on the generic path while allowing api_token, console_session, and pilot_tunnel.
* feat(rbac): make Settings authorization permission-aware
Align Settings visibility and mutations with the existing permission matrix so Node Admin can edit node-scoped operational settings while system and credential surfaces stay Admin-protected.
* fix(rbac): tighten settings permission buckets and tests
Collapse settings key permission maps into one source of truth, and cover mixed PATCH atomicity plus image-update enabled writes.
* fix(rbac): tighten Settings scoped grants and CI assertions
Empty settings PATCH fails closed, node:manage is scoped to the active
node, system-only Settings stay hidden without system:settings, and
Check updates / webhooks mutate gates follow the permission matrix.
* fix(rbac): defer Settings section fallback until authz is ready
Keep deep links to permission-gated sections (e.g. license) intact while
can() is still fail-closed during permission metadata load.
* feat: expose Community audit log via system:audit navigation
Gate the Audit view on the system:audit permission instead of paid tier,
so Community admins can open the existing 14-day recent-activity window.
Export, anomaly flags, and stats remain Admiral-only.
* test: clarify synthetic Community admin mock lacks system:audit
Document that mockCommunityAdmin is a gate-isolation helper, not the
real Admin permission matrix where system:audit is always present.
* feat(rbac): make Settings authorization permission-aware
Align Settings visibility and mutations with the existing permission matrix so Node Admin can edit node-scoped operational settings while system and credential surfaces stay Admin-protected.
* fix(rbac): tighten settings permission buckets and tests
Collapse settings key permission maps into one source of truth, and cover mixed PATCH atomicity plus image-update enabled writes.
* fix(rbac): tighten Settings scoped grants and CI assertions
Empty settings PATCH fails closed, node:manage is scoped to the active
node, system-only Settings stay hidden without system:settings, and
Check updates / webhooks mutate gates follow the permission matrix.
* fix(rbac): defer Settings section fallback until authz is ready
Keep deep links to permission-gated sections (e.g. license) intact while
can() is still fail-closed during permission metadata load.
* docs(settings): clarify Notifications channels vs routing authz
Channels use node:manage via /api/agents; routing and mute stay Admin-only.
* feat(rbac): make stack-scoped grants node-specific
Qualify stack role assignments as (nodeId, stackName), migrate legacy rows to the default node, and forward bound multi-action evidence on Proxy/Pilot hops so scoped users keep least-privilege remote access without shipping the full grant table.
* fix: mirror scoped-stack-auth-evidence capability to frontend, sanitize node id in role assignment log
Backend added the scoped-stack-auth-evidence capability without the
matching frontend entry, failing the capability parity test. The role
assignment log also interpolated the node id without sanitizeForLog,
unlike the rest of the line.
* fix(rbac): honor node-wide scopes and fix proxied DELETE cleanup
Node-scoped grants now authorize that role's stack actions on the same node in the backend resolver, frontend can(), and remote evidence. Proxied DELETE cleanup uses the gate-stashed route because pathRewrite mutates req.path before proxyRes. Add proxy integration coverage and drop the stale scoped-permissions screenshot.
* fix(rbac): preserve node-qualified grants during repair
CVE-2026-56852 (norm.Iter infinite loop on crafted input) affects the
transitive x/text dependency in both the docker CLI and docker-compose
binaries bundled into the image.
* feat(fleet): reapply Compose configuration without a version update
Add a distinct Fleet Reapply configuration path so Compose-managed nodes can recreate Sencho from the current on-disk project when already up to date, without pulling or rewriting the image reference.
* fix(fleet): confirm remote reapply and close concurrent tracker race
Require confirmation for remote compose reapply, and lock dispatch before the remote POST so a second request cannot overwrite a successful in-flight tracker.
* fix(ui): icon-only Reapply control so Up to date badge can breathe
Collapse the Node updates Reapply label into a tooltip so the status pill no longer wraps in the Status column.
* feat(editor): Save & Reapply self-stack via fleet compose reapply (#1726)
* feat(editor): Save & Reapply self-stack via fleet compose reapply
Eligible admins can apply on-disk Compose edits to Sencho's own stack from the editor using the same confirm, dispatch, and reconnect path as Fleet Node Updates.
* fix(editor): gate Save & Reapply label to self-stack only
Ordinary stacks were labeled Save & Reapply whenever the node was
reapply-eligible. Require the selected file to be the self-stack for the
toolbar label and diff confirm CTA.
* fix(ui): move compose diff action label helper out of dialog module
Keep ComposeDiffPreviewDialog component-only so react-refresh Fast Refresh
lint passes after the Save and reapply stacked merge.
When every finding is acknowledged, the Doctor summary uses the success All Clear banner with distinct copy instead of the muted acknowledged chip, so operators can see at a glance that no active findings remain.
* fix(compose-doctor): resolve effective healthcheck coverage
Compose Doctor now classifies healthcheck coverage from the Compose model, running containers, and local images so image-provided HEALTHCHECKs are not false positives. Update Guard shares the same presence helper so test NONE is not treated as active.
* fix(compose-doctor): fix healthcheck project label and empty compose HC
Use the Compose project name for runtime container listing so stacks whose name: differs from the directory still get runtime evidence. Treat empty or timing-only healthcheck objects as absent rather than active.
* fix(compose-doctor): treat inherited healthcheck as All Clear note
Inherited image healthchecks no longer block All Clear; they surface under a notes section and cannot be acknowledged.
Intersect stack_alerts with FileSystemService.getStacks() so Configuration
Status (dashboard and fleet local row) counts only rules for stacks that
exist on that node, not orphaned names in the global table.
* feat(schedules): auto-update stacks by Stack Label
Add a reusable selector_type/selector_value on scheduled tasks so admins
can schedule image updates against live Stack Label membership across the
fleet or one node, reusing fleet label resolution and the existing
auto-update orchestrator.
* fix(image-updates): sanitize auto-update execute failure logs
Use a static format string and sanitizeForLog so CodeQL no longer
flags user-controlled stack names and error text in the execute catch.
* fix(ui): space Scope label from fleet/node segmented control
Match the Schedule row layout so the inline SegmentedControl no longer
sits flush against the Scope label.
* fix(ui): remove redundant wrapper around Scope segmented control
* feat: add node-scoped opt-out for image update detection
Operators who use an external update authority can disable Sencho registry
polling per node without losing explicit stack Update, pull, or redeploy.
* test: fix mocks and lint for image-update checks opt-out
Scheduler tests need isChecksEnabled on the ImageUpdateService mock, and the UpdatesSection older-node fixture must not leave an unused binding.
* fix: gate update-preview and recheck when detection is off
Anatomy was still calling stack update-preview (and contacting registries)
while checks were disabled. Short-circuit those routes and skip recheckStack
writes so disabled nodes stay quiet until detection is re-enabled.
* feat(auth): add SSO-only authentication mode
Let administrators disable interactive local password login when SSO is configured, with backend enforcement, activation safeguards, and host CLI recovery.
Closes#1709
* fix: resolve CI failures in auth mode PR
- Add useLicense mock to SSOSection test to prevent crash from
AuthenticationModePanel rendering without LicenseProvider
- Remove username from authMode console.log calls that CodeQL flags
as clear-text logging of sensitive information
* fix(auth): keep SSO-only on named disableSso and fail-closed login
Named provider disable no longer reverts authentication_mode. Login initializes localLoginEnabled false so a status fetch failure cannot reveal the password form. Center a single OIDC provider button on the login card.
* fix(auth): move SSO-only authentication mode from Admiral to Community tier
Security-hardening features belong on the Community tier per the existing
Community rebalance. The reporter of #1709 noted that disabling local
password login after configuring SSO is a basic security measure, not an
enterprise governance feature. LDAP provider configuration remains
Admiral-gated via requireTierForSsoProvider.
* fix(ui): keep SSO Active badge and ON toggle in sync
Provider cards mounted before config fetch finished with enabled:false, so a saved Active provider showed OFF until the local draft was resynced. Drive both the badge and TogglePill from the synced local config.
* feat(auth): auto-redirect to sole OIDC provider under SSO-only
When authentication mode is SSO only and exactly one OIDC provider is enabled (no LDAP), skip the login chooser and send the browser to that provider's authorize URL. Returning sso_error stays on the login page so the failure message remains visible.
* fix(ui): move oidcAutoRedirectUrl out of Login for fast refresh
Exporting the helper alongside the Login component tripped react-refresh/only-export-components and failed Frontend lint CI. Keep Login as a component-only module and colocate the helper with its unit tests under lib/.
* feat: live-refresh stack detail container and health state
Keep the open stack's container cards in sync with Docker via state-invalidate events and a visibility-aware poll, without reloading compose, env, or logs.
* fix: remove unused _ms parameter from visibilityInterval mock
Fixes the @typescript-eslint/no-unused-vars ESLint error in CI lint job.
* fix: stop stack detail live-refresh when leaving the editor
Gate poll and invalidate handling on editor visibility, refresh the
current selection after a mid-flight stack switch, and skip starting
visibilityInterval when the tab is already hidden.
* fix: avoid return in finally for stack detail live-refresh
Satisfy no-unsafe-finally by gating the trailing refresh with a positive
condition instead of early returns inside the finally block.