* feat(auto-heal): restart crashed containers and harden the heal loop
Auto-Heal now restarts containers that crash (non-zero exit) and stay down
past the policy threshold, in addition to those that fail their Docker
healthcheck. Crash detection reuses the container event classifier so a
container that exits cleanly or that an operator stopped is never restarted;
only classified crashes set the heal signal.
Also hardens the existing loop:
- A paid controlling instance refreshes proxied remotes' entitlement on a
background interval so a remote node's policies keep evaluating between
operator visits instead of lapsing a few minutes after the sheet was last
opened. A node that stays unreachable surfaces a warning.
- Overlapping policies (all-services plus a service-specific one) restart a
given container at most once per evaluation pass, so the hourly cap holds.
- A failed restart now counts toward the cooldown and hourly cap, so a broken
setup is retried on the cooldown interval rather than every poll.
- Diagnostic logging behind developer mode for evaluation, heal decisions,
timing, and lease refresh.
* docs(auto-heal): document crash healing and refresh troubleshooting
Cover the two heal conditions (unhealthy and crashed), note that clean exits
and operator stops are never restarted and that crash healing acts on crashes
observed while Sencho is running, and update the troubleshooting and tab
visibility entries accordingly.
* fix(auto-heal): close stale crash-signal race and harden lease refresh
A crash signal could outlive the crash it described. The exit classifier is
deferred 500ms, so an immediate restart could let it stamp the crash marker
after the container was already running, and a later clean or operator-initiated
exit did not clear it; the next poll could then restart a container that had
exited cleanly. Now a clean or intentional exit always clears the marker, a die
that a start has superseded is not stamped, and the die's own time is captured at
arrival rather than at the deferred classification so the supersede check is
accurate.
Also:
- Crash state survives the event service's idle-prune window, so crash healing
works for any configured threshold rather than only short ones.
- An exited or dead container is matched before any health-text parsing, so it
can never fall into the healthcheck path.
- A remote with no reachable proxy target counts toward the lease-refresh
failure warning instead of being silently skipped.
* feat(db): add auto_heal_policies and auto_heal_history schema and CRUD
Adds two new SQLite tables (auto_heal_policies, auto_heal_history) to
DatabaseService.initSchema() and exposes CRUD methods: getAutoHealPolicies,
getAutoHealPolicy, addAutoHealPolicy, updateAutoHealPolicy,
deleteAutoHealPolicy, recordAutoHealHistory, getAutoHealHistory,
incrementConsecutiveFailures, resetConsecutiveFailures, setPolicyEnabled.
Also adds AutoHealPolicy and AutoHealHistoryEntry TypeScript interfaces.
* feat(events): track health-status duration and expose state accessors
- Add healthStatus and unhealthySince fields to InternalContainerState
- onHealthStatus now records unhealthySince timestamp on first transition
to unhealthy, and clears it when the container recovers or restarts
- onStart resets both fields so a restarted container begins from 'starting'
- Add listContainerStates() and getContainerState() public accessors for
use by the upcoming AutoHealService evaluator
* fix(auto-heal): key allowlist in updateAutoHealPolicy, cascade delete, extract ContainerHealthSnapshot
* feat: add AutoHealService evaluator singleton
Polls every 30 s, matches containers to enabled policies via Compose
labels, and restarts containers that have been unhealthy beyond the
configured threshold. Enforces cooldown, per-hour rate cap, and
recent-user-action suppression; auto-disables policies after repeated
consecutive failures. Also adds DockerEventManager.getService() accessor
required by the evaluator.
* fix(auto-heal): prune stale restartTimestamps, guard undefined policy id
- Prune restartTimestamps entries for containers no longer running after
each container list fetch, preventing unbounded map growth from dead
container IDs.
- Guard against policies with undefined id at the start of the per-policy
loop; warn and skip rather than proceed with a non-null assertion.
- Extract handleAutoDisable private helper to bring executeHeal under 30
lines and isolate the auto-disable side-effect sequence.
- Move ContainerInfo type to module scope.
* feat: add auto-heal API routes and wire AutoHealService lifecycle
Registers five REST endpoints under /api/auto-heal/policies (list, create,
patch, delete, history) with requirePaid + requireAdmin guards and Zod
validation. Wires AutoHealService.start()/stop() into the server startup
and graceful-shutdown blocks alongside MonitorService.
* test: add AutoHealService and DatabaseService auto-heal unit tests
- 15 unit tests for AutoHealService.shouldHeal covering all decision branches
(healthy state, duration threshold, user-action suppression, cooldown,
rate limiting, and correct skipReason values)
- 13 integration tests for DatabaseService auto-heal CRUD: policy round-trip,
stack-name filter, partial update, cascade delete, history ordering/limit,
consecutive failure counters, and setPolicyEnabled toggle
* fix: log AutoHealService shutdown errors consistently
* fix(api): requireAdmin-first guard order and try/catch on auto-heal routes
* feat(ui): add StackAutoHealSheet component
* feat(ui): add Auto-Heal context menu item to EditorLayout
* fix(ui): StackAutoHealSheet label, token, a11y, and useEffect fixes
- Rename 'All services in stack' to 'All services' in combobox options and placeholder
- Replace text-green-600 with text-success design token in actionColorClass
- Add htmlFor/id pairs to all four numeric form inputs for accessibility
- Inline fetch logic into useEffect, removing stale closure risk and eslint-disable comment
- Remove now-unused fetchPolicies and fetchServices standalone functions
- Update 'Auto-disable after' label to 'Auto-disable after (failures)' for clarity
- Add toast.error in policy fetch failure path; services fetch silently skips as before
* docs: add auto-heal-policies feature documentation
* test(e2e): add auto-heal policies CRUD spec
* fix(docs): correct auto-heal-policies nav position in docs.json