Commit Graph

3 Commits

Author SHA1 Message Date
Anso e2fc3a58a0 fix: reduce prune estimate work and add managed-scope timeout (#1768)
* fix: reduce prune estimate work and add managed-scope timeout

estimateSystemReclaim previously called getDiskUsageClassified,
which walks the full classified-resources pipeline (6+ Docker
API calls and filesystem I/O) under the 8 s timeout, but only
reads the three reclaimable* fields that getDiskUsage() (a single
docker.df() call) already provides. Switch to getDiskUsage() so
the timeout actually bounds the work the comment describes.

Additionally, the managed-scope estimateManagedReclaim path had
no timeout on either the remote route or the fleet local path.
Wrap both call sites in withTimeout so a slow daemon surfaces
the actionable 'Docker daemon is busy' message within 8 s instead
of hanging until the hub's 15 s fetch abort fires.

* fix: skip getStacks() for all scope in prune estimate route

The remote handler unconditionally walked the compose directory before
starting the 8 s estimate timer, but for 'all' scope the knownStackNames
parameter is now unused (estimateSystemReclaim uses only docker system df).
Mirror the fleet route's conditional so the walk only happens for managed
scope, where estimateManagedReclaim genuinely needs stack names.

Found during QA: on a Pilot node with real tunnel latency, this unbounded
walk added latency outside the timeout budget.

* fix: raise prune estimate budget to 12s for large image stores

docker.df() cost scales with image-store size: measured ~7.4s on a
34GB / 96-image store, alone nearly exhausting the previous 8s budget
before tunnel transport overhead. A healthy Pilot node could flip to
'Docker daemon is busy' at idle load.

Raise PRUNE_ESTIMATE_TIMEOUT_MS and FLEET_DF_TIMEOUT_MS to 12s, which
sits strictly below the hub's 15s AbortSignal.timeout on the fleet
estimate fetch, keeping the remote 503 the actionable failure. The
MonitorService janitor keeps its own 8s budget for destructive paths.

Found in QA pass 2: single-target estimate failed at ~8.05s on a node
where docker.df() alone takes ~7.4s.
2026-08-05 13:01:14 -04:00
Anso 44d6078241 feat(fleet): show itemized prune plans (#1734)
* feat(fleet): itemize prune review plans

Build and display fingerprint-bound prune candidates for every reviewed
fleet node. Preflight all node plans before mutation and preserve detailed
removed, skipped, failed, and partial outcomes.

Add safe resource metadata projection, managed ownership attribution,
runtime contract validation, transport parity coverage, and operator docs.

Closes #1724

* fix(security): harden stack path lookup

Use a Map for Compose working-directory ownership resolution so untrusted
path strings cannot become object property writes.

* fix(fleet): harden prune execution safeguards
2026-07-29 14:30:18 -04:00
Anso a51547a158 fix(monitor): decouple janitor disk-usage check from 30s cycle (F-6) (#1164)
* fix(monitor): decouple janitor disk-usage check from 30s cycle (F-6)

`docker system df` (called by the MonitorService janitor check) can take
30+ seconds on Docker Desktop with many volumes. Running it on the 30s
evaluate cycle compounded with the per-container stats fan-out and pushed
the cycle to 140s+, blocking subsequent monitoring work.

This change:

- Moves the janitor disk-usage check into its own 15-minute cycle with
  a tight 8s timeout. A circuit breaker opens after 3 consecutive
  timeouts (60-minute cooldown) so a sick daemon stops pinning Dockerode
  sockets every tick. The first janitor tick is deferred 45 seconds past
  boot to avoid head-of-line collision with the initial monitor cycle's
  stats fan-out.
- Adds a paired 8s `withTimeout` wrap to the admin prune-estimate routes
  (`/api/system/prune/estimate` and the dry-run path of
  `/api/system/prune/system`) so a slow df does not hang the admin tab.
  Both routes respond 503 with code `docker_df_slow` on timeout.
- Factors `withTimeout` and `TimeoutError` into `utils/withTimeout.ts`
  so the route layer does not have to import from a service module.
- Adds 10 unit tests covering the decoupling guardrail, breaker
  open/close, cooldown, threshold gate, the 100 MB reclaimable floor,
  re-entrancy, recovery logging, non-timeout error handling, and the
  full timer-cleanup contract of `stop()`.
- Adds 4 integration tests for the prune routes covering the 503
  timeout response, the success path, and the non-timeout 5xx path.

* fix(fleet,monitor): extend F-6 timeout to fleet prune routes; close breaker-recovery log gap

Codex audit findings on PR #1164:

Major. The fleet routes that fan out prune-estimate work on local nodes
(`POST /api/fleet/labels/fleet-prune` dry-run path and
`POST /api/fleet/prune/estimate`) called `estimateSystemReclaim` without
a timeout, so a slow local Docker daemon could still hang the fleet
admin tab even though the system-maintenance routes were already
bounded. Wrap both call sites with the shared `withTimeout(..., 8s)`
and surface a "Docker daemon is busy" message via the per-target and
per-node error channels the routes already used for other failures.
The destructive (non-dry-run) prune path stays unwrapped because it
calls `pruneSystem` / `pruneManagedOnly`, not `df`.

Minor. The janitor circuit breaker zeroed `janitorConsecutiveTimeouts`
when it opened, so a successful call after a full breaker-open cooldown
slipped past the `if (counter > 0)` recovery-log branch and never
emitted `[Monitor] Janitor disk-usage check recovered`. The operator
observability signal was missing exactly when it mattered most.
Extend the predicate to also trip on `janitorBreakerUntil > 0` (which
stays set to its past timestamp after cooldown until the next success
clears it), so recovery logs symmetrically for both partial-failure
and post-breaker recovery paths. Added a dedicated test.

Three new integration tests cover the fleet routes (timeout, success,
and the estimate endpoint's per-node unreachable shape).
2026-05-23 00:04:05 -04:00