Commit Graph

8 Commits

Author SHA1 Message Date
Anso bb7c76ba46 fix(image-updates): normalize docker.io host aliases to the registry API host (#1706)
* fix(image-updates): normalize docker.io host aliases to the registry API host

parseImageRef kept a literal `docker.io` or `index.docker.io` host when the
user wrote an explicit registry prefix, so every downstream request hit the
marketing domain instead of registry-1.docker.io and failed with an
unhandled 3xx. Normalize both aliases before the library/ auto-prefix check
so explicit and implicit Docker Hub refs resolve identically.

* test(image-updates): pin docker.io alias parity and drop redundant coverage

Add the index.docker.io namespace-omitted parseImageRef case and a
buildRollbackTarget assertion for an explicit docker.io/ ref, and drop
the compareLocalToRemoteTag traefik-shape test that only duplicated
existing attestation-manifest coverage. Also clarify the canonicalRegistry
doc comment now that it and parseImageRef normalize to different forms.

* docs(image-updates): correct canonicalRegistry comment's call-path example

* test(image-updates): collapse docker.io alias cases into a single it.each
2026-07-26 01:56:12 -04:00
Anso 6688da97b1 fix(image-updates): match any local RepoDigest against the remote tag (#1695)
* fix(image-updates): match any local RepoDigest against the remote tag

Docker can list a stale multi-arch index digest ahead of the current one
on the same image. Selecting only the first RepoDigest caused false
same-tag rebuilds (for example redis:8.8.0) even when another digest
equaled the registry primary. Compare every matching candidate and keep
fail-closed behavior for empty, unknown-platform, and classification errors.

Fixes #1684

* fix(image-updates): surface digest verification failures to operators

Carry comparator errors into update-preview as check_error / verification_failed,
prefer failed checks over sticky has_update in Fleet and the sidebar, and keep
Update Guard from claiming no pending update when verification failed.

* fix(e2e): align sidebar truncation spec with check-failed precedence

StackRow now shows the check-failed icon over a stale update dot, but
this spec still asserted the old precedence and failed deterministically
in CI on every attempt.

* fix(fleet): treat verification-only previews as non-actionable

Fresh update-preview wins over sticky fleet booleans: disable Apply, exclude
from ready counts, and move verification-only stacks into the check-failures
advisory (including remote-labeled names).

* fix(fleet): move preview actionability helpers out of the view

Exporting non-components from AutoUpdateReadinessView tripped react-refresh
lint in CI. Keep the helpers in a shared lib module and drop an unused mock arg.

* fix(fleet): drop sticky cards when fresh preview clears the update

A successful no-update preview now removes the pending Fleet card instead of
leaving Apply enabled. Verification-only stacks still go to the advisory, and
empty-state copy no longer claims all-clear while checks remain unresolved.

* fix(image-updates): hold full-stack apply for review when another image fails verification

A confirmed update or rebuild on one image previously left the whole stack
fully actionable even when a different image in the same stack failed digest
verification: Anatomy claimed "safe to apply", Update Guard reported ready,
and Fleet's full-stack Apply stayed enabled, all while showing the
verification-failure text right next to those claims.

isActionableUpdatePreview now requires no verification failure anywhere in
the stack; a new isReviewRequiredUpdatePreview flags the mixed state so Fleet
still surfaces the card (not silently cleared) with Apply now disabled and a
"Review · unverified" badge. Anatomy's banner says "review required" instead
of a bump-based safety claim and withholds its Apply button. Update Guard's
pending-update signal downgrades from ok to attention. Per-service apply
(Fleet's per-image row) is deliberately left enabled since a service-scoped
update to the confirmed image does not touch the unverified one.

* fix(image-updates): treat rebuild_available symmetrically with has_update in Update Guard

updatePreviewSignal only downgraded to 'attention' inside the has_update
branch, so a rebuild-only stack (has_update false, rebuild_available true)
with a sibling verification failure fell through to the plain
verification-only 'unknown' branch and never mentioned the pending rebuild,
inconsistent with isReviewRequiredUpdatePreview on the frontend which treats
has_update and rebuild_available the same way.

Also adds desktop-card coverage for the mixed state (previously only the
mobile card was exercised) and locks in blocked/major-bump precedence over
the new review-required badge/banner in both Fleet and Anatomy.

* fix(image-updates): derive the mixed-verification review-hold from per-image detail, not the stack aggregate

has_update and check_error are independent per image: a tag-based update can
be confirmed via the registry's tag list even when that same image's own
digest comparison against the current tag errored (already covered by an
existing update-preview-service test). The stack-level verification_failed
and has_update flags can therefore both be true for the SAME single image,
which the previous review-hold treated identically to a genuinely different
image failing verification: Update Guard said "another image failed digest
verification" and Fleet told the user to "apply the confirmed service
individually" on a single-service stack where no such affordance exists.

isReviewRequiredUpdatePreview (and isActionableUpdatePreview) now walk the
preview's images to require a pure failure image (check_error, no has_update
of its own) alongside a genuinely different confirmed image or rebuild,
falling back to the old aggregate-only judgment when per-image detail is
unavailable. StackAnatomyPanel now imports the shared helper instead of
hand-rolling the same predicate, so Fleet and Anatomy cannot drift apart.
Backend updatePreviewSignal gets the same per-image treatment via a new
optional images parameter, threaded through from UpdateGuardService.

* fix(image-updates): fail closed on platform-unavailable indexes and legacy previews, allow anonymous tag listing

Four independent gaps from the same QA pass, all in the digest/tag
verification path this PR introduced or touches:

- compareLocalToRemoteTag now distinguishes a remote index with no descriptor
  at all for the local platform (including an empty or fully-filtered index)
  from a genuine mismatch: the former returns an error instead of reporting a
  speculative update. A node cannot pull a platform the index does not offer.
- selectLocalRepoDigests no longer falls back to a sole unrelated-repository
  RepoDigest when nothing matches the configured repo; comparing against a
  registry state that has nothing to do with the declared image risks a
  false update. Returns unresolved instead of guessing.
- isClearedUpdatePreview no longer treats a preview with verification_failed
  missing entirely (not merely false) as proof the stack is clean. The
  current backend always includes this field, so its absence identifies an
  older remote node's response, which cannot vouch for a clean result the
  way an explicit false can.
- listRegistryTags (and the underlying listRegistryTagsResult) no longer
  short-circuits to an empty list whenever no registry credentials are
  configured. getAuthToken already resolves anonymous tokens for public
  repositories; skipping it meant tag-based update detection silently never
  fired for any public image without a stored credential.

* fix(image-updates): correct platform-check overreach, add cache and advisory for prior fixes

Addresses code-review findings on the previous commit:

- The platform-unavailable check fired too eagerly: an index whose runnable
  descriptors legally omit platform (OCI-permitted, routed to exactDigests)
  has real pullable content, so it must not be confused with a genuinely
  empty or fully-filtered index. Now only errors when both platform-labeled
  descriptors and exactDigests are empty.
- listRegistryTags is now cached (15 min TTL): anonymous listing has no other
  rate limiting, and Fleet fans this out across every image on every reload.
- A legacy preview (kept rather than cleared) now also pushes a check-failure
  advisory entry explaining why, instead of rendering as an unexplained
  pending card.
- Corrected docstrings that described the old sole-unmatched-digest fallback
  and inverted how anonymous registry auth actually resolves.

Adds coverage for: nested-index and attestation-only-filtered platform
unavailability, a platform-less-but-populated index staying a match, the
unrelated-repo digest rejection wired through the real preview-computation
path (not just the registry-api unit), and legacy-preview interaction with
an actionable has_update:true.

* fix(image-updates): fail closed on mixed platform indexes, stop caching tag-list failures

Addresses a second review round on the previous commit, including an
empirically-verified regression:

- The exactDigests fallback was unconditional: an index mixing a
  platform-labeled descriptor for a DIFFERENT platform with an unlabeled leaf
  let that leaf stand in as this platform's content, reporting a speculative
  update for a genuinely incompatible platform. Now an unlabeled leaf is only
  trusted when it is the ONLY kind of descriptor in the index (nothing else
  claims a different platform); a mixed index errors instead.
- listRegistryTags was caching failed lookups for the full 15-minute TTL
  (a 429, an unreachable registry, or credentials not yet configured all
  looked identical to a real empty tag list). The fetcher now throws on
  failure so only a success is ever cached; CacheService's existing
  stale-on-error fallback still serves the last good list when one exists.
- The manual "Recheck" action now also drops the tag-list cache, so a newly
  published tag is visible immediately instead of waiting out the TTL.
- The legacy-preview advisory no longer fires when the same preview is
  already actionable on its own terms (a remote's own confirmed
  has_update/rebuild_available): pairing a "could not be checked" banner
  with an enabled Apply button next to it contradicted itself.
- Corrected docstrings and a self-contradictory inline comment left over
  from the prior fix.

New coverage: the exact mixed-index regression this round found and fixed,
cache hit/no-repeat-fetch and failure-not-cached behavior, and the
legacy-preview-plus-already-actionable non-contradiction.

* fix(image-updates): add digest_error unmasked field, reorder RiskBadge, fix actionability gates

Add digest_error to UpdatePreviewImage as an always-populated field that
is independent of check_status masking: a confirmed tag-based update on
the same image resolves check_status to 'ok' and nulls check_error, but
digest_error stays set since the image's current tag content was never
verified. Switch hasUnverifiedOtherImage to read digest_error.

Reorder RiskBadge precedence so reviewRequired is checked before
uncertain (both derive from check_error, so uncertain was unreachable).

Fix self-contradiction in test fixtures (digest_update:true +
check_error). Add masked-tag regression test with two-image fixture.
Simplify hasUnverifiedOtherImage per code review feedback.
2026-07-25 21:57:30 -04:00
Anso 0daddfde00 fix: reconcile sticky update indicators with Anatomy preview (#1698)
* fix: reconcile sticky update indicators with Anatomy preview

Sidebar, Updates filter, and Fleet treated retained partial/failed
scanner has_update as confirmed. Keep raw state for retention/notifications,
project confirmed-only to APIs, show distinct incomplete indicators, and
clear sticky rows only after an authoritative-negative preview.

Closes #1685

* test: align sidebar truncate E2E with failed-over-retained precedence

Purple update indicators are confirmed-only; hasUpdate with a failed
check correctly shows the failed trailing icon.

* fix: clear confirmed update rows on authoritative-negative preview

Address audit SF-1/SF-2/SF-3: observation-watermark clears for older
ok+has_update rows (DB + memory gens), Fleet checkability parity with
backend not_checkable, and Updates chip confirmed-only regressions.

* fix: tombstone equal-generation writers on preview clear

Advance the per-stack write generation when clearing at the observation
watermark so a scanner reserved before preview cannot recreate the row
after an authoritative-negative reconcile.

* fix: clear sticky updates with digest and tag preview parity

Share detection across scanner and preview, keep GET read-only with POST reconcile, gate Apply to digest and rebuild updates, and invalidate the hub fleet cache on clear.

* test: set digestUpdate on auto-update checkImage mocks

Scheduler and execute routes now gate Compose on digest drift; fixtures that expect an apply need digestUpdate so they exercise the update path.

* fix: clear unused lint errors on sticky update branch

Drop unused partial helper and fleet invalidate import; keep the CacheService inflight self-ref as let with an eslint exception so tsc stays green.

* fix: use inflight holder for CacheService prefer-const

Keep generation-aware ownership without a let self-reference that fights ESLint and tsc.
2026-07-25 15:42:19 -04:00
Anso 66ec4ebdd2 fix(image-updates): treat multi-arch child digests as up to date (#1641)
* fix(image-updates): treat multi-arch child digests as up to date

Floating tags like redis:8-alpine can store a platform child digest locally
while the registry tag resolves to the parent index. Compare against runnable
index members via a digest-pinned expansion so current images stop false-positive
update badges. Fixes #1630.

* fix(image-updates): preserve UTF-8 in capped GET and fail closed on nested indexes

Accumulate raw Buffer chunks before hashing or decoding so multibyte UTF-8
cannot corrupt content digests. Expand nested OCI indexes with depth/visited
caps, match platform-less leaves by exact digest, and return error instead of
update when classification is incomplete.

* fix: prefer-const lint error in registry-api test

* fix(image-updates): align multi-arch checkNode tests with 2-arg signature

After rebasing onto main (#1640), checkNode no longer takes nodeName. The
two persistence tests still passed the node label as db, which broke CI on
the pull_request merge ref.

* fix(image-updates): guard tag/repo components before registry URL construction, dismiss CodeQL false positive

Add defense-in-depth validation in probeManifestForRef that rejects tag
strings containing URL-injection characters (/ ? # \ null) and repo paths
with .. segments before they reach the outbound HTTPS request. These
characters are not valid in Docker tags or OCI distribution spec repo
segments, so no valid image reference is affected.

Exclude js/request-forgery on registry-api.ts via codeql-config.yml.
Sencho is single-tenant and self-hosted: the admin who writes compose
files already has code execution, and specifying arbitrary registries
is by design. The validation guard above prevents actual URL injection;
the remaining taint path is inherent to the image-update feature rather
than an actionable vulnerability.

Closes CodeQL alerts #531 and #532.
2026-07-17 08:47:22 -04:00
Anso f2b5c68d84 feat: add build-aware compose stack updates (#1561)
Detect services with build: in the update preview and run compose build --pull
plus pull --ignore-buildable when Update is triggered on those stacks, while
keeping the existing pull-only path for image-only stacks.
2026-07-05 04:45:45 -04:00
Anso 5f1baa7522 fix: harden deploy/update concurrency and node-targeting safety (#1390)
* fix: harden deploy/update concurrency and node-targeting safety

Release stabilization for deploy/update operational safety.

Per-stack operation locking is now global. Background lifecycle paths
(scheduler auto stop/down/start/backup/update, webhook execute, Git source
auto-deploy, image auto-update, label bulk actions, fleet snapshot redeploy,
and mesh redeploy) acquire the per-node, per-stack lock through a new
StackOpLockService.runExclusive helper and skip rather than race a manual
deploy/update/rollback/backup on the same stack and node. Skips surface
honestly (a failed scheduled run, a recorded webhook failure, a per-stack
batch result, or a thrown error) instead of a silent no-op.

Update readiness and policy-bypass now run against the node captured when the
dialog opened, not the live active node, so switching nodes while a dialog is
open cannot retarget the update or the bypass retry.

Rollback readiness no longer presents a moving-tag or unpinned image as a ready
image revert. Restoring files does not revert a moving tag, so those stacks
read as partial, and the rollback success message states that the compose and
env files were restored.

* fix: lock blueprint reconcile against manual ops and correct rollback wording

Follow-up to the deploy/update safety hardening, closing two more gaps from a
verification pass.

BlueprintService.deployLocal and withdrawLocal called ComposeService directly,
so blueprint reconciliation could race a manual deploy/update/rollback/backup on
an owned stack. Both now run their compose lifecycle call through
StackOpLockService.runExclusive and skip (recorded as a failed reconcile,
retried on the next cycle) on conflict. The withdraw holds the lock across both
the compose down and the directory delete so neither races a manual operation.

The runtime rollback messages overstated recovery: a rollback restores the
compose and env files and recreates containers, but does not revert an image
behind a moving tag. The auto-rollback deploy-progress output, the recovery
panel and chip, the failure toasts, and the manual rollback route message now
state that the compose and env files were restored, with the matching OpenAPI
example and atomic-deployments doc updated.

* fix: acquire stack lock before blueprint deploy mutates compose and marker files

Local blueprint deploy wrote the compose and marker files and ran the policy
assert before acquiring the per-stack lock; the lock only wrapped the deploy
itself. A reconcile could therefore rewrite an owned stack's files while a
manual deploy/update/rollback/backup was running. The lock now wraps the whole
critical section (create, write compose, write marker, policy assert, deploy),
so on conflict nothing is written and the reconcile records a failed outcome.

Adds a test asserting a deploy under a held lock records failed, writes no
marker file, and leaves the manual lock untouched.

* fix: make remote blueprint apply atomic under the receiving node's stack lock

Remote blueprint deploy wrote the compose and marker files to the target node
via separate HTTP calls and only locked on the final deploy, so the file writes
could race a manual operation on that node. A node's operation lock is
process-local and cannot be held by the hub across HTTP calls, so the locked
create/write/deploy now runs on the receiving node.

The locked critical section is extracted into BlueprintService.applyLocalUnderLock
and exposed via POST /api/blueprints/apply-local. The hub posts the blueprint to
that endpoint in one call; the receiving node runs create + write compose+marker
+ deploy under its own per-stack lock. Older nodes without the route answer 404
and fall back to the legacy multi-call flow. The endpoint is gated by paid tier
and the same per-stack stack:edit and stack:deploy permissions as the
PUT-compose + deploy it bundles, validates the stack name, compose size, and
marker structure, and returns 409 on a lock conflict without writing anything.

Adds tests for the atomic single-call path, the 404 legacy fallback, the 409
lock-conflict mapping, the route validation and permission paths, and the
write-compose-then-marker-then-deploy ordering of the shared locked apply.

* fix(deps): bump undici to 7.28.0 to clear high-severity advisory

The frontend CI npm audit gate (--audit-level=high) failed on a transitive
undici 7.25.0 (a dev-only dependency via jsdom): TLS certificate validation
bypass (GHSA-vmh5-mc38-953g) and cross-user cache information disclosure
(GHSA-pr7r-676h-xcf6). Bumping undici within jsdom's existing ^7.25.0 range to
7.28.0 clears the high-severity advisory and unblocks the frontend job. Lockfile
only; no direct dependency or source change.
2026-06-18 13:38:19 -04:00
Anso 584cda7182 fix(auto-update): label same-tag rebuilds as 'Rebuild available' instead of '10.11 -> 10.11' (#766)
When a registry pushes a new build of an image at the same tag (digest
changes, tag does not), the preview service set next_tag to the same
string as current_tag and the readiness view rendered '10.11 -> 10.11',
which reads as a UI bug.

Add an update_kind field to UpdatePreviewSummary that distinguishes:
- 'tag'    - at least one image has a strictly newer tag
- 'digest' - the only updates available are same-tag rebuilds
- 'none'   - nothing to apply

The frontend now branches on update_kind and renders 'Rebuild available'
next to the current tag for the digest case, leaving the version-arrow
diff for genuine tag bumps.

Three new buildSummary cases lock in the kind classification.
2026-04-24 23:37:32 -04:00
Anso 95278843cf feat(schedules): next-24h timeline + merge auto-update into schedules (#681)
* feat(backend): add stack update-preview endpoint for readiness board

Adds GET /api/stacks/:stackName/update-preview that returns per-image
semver diff, bump classification, and a stack-level summary powering
the Auto-Update readiness board.

- New UpdatePreviewService parses compose images, inspects local
  digests, fetches remote digests and tag lists, and finds the
  highest compatible semver tag.
- Major bumps are flagged blocked until human review; unknown bumps
  rank below real semver so they cannot mask a major.
- Rollback target is reconstructed through parseImageRef to preserve
  registry ports and drop the Docker Hub library/ prefix.
- Registry helpers (httpGet, auth token, digest, tag list, ref parse)
  are extracted into registry-api.ts and shared with ImageUpdateService.
- 28 Vitest cases cover parse, selection, bump math, digest rebuilds,
  blocked policy, and rollback target construction.

* feat(schedules): next-24h timeline, merge auto-update crud, add readiness board

Replace the flat task table with a Timeline view as the default, showing the
next 24 hours of scheduled work across four lanes (Restart, Update, Scan,
Prune) with a live now rail and per-firing pills. The All tasks tab preserves
the existing CRUD surface.

Merge Auto-update Stack into Schedules as a first-class action and replace the
standalone Auto-Update Policies view with a per-stack Readiness board that
surfaces version diffs, risk tags, changelog previews, and rollback targets
sourced from the stack update-preview endpoint.
2026-04-18 17:48:02 -04:00