Commit Graph

167 Commits

Author SHA1 Message Date
Anso 1702dabb7a fix(fleet): add auth middleware, input validation, and design system compliance (#536)
Add authMiddleware to all 13 fleet endpoints that were previously
accessible without authentication. Add NaN validation for parseInt
params, stackName validation on snapshot restore, and description
length cap on snapshot creation. Clean up updateTracker entries on
node deletion to prevent memory leaks.

Replace hardcoded colors with design system tokens, swap Select for
Combobox, replace overflow-y-auto with ScrollArea, fix card styling
(shadow-card-bevel, border tokens). Fix stale container data by
always refetching on stack expand with a loading guard against
concurrent requests.

Add operational logging for state-changing fleet operations and
diagnostic logging gated behind Developer Mode. Add 20 fleet tests
covering auth enforcement, input validation, tier gating, and
snapshot CRUD lifecycle.
2026-04-12 20:28:18 -04:00
sencho-token-app[bot] bbaee7f7f0 docs: refresh screenshots (#535)
Co-authored-by: AnsoCode <18150933+AnsoCode@users.noreply.github.com>
2026-04-13 00:00:27 +00:00
Anso cd3d7b23be feat(app-store): add port conflict indicator to deploy sheet (#533)
Show a pulsating warning dot next to host ports that are already in
use by a running container. Hovering over the dot reveals which
Sencho-managed stack or external app occupies the port.

Adds GET /api/ports/in-use endpoint that returns a map of bound host
ports with ownership info, and a getPortsInUse method on
DockerController that reuses the existing container-to-stack
resolution logic.
2026-04-12 19:57:38 -04:00
sencho-token-app[bot] 7a2099b56e docs: refresh screenshots (#532)
Co-authored-by: AnsoCode <18150933+AnsoCode@users.noreply.github.com>
2026-04-12 23:32:37 +00:00
sencho-token-app[bot] 36f2e8a7bc docs: refresh screenshots (#529)
Co-authored-by: AnsoCode <18150933+AnsoCode@users.noreply.github.com>
2026-04-12 19:37:01 +00:00
Anso 4909c35e50 fix(resources): harden Resource Explorer with auth, validation, design, and UX fixes (#527)
- Sanitize error messages in all delete/prune/create/inspect endpoints
  to prevent Docker internals from leaking to the frontend
- Add CIDR, IPv4, and Docker resource ID input validation
- Add requirePaid gate to network topology endpoint
- Add invalidateNodeCaches after image/volume/network mutations
- Fix design system violations: card borders, destructive button variant,
  visible DialogDescription, overflow-auto replaced with ScrollArea,
  hardcoded Tailwind colors replaced with tokens
- Gate purge button behind isAdmin to prevent silent 403s
- Fix shared inspect loading state to be per-network-row
- Parse error response bodies for meaningful toast messages
- Add clipboard API fallback for non-HTTPS contexts
- Render Options section in network inspect sheet
- Add operational and diagnostic logging for resource operations
- Extend validation and DockerController test suites
- Update docs with Options field in network inspect
2026-04-12 15:31:35 -04:00
sencho-token-app[bot] a5cb316ca9 docs: refresh screenshots (#525)
Co-authored-by: AnsoCode <18150933+AnsoCode@users.noreply.github.com>
2026-04-12 18:33:50 +00:00
sencho-token-app[bot] 8e91f91622 docs: refresh screenshots (#522)
Co-authored-by: AnsoCode <18150933+AnsoCode@users.noreply.github.com>
2026-04-12 09:48:00 +00:00
sencho-token-app[bot] 3ad1ab5c84 docs: refresh screenshots (#519)
Co-authored-by: AnsoCode <18150933+AnsoCode@users.noreply.github.com>
2026-04-12 08:28:47 +00:00
Anso 9db97107aa fix(dashboard): harden real-time dashboard with bug fixes and design compliance (#517)
* fix(dashboard): harden real-time dashboard with bug fixes and design compliance

- Fix host alert spam: add 5-minute cooldown for CPU/RAM/disk threshold
  alerts, preventing duplicate notifications every 30s during sustained
  breaches. Extract shared dispatchWithCooldown helper (also used by
  Docker janitor alerts).
- Fix memory metric inflation: subtract filesystem cache from stored
  memory_mb values, matching the existing calculateMemoryPercent logic.
- Fix crash detection reliability: replace fragile 'seconds ago' string
  matching with a tracked Set of alerted container IDs. Containers are
  only alerted once per crash event, with automatic cleanup when they
  start running again or after a 1-hour TTL.
- Fix health status bar: exited containers now trigger 'degraded' state
  independently of unread error notifications.
- Fix CPU chart Y-axis: auto-scale when aggregate container CPU exceeds
  100% instead of silently clipping at the hardcoded domain ceiling.
- Fix grammar: 'actives' to 'active' in container count label.
- Add shadow-card-bevel to all dashboard cards per design system.
- Update dashboard docs to reflect revised health status thresholds.

* test(dashboard): update monitor service tests for new alert signatures

- Add container Id fields to crash detection test fixtures
- Update host alert assertions to match dispatchWithCooldown 3-arg call
- Fix unhealthy container test to use State: 'unhealthy' instead of
  State: 'running' (running containers are now skipped in crash detect)
2026-04-12 04:25:41 -04:00
sencho-token-app[bot] c7f041957c docs: refresh screenshots (#516)
Co-authored-by: AnsoCode <18150933+AnsoCode@users.noreply.github.com>
2026-04-12 07:22:13 +00:00
Anso 4950cd0bd0 fix(updates): scan all filesystem stacks for image updates (#514)
ImageUpdateService previously discovered stacks by iterating Docker
containers, which meant stacks without containers (e.g. after
docker compose down) were silently excluded from update checks.

Switch to a hybrid discovery approach: enumerate stacks from the
filesystem via FileSystemService.getStacks(), parse compose files
for image refs with .env variable resolution, then augment with
container-based image discovery for running stacks.

Also cleans up stale stack_update_status entries when stacks are
deleted or no longer exist on disk, and replaces the plain
update-available tooltip with an animated cursor follow pattern.
2026-04-12 03:17:37 -04:00
sencho-token-app[bot] f26ac16bc0 docs: refresh screenshots (#513)
Co-authored-by: AnsoCode <18150933+AnsoCode@users.noreply.github.com>
2026-04-12 06:29:22 +00:00
sencho-token-app[bot] 0dc996264d docs: refresh screenshots (#512)
Co-authored-by: AnsoCode <18150933+AnsoCode@users.noreply.github.com>
2026-04-12 06:25:55 +00:00
sencho-token-app[bot] d62f6244b2 docs: refresh screenshots (#508)
Co-authored-by: AnsoCode <18150933+AnsoCode@users.noreply.github.com>
2026-04-12 05:17:18 +00:00
sencho-token-app[bot] f3fcf924ad docs: refresh screenshots (#503)
Co-authored-by: AnsoCode <18150933+AnsoCode@users.noreply.github.com>
2026-04-11 01:39:22 +00:00
sencho-token-app[bot] f33c12fb36 docs: refresh screenshots (#500)
Co-authored-by: AnsoCode <18150933+AnsoCode@users.noreply.github.com>
2026-04-11 00:44:03 +00:00
sencho-token-app[bot] 9a861f0a76 docs: refresh screenshots (#497)
Co-authored-by: AnsoCode <18150933+AnsoCode@users.noreply.github.com>
2026-04-10 21:41:15 +00:00
sencho-token-app[bot] fea8b98cb5 docs: refresh screenshots (#494)
Co-authored-by: AnsoCode <18150933+AnsoCode@users.noreply.github.com>
2026-04-10 20:39:05 +00:00
Anso e654d92d24 docs: refresh screenshots (#485) 2026-04-10 14:04:50 -04:00
Anso bc3b827b20 docs: refresh screenshots (#474) 2026-04-10 12:03:27 -04:00
Anso c09673d8bb docs: refresh screenshots (#471) 2026-04-10 10:07:49 -04:00
Anso 9b4b6c2f1e docs: refresh screenshots (#462) 2026-04-09 19:17:27 -04:00
Anso 7b6cd79e92 docs: refresh screenshots (#459) 2026-04-09 13:30:56 -04:00
Anso 1d883f4675 docs: refresh screenshots (#456) 2026-04-09 12:55:56 -04:00
Anso 53d9316f20 docs: refresh screenshots (#453) 2026-04-09 10:06:17 -04:00
Anso cc23732727 docs(topology): update network topology docs with new features and screenshots (#450)
Document dagre auto-layout, enriched container nodes, click-to-logs,
and system network toggle. Refresh screenshots to show current UI.
2026-04-08 21:42:40 -04:00
Anso 5d6c87f3b0 docs: refresh screenshots (#449) 2026-04-08 20:17:55 -04:00
Anso 1524b64847 docs: refresh screenshots (#446) 2026-04-08 17:16:08 -04:00
Anso 8d5bde6a8c docs: refresh screenshots (#443) 2026-04-08 15:11:58 -04:00
Anso 6fff2c2d35 fix(fleet): resolve self-update compose file access and improve completion detection (#441)
The self-update feature failed on remote nodes because SelfUpdateService
ran `docker compose -f <host_path>` inside the container, where the host
compose file path does not exist. The fix splits the update into two
steps: (1) pull the latest image directly via `docker pull`, and (2)
spawn a short-lived helper container that mounts the compose directory
from the host and runs `docker compose up --force-recreate`.

Additional changes:
- Use execFileSync/execFile with argument arrays instead of shell strings
  to eliminate shell injection surface from Docker label values
- Add Signal 4 completion detection: mark update as completed when the
  remote version matches the gateway version (with 15s elapsed guard)
- Extend early failure heuristic from 90s to 3 minutes for slow pulls
- Distinguish "node unreachable" from "node lacks self-update capability"
  in error messages; use silent skip in update-all to avoid res crashes
- Add requireAdmin guard to POST /api/system/update
- Handle comma-separated compose config file paths (multiple -f flags)
- Update fleet docs with self-update mechanism, troubleshooting entries
2026-04-08 14:59:11 -04:00
Anso 24011aea9e docs: refresh screenshots (#430) 2026-04-08 10:58:38 -04:00
Anso 0d6e5367b1 docs: refresh screenshots (#422) 2026-04-07 01:57:04 -04:00
Anso c79f58d8dd docs: refresh screenshots (#415) 2026-04-06 23:16:33 -04:00
Anso 7f1d32f616 docs: refresh screenshots (#412) 2026-04-06 22:14:19 -04:00
Anso 8ba4532995 fix(fleet): resolve version detection using package.json over stale generated constant (#410)
* fix(fleet): resolve stuck update states and improve update UX

The fleet node update flow had several bugs: the in-memory update tracker
never cleared terminal states (timeout, failed, completed), leaving nodes
permanently stuck with no way to retry or dismiss. The Recheck button
only re-fetched stale state without clearing it, and the POST trigger
rejected retries with 409 even after timeout.

Backend fixes:
- Add DELETE endpoints (single node + batch) to clear tracker entries
- Fix 409 race: detect expired timeouts and clear terminal states before
  re-triggering
- Populate error messages in the tracker for timeouts and failures
- Include error field in the update-status API response
- Auto-expire completed entries after 60 seconds

Frontend fixes:
- Add retry (RotateCcw) and dismiss (X) buttons on failed/timed-out badges
- Show error details via animated cursor hover (CursorFollow pattern)
- Recheck button now batch-clears all terminal states before fetching
- Recheck shows loading spinner and disables while checking
- Extract NodeCardProps interface for readability

* fix(fleet): detect update completion via process start time

Remote nodes that cannot report their version (e.g. older builds)
caused updates to always time out because completion detection
relied solely on version comparison. The gateway now tracks the
remote node's process start time from /api/meta and detects
container restarts by comparing it across polls.

Also extracts a createTracker() factory to eliminate repeated
object construction across 5 call sites.

* docs: add troubleshooting for first-update timeout on old nodes

Adds a new troubleshooting entry explaining why the first remote
update on nodes running pre-v0.40.0 always times out (neither
version nor process start time can be detected). Documents the
fix: dismiss, recheck, and confirm the node updated.

Also adds a screenshot of the timed-out state with retry/dismiss
buttons to the remote updates feature page.

* fix(fleet): detect update completion via offline detection and error reporting

The update completion detection relied on version change and process
start time, both of which fail on nodes running older Sencho versions
that report "unknown" and lack the startedAt field. This caused every
update to time out after 5 minutes.

Add three-signal detection: version change, process restart (startedAt),
and offline/online detection (node went unreachable during update and
came back). Also add a 90-second early failure heuristic for when the
remote image pull fails silently, and surface pull errors from
SelfUpdateService via /api/meta so the gateway can report them
immediately.

* fix(deps): bump vite to 8.0.5 to resolve high severity vulnerabilities

Fixes GHSA-4w7w-66w2-5vf9, GHSA-v2wj-q39q-566r, GHSA-p9ff-h696-f583.

* fix(deps): bump vite in backend lockfile to resolve audit failures

Vitest pulls in vite as a transitive dependency. Bumps to 8.0.5.

* fix(fleet): resolve version detection using package.json over stale generated constant

resolveVersion() previously returned the build-time SENCHO_VERSION
constant without checking the root package.json. When a branch fell
behind a release-please version bump, the generated constant was stale,
causing remote nodes to show "unknown" version and false "Update
available" badges. The function now walks up to the root package.json
first (authoritative source) and falls back to the generated constant
only if the walk fails.
2026-04-06 22:05:06 -04:00
Anso fb998fd0c0 docs: refresh screenshots (#409) 2026-04-06 20:30:16 -04:00
Anso cc2da99d6f fix(fleet): resolve stuck update states and improve detection (#405)
* fix(fleet): resolve stuck update states and improve update UX

The fleet node update flow had several bugs: the in-memory update tracker
never cleared terminal states (timeout, failed, completed), leaving nodes
permanently stuck with no way to retry or dismiss. The Recheck button
only re-fetched stale state without clearing it, and the POST trigger
rejected retries with 409 even after timeout.

Backend fixes:
- Add DELETE endpoints (single node + batch) to clear tracker entries
- Fix 409 race: detect expired timeouts and clear terminal states before
  re-triggering
- Populate error messages in the tracker for timeouts and failures
- Include error field in the update-status API response
- Auto-expire completed entries after 60 seconds

Frontend fixes:
- Add retry (RotateCcw) and dismiss (X) buttons on failed/timed-out badges
- Show error details via animated cursor hover (CursorFollow pattern)
- Recheck button now batch-clears all terminal states before fetching
- Recheck shows loading spinner and disables while checking
- Extract NodeCardProps interface for readability

* fix(fleet): detect update completion via process start time

Remote nodes that cannot report their version (e.g. older builds)
caused updates to always time out because completion detection
relied solely on version comparison. The gateway now tracks the
remote node's process start time from /api/meta and detects
container restarts by comparing it across polls.

Also extracts a createTracker() factory to eliminate repeated
object construction across 5 call sites.

* docs: add troubleshooting for first-update timeout on old nodes

Adds a new troubleshooting entry explaining why the first remote
update on nodes running pre-v0.40.0 always times out (neither
version nor process start time can be detected). Documents the
fix: dismiss, recheck, and confirm the node updated.

Also adds a screenshot of the timed-out state with retry/dismiss
buttons to the remote updates feature page.

* fix(fleet): detect update completion via offline detection and error reporting

The update completion detection relied on version change and process
start time, both of which fail on nodes running older Sencho versions
that report "unknown" and lack the startedAt field. This caused every
update to time out after 5 minutes.

Add three-signal detection: version change, process restart (startedAt),
and offline/online detection (node went unreachable during update and
came back). Also add a 90-second early failure heuristic for when the
remote image pull fails silently, and surface pull errors from
SelfUpdateService via /api/meta so the gateway can report them
immediately.

* fix(deps): bump vite to 8.0.5 to resolve high severity vulnerabilities

Fixes GHSA-4w7w-66w2-5vf9, GHSA-v2wj-q39q-566r, GHSA-p9ff-h696-f583.

* fix(deps): bump vite in backend lockfile to resolve audit failures

Vitest pulls in vite as a transitive dependency. Bumps to 8.0.5.
2026-04-06 20:09:55 -04:00
Anso 9e5c8307bc docs: refresh screenshots (#404) 2026-04-06 03:11:11 -04:00
Anso a55d1245f8 fix(fleet): resolve version detection pipeline for Docker builds (#402)
* fix(fleet): resolve version detection pipeline for Docker builds

The Dockerfile backend-builder stage was missing a COPY of the root
package.json, causing generate-version.js to fall back to "0.0.0-dev"
at build time. At runtime, the filesystem walk also failed (root
package.json not in the final image), producing the string "unknown"
which the frontend rendered as "vunknown".

Changes:
- Dockerfile: copy root package.json into backend-builder stage
- CapabilityRegistry: return null (not "unknown") for unresolvable
  versions; add isValidVersion() type guard; normalize remote meta
  responses to strip "unknown"/"0.0.0-dev" sentinel values
- Fleet endpoints: hoist gateway version validation outside per-node
  loops; treat unresolvable remote versions as "potentially outdated"
  instead of silently marking them up to date
- FleetView: guard all version display points (card badge, update
  button, gateway label, modal columns) via shared formatVersion()
- EditorLayout, CapabilityGate: use shared isValidVersion utility
- New frontend/src/lib/version.ts shared utility
- Docs: add troubleshooting section for version display edge cases
- Screenshots: updated Fleet Overview and Node Updates modal

* docs: update fleet node updates screenshot with live remote node
2026-04-06 03:03:56 -04:00
Anso 9002ff95e1 docs: refresh screenshots (#401) 2026-04-06 01:52:57 -04:00
Anso cd3d670ab4 docs: refresh screenshots (#398) 2026-04-06 00:51:31 -04:00
Anso 89da454db4 docs: refresh screenshots (#395) 2026-04-05 23:46:40 -04:00
Anso 35a4306104 docs: refresh screenshots (#394) 2026-04-05 23:42:07 -04:00
Anso 52c0196cd3 docs: refresh screenshots (#390) 2026-04-05 22:55:49 -04:00
Anso 2bc9ed8273 docs: refresh screenshots (#387) 2026-04-05 21:58:19 -04:00
Anso a613498c2c docs: refresh screenshots (#384) 2026-04-05 20:07:11 -04:00
Anso ebb696e382 docs: refresh screenshots (#381) 2026-04-05 18:53:31 -04:00
Anso 4163af2ee7 docs: refresh screenshots (#378) 2026-04-05 17:34:32 -04:00
Anso f841c402b2 fix(licensing): resolve Admiral variant detection and lifetime license handling (#376)
* fix(licensing): resolve Admiral variant detection and lifetime license handling

The Lemon Squeezy variant name for Admiral licenses contains "Admiral"
(not "Team"), but getVariant() only checked for "team" and "personal".
This caused Admiral licenses to be misidentified as Skipper, locking all
Admiral-exclusive features.

- Map "admiral" variant names to internal "team" value, "skipper" to "personal"
- Add isLifetime field to LicenseInfo API response
- Hide "Manage Subscription" button for lifetime licenses (no billing portal)
- Show "Duration: Lifetime" instead of empty renewal date
- Hide upgrade cards for active Admiral users
- Add 23 unit tests covering variant resolution, tier computation, and lifetime detection
- Add troubleshooting entries for wrong tier label, locked features, and billing portal errors

* fix(licensing): address code review findings

- Fix nested ternary in LicenseSection JSX; restore conditional rendering
  to avoid showing an empty "N/A" row for non-subscription states
- Clean up test file: use shared svc variable, remove redundant comments,
  add trialDaysRemaining assertions, rename describe block
2026-04-05 07:00:21 -04:00